Portal Factory Build Board

Portal Factory · the full board

Portal Factory Build Board

The path from tonight's code to a factory that is ready to sell — what is done, what is running, what is next, and what needs a person. For Alyssa and Dave.

updated ET

Reader settings

Text size
A Claude Code session is running Afternoon session, Sept 12 (Queue 6) — Alyssa released the boards' fix lists and ruled a fixed night sky for the games; the designer and the critic are on both now.
started
Sep 12, 1:00 PM
runs until
Sep 12, 9:47 PM
checks in every
13 min
spent since the last brief
$0.01 (Console — Console, Sept 12 UTC: $0.01 (includes last night's $0.002); this session's own ledger $0; cap $5)
last heard from
Sep 12, 3:14 PM

Progress to sellable

54%of the path to sellable
6 done 2 in progress 2 waiting on a decision 1 blocked 2 not started
  • B1 reviewed & timed
  • B2 end-to-end walk
  • B3 live domain
  • B4 a new portal
  • B5 first payment

Done counts 1, in progress counts ½. Targets: a demo to sell from by September 24; a factory, not a demo, by October 1. Revenue floor the plan is built to clear: $345,000 a year.

Now · next · needs a person

In progress

  • Two design loops released by Alyssa: the three boards' remaining fix lists (one designer for all three, so they match) and the arcade with its fixed night sky.
  • The MK11 regeneration waits for her mark-up and her words tonight.

Next in line

  • Alyssa: walk the fixture portal as learner then manager, and time it
  • Regenerate MK11 Onboarding once approved — through the pipeline's Batch API (about $5–6, not $10) — and compare with shape:check against the export
  • Reconcile 9 stale 'running' generation runs left by earlier sessions
  • Wire the design critic into the review loop (only Week-2 design item left)

Needs a person

  • AlyssaMark up the 33-vs-60 side-by-side (KEEP / MERGE / CUT / MOVE), then 'Regenerate MK11 through the pipeline' — $5–6, stops itself at ~$1.50 if the shape is wrong.
  • AlyssaTonight by 8:55 PM: the 25-minute homework in your email and the Google Doc (mark-up, four sentences to say, one click). Then paste the start prompt at 10 PM. No mark-up = $0 items only tonight.
  • AlyssaTwo minutes: open the practice coach in Chrome, press the mic, say a line — tell me if you heard the customer.
  • AlyssaSay 'Apply it' for the one-line protected leak-test fix (second brief).
  • AlyssaFifteen minutes, $0: an Anthropic Organization + Admin API key so spend reconciles itself.
  • AlyssaNameservers → Cloudflare and the Fly login: put it back on the list, park Week 3, or hand it to Dave — say which.

The 13 steps to sellable

  1. Step 1: Zero-cost draft loop

    Done Claude

    Done when A shape proposal is produced and gated with $0 API spend

    Shape gate lands on every outline before a dollar is spent.

  2. Step 2: Fix the ruler

    Done Claude

    Done when A 60×2-minute run refuses; each gate proven to fail first

    Per-archetype caps, 5-minute floor, front matter as furniture, phase balance. Three wrong versions of one rule caught and removed the same night.

  3. Step 3: Persist phase gates; chips instead of counts

    Done Claude

    Done when Gate subtitles visible in a browser

    Seen in a real browser, light and dark, 22:39 Sept 10.

  4. Step 4: M60 + spend-cap metering

    Done Claude

    Done when 4 tests green; coach runs on MK11

    Practice cap now meters only roleplay; 18-turn target with 4 turns of grace.

  5. Step 5: Regenerate MK11 Onboarding against the new ruler

    Waiting on a decision Claude

    Done when ~30 lessons, ≥5 min each, opens with an orientation

    Costs about $5–6 through the Batch API (the scripts had been paying full price). Needs Alyssa's go-ahead — it replaces the current 60 draft lessons, which are exported.

  6. Step 6: You review it — and we time it

    Waiting on a decision Alyssa

    Done when BENCHMARK 1 — the review-hours number exists

    Follows item 5.

  7. Step 7: Agent roles + morning brief

    Done Claude

    Done when 8 role files; a brief you actually read

    8 role files exist; the morning brief for Sept 11 was written and handed over (docs/briefs/2026-09-11.md).

  8. Step 8: Factory UI: one bucket per client

    Done Claude

    Done when You can run a portal without me

    Test tenants hidden by default; /clients shows every client with its stage and the next step, and each client has a bucket page with counts, deep links and each course's minutes per lesson.

  9. Step 9: Design pass: brand on every surface; critic loop

    In progress Claude

    Done when No page ships default-styled

    Learner and manager surfaces on tokens; the design critic is now a standing step in the loop (ran four times Sept 12, sent two items back). Not yet one header treatment across every learner page — the critic's open note.

  10. Step 10: Manager + learner end-to-end walk

    In progress Claude, then Alyssa

    Done when BENCHMARK 2 — a learner completes a lesson; a manager sees it

    Learner walk passed Sept 6. Manager side walked tonight: a completion shows on the roster and the learner page through the real doors. The benchmark itself is Alyssa walking it.

  11. Step 11: Deploy to the real domain

    Blocked Alyssa, then Claude

    Done when BENCHMARK 3 — both hostnames resolve

    Blocked on nameservers → Cloudflare and a Fly login. Both need Alyssa at a keyboard.

  12. Step 12: Generate a NEW portal end-to-end as proof

    Not started Claude

    Done when BENCHMARK 4 — a second client-ready portal from a folder

  13. Step 13: First paying client

    Not started Alyssa

    Done when BENCHMARK 5 — payment

Week by week

Week 1 — Build the ruler Days 1–5

  • Done: Zero-cost draft loop; per-archetype gates; minutes floor; furniture; phase balance
  • Done: Persist phase gates; backfill MK11's six; chips replace counts
  • Done: M60; spend-cap metering
  • Not started: Regenerate MK11 against the ruler (~$10, needs go-ahead)
  • Not started: Alyssa reviews it; we time it — BENCHMARK 1

Week 2 — Make it presentable Days 6–10

  • Done: Factory UI: one bucket per client; hide test tenants
  • Done: 8 agent role files; the morning brief
  • In progress: Design pass: brand tokens on every surface; Resources as cards; critic loop
  • In progress: Manager + learner end-to-end walk — BENCHMARK 2 (both halves reachable; Alyssa's walk pending)

Week 3 — Prove it in production Days 11–15

  • Not started: Nameservers, certificate, deploy (needs Alyssa)
  • Not started: Both hostnames resolve; Cloudflare Access on the Factory — BENCHMARK 3
  • Not started: A brand-new portal end-to-end from a folder — BENCHMARK 4
  • Not started: Buffer: defect tail, module edit form, object-storage backup

Numbers

Tests passing 1,400 0 failing · 2 skipped · last full run source: npm run test log, 2026-09-12T07:01:49.754Z
Security suite 95 / 95 tenant-isolation and leak checks source: npm run test:leak log, 2026-09-12T01:44:00.246Z
Commits, last 24 hours 42 473 on every branch all time · 0 unpushed source: git log
Ledgered model spend, last 24 hours $0.00 2 logged calls · loop cap $0.00 · the Console is the bill, this is what the ledger wrote down source: generation_runs, priced by delivery
Real courses that pass the ruler 1 of 6 refused: MK11 · Product Knowledge (Op · MK11 · Product Knowledge · MK11 Onboarding +2 source: shape gate over every non-test course in the database
Runs stuck at 'running' 9 older than a day · never finished · the owner decides whether to mark them abandoned source: pipeline_runs
What MK11 as built cost $19.02 60 lessons · 459 calls · priced by delivery · through the Batch API it would be about half source: generation_runs for that course, re-priced
Spend on courses never published 35% $28.41 of $80.37 in the ledger · the number the ruler exists to bring down source: generation_runs joined to courses.status
Practice call, per turn $0.01 buyer + observer per turn, 140 calls measured · 18-turn target ≈ $0.14 a call source: generation_runs, roleplay stages

Every number here is measured when the board is rebuilt and names its source. Nothing in this section is typed by hand.

Hours on the project

Claude Code on the factory139.7 hmeasured from 14 session transcripts on this Mac since 2026-07-12 · this week 20.1 h
Of that, time around commits93.2 hsince 2026-08-01 · 473 commits in 83 stretches · this week 12.3 h
Claude Code, every project214.4 hall sessions on this Mac, all products, since April
Alyssa, all in300 hself-reported Sept 11: over 300 hours across every chat, Cowork session and Claude Code run — the clocks below cannot see the chats on claude.ai, so this is the whole

Time around commits, by week

07/2711.3 h
08/0310.8 h
08/1721.5 h
08/249.5 h
08/3127.8 h
09/0712.3 h

Two clocks, both measured on this Mac, both floors. Sessions: every Claude Code transcript carries timestamps; gaps over 30 minutes end a stretch, and a session counts for the factory when its content is about the factory. Commits: the same idea applied to git, a subset of the first. Neither can see chats on claude.ai, Cowork, reading, Google Docs, or the hand-built July portal — that is why the fourth tile is Alyssa's own number, and why it is the biggest.

Total project spend

Project total spent$247.42$298.20 paid in − $50.78 still on the account · tax included
Paid into the API account$298.209 purchases, August 20 – September 7 · $280.00 of credit + $18.20 tax
Left on the account$50.78console balance, read Sep 12, 11:52 AM ET · 18% of credit remains
Console, last 30 days$262.99portal-factory $253.58 · console (Playground) $9.40 · the factory's own ledger caught $105.45 of it · this week the console says $38.05, the ledger $1.26

By month (the console; the ledger beside it)

Aug$134.95 · ledger $32.80
Sep$128.04 · ledger $47.57

By kind of work

outlining$46.80
quizzes & exams$12.55
writing lessons$12.51
other$5.47
grounding checks$2.09
practice calls$0.95

API account top-ups (Alyssa)

  • August 20$31.95
  • August 20$21.30
  • August 21$53.25
  • August 22$21.30
  • August 25$53.25
  • September 4$21.30
  • September 4$21.30
  • September 5$21.30
  • September 7$53.25

Where the difference went (console minus ledger)

  • Aug 19–21: $68.81 on the factory key before the ledger existed — and the two database resets erased the rows that did exist.
  • A nightly cloud eval ran every day at 09:00 UTC with the real key, on a database that dies with the run, and FAILED every day since Sep 5 — about $1–4 a day, plus five push-triggered runs on Sep 6–7. Switched off Sept 11; re-enable only after a green on-demand run.
  • Sep 5–7 overnight work under-recorded: interactive calls priced at the batch rate (fixed), rows deleted at test teardown, 147 failed calls with no token count.
  • Almost no batch discount was ever earned: $2 off $260 — the pipeline ran interactively nearly every time. The stage scripts now default to the Batch API.

The console is the bill; the ledger is what the factory managed to write down. Money Alyssa put into the API account (credit purchases, tax included), entered by her on Sept 11. Ledger spend is measured every time the board is rebuilt (last measured Sep 12, 3:14 PM ET). Subscriptions, the domain and hosting are not in any ledger the factory can read — they are Alyssa's to enter, and the total above adds them in when she does.

Session briefs and Alyssa's notes

What Claude reported at the end of each session, Alyssa's notes on that report, and the AI team's answer — so anyone reading can see what was done and what the owner said about it. Newest first.

Daytime brief — Sept 12 (Queue 5) September 12 · Claude (cfo + plain-english)

What Claude did

  • Spent $0.00 this session (ledger: no rows; Console Sept 12 UTC $0.01 = last night's proof plus ~0.8¢ unexplained, too small to chase). Credits $50.78 after Alyssa's top-up.
  • Done: the MK11 'who we are' source (audited, ingested); survey allowed links; brand steward + measured palette tool; edit lesson text + update with new material; 27 screenshots in the gallery; the simple board at /simple.
  • Stopped: the arcade fix list after two design-critic strikes — the rest is 1c and needs a ruling (fixed night sky vs brand colours).
  • Needs you: 'Regenerate MK11' after mark-up; 'Apply it'; org + Admin key; delete the eval secret; nameservers/Fly; tutor prompt (11b); 'Do 1c'; one click on the office-address finding; add https://mk11team.com/ to MK11's survey links.

Morning brief, Sept 11–12 overnight (Queue 4) September 12 · plain-english over the cfo report, 3:10 AM

What Claude did

  • Cost: $0.002 (ledger, two proof calls); Console 30-day total unchanged at $262.98 when re-read at 3:05 AM; credits $30.79 at the last full read.
  • Built: Chill Study Mode (six games on the client's own content), the progress rail, study time, Where you stand, leaderboard, manager depth, messages both ways, announcements, voice on the practice coach, ten measured palette presets, find-in-your-documents + Listen, money plumbing, stale runs closed, the design critic as a standing step, the regeneration run plan, the 33-vs-60 side-by-side.
  • Blocked for Alyssa: the arcade's dark-mode contrast fix list (1b); the paid outline (rule 18 at 30–40); the tutor's written answer (prompt + embeddings).
  • Seven decisions and six red flags are in the brief, each with options and a recommendation.

Alyssa's notes, and the answers

On: NEEDS YOU 1 — rule 18

Alyssa: Approve rule 18 at 30–40 lessons
Claude: Done at 8:32 AM: the rule is in the outline prompt (as rule 20; 18 and 19 already existed), version 12. The band itself comes from the ruler, so the prompt and the checker cannot disagree. The first real outline under it is the run plan's ~$1.50 step and waits on your mark-up and 'Regenerate MK11 through the pipeline'.
done

Morning brief, Sept 10–11 overnight + Sept 11 afternoon September 11 · vp-review (Fable), answered by Claude at 9:45 PM

What Claude did

  • Overnight: 22 items closed on $0.003 of LEDGERED spend — the ledger's number, not the bill's. See the first note below for the honest version.
  • The Build Board, the shape ruler (shape:check), the content audit, the reference snapshot, the Clients bucket, the manager walk, the nightly-eval finding.
  • Afternoon: Resources cards, the client-facing audit document, CI unit tests green on GitHub for the first time since Aug 22, the manager screens' design pass.
  • Evening, on Alyssa's notes: the paid-API guard for tests, the 30–40 lesson band, the CFO / plain-English / factory-ops roles, spend:audit, the 60-draft review page, this section.

Alyssa's notes, and the answers

On: THE ONE THING — API credit

Alyssa: It looks like about $5 was spent last night as I only have $30 left this morning. Give an honest breakdown of how it was spent and whether it was valuable or wasteful. Maybe a CFO agent that audits after every run. This spend did not show up on my dashboard; I want it live and accurate including spend. Why don't my Console numbers match the $0.003 you gave me? If Claude is misunderstanding something this small, what else is it getting wrong?
Claude: You are right, and the $0.003 was wrong in the way that matters. It was the LEDGER's number — what the factory's own code wrote down — and I had found the same night that the ledger misses money. I reported it as if it were the bill. That is the thing I must never do again, and the board now refuses to show spend from the ledger alone. What actually went out (read from your Console tonight, by day): Sept 10 $3.96, Sept 11 $4.30. Your balance went $35.08 → $30.79, a drop of $4.29, which matches. Where the $4.30 went: almost all of it was the nightly cloud eval running AGAIN at 9:26 AM on Sept 11. It authored 22 lessons and 134 quiz questions on three test corpora, then failed its own bar — a $4 rehearsal nobody watched. WASTED. The overnight build loop itself spent under about $0.30 (a few small judge calls the test suite made; the ledger saw $0.003 of that because a test's clean-up deleted the rest). Why it ran again after I said it was off: I commented out the schedule on our working branch. GitHub runs schedules from the DEFAULT branch (main), which I am not allowed to push to. So my 'off' changed nothing. Tonight I disabled the workflow at the repository level (gh workflow disable) — that switch does not depend on any branch. It has been paying about $4 a day since Sept 8 and about $1.10 a day from Aug 26 to Sept 2, failing every time: roughly $25–35 over the month. The CFO exists as of tonight (a role file in the repo) and there is a $0 command, spend:audit, that prints Console vs ledger by day and goes RED when they disagree. Run tonight it says: last 10 days, Console $126.70, ledger $47.27 — $79 the ledger never saw. That is the honest gap, in one line, every session from now on.
fixed

On: Regenerating MK11 — how do I say yes, and does opening the portal cost money?

Alyssa: What do I need to do to say yes to regenerating MK11? Is it on the portal factory or a command here? Does opening and logging into the portal cost money? Is that where the ghost budget went? If so, make the portals Cloudflare links. Is this the 60 lessons planned to ship? I think 30–40. Please combine some lessons. Are you comparing against my manual MK11 build? Can I review free drafts before the factory spends money?
Claude: Saying yes is a sentence to me here in Claude Code: 'regenerate MK11 through the pipeline'. I then run it through the Batch API (about $5–6), and I will show you the outline for free first — see below. Opening or logging into a portal costs nothing. Pages are served from your Mac's database; no model is called when someone reads a lesson, signs in, or opens the roster. The only things that spend are generation, grounding checks, the practice coach, and the evals. Your ghost budget was the nightly eval and the test suite, not the portals. Cloudflare links for the portals: not yet possible as a static page. A portal is a live app with a database (sign-in, progress, seats), so it needs a host — that is the Fly + nameservers item in NEEDS YOU. The fixture portal is reachable only on your laptop for now. The 60 lessons are NOT planned to ship. They are the drafts the factory made on Sept 4 — never published, and the ruler now refuses that shape. Your 30–40 rule is now in the code: the onboarding band is 30–40 lessons (your build has 33), with 8–12 minutes each. Your manual build IS the reference. The curriculum-architect role opens with your 6-phase, 33-lesson structure and grades phase weighting against it; shape:check measures the count and minutes; and I re-read your manual MK11 portal tonight and your screenshot of the journey. The lesson-by-lesson comparison (does the new outline put Energy 101 before the bill lessons, the way you did?) is the architect's written check before any authoring. Free review before money: yes, and it is the new order of operations. The outline (titles, phases, minutes) is the executive-function part and it is cheap — about $1 on the outline stage, or $0 to review the shape proposal. Lesson BODIES are the expensive part and come last. So: outline → you mark it up in a Google Doc → I aim the prompt at your marks → then authoring. Tonight you have the 60 drafts as a document to mark up (link in chat).
needs Alyssa

On: Pixel proof for the chips and shop CTA

Alyssa: The design is a vast improvement. Document my screenshots and yours in one Google document as a continuous before-and-after record. Give me a written description of what changed and instructions with screenshots for how each button and feature is used.
Claude: Done in two pieces, because Google Docs drops images when I import them: a Drive folder 'Portal Factory — Before & After' holds the record and the written guide; the screenshots themselves live on one page every session adds to (the 'Portal Factory Screens' link in chat). The guide walks every button on the learner side (Continue, the journey chips, Mark complete, Resources, the trainer card) and the manager side (Invite, Export CSV, search and progress filter, Reset link, Deactivate). Drop your own screenshots into the folder and they become part of the record.
done

On: shape:check — the $80 of rework

Alyssa: How do we eliminate this rework problem? I can't afford to waste this much money. Deploy as many agents as needed to monitor and fix waste. I want someone in charge of optimizing factory processes for time, quality and cost.
Claude: The $80 was three courses generated BEFORE a ruler existed; each was refused by the ruler after the money was spent. The fix is order, not effort: shape at $0 first, your review of structure for free second, authoring last, on the Batch API. That order is now written into a role whose only job is it — factory-ops (a role file in the repo) — and it will not let a paid run start until the shape has passed at $0. The CFO grades every dollar afterwards. Both run at the end of every session and before any paid run. Honest limit: a role is a set of instructions I follow, not a person who works while I am off. The guarantee is that no run starts without those two checks, in writing, in the log.
in progress

On: Manager half of Benchmark 2 — how do I walk it, and is the Build Board the factory's interface?

Alyssa: How do I do this? Provide the link, credentials and an explanation. Is this the portal factory's new interface?
Claude: Two different things. The Build Board (this page) is a progress report for you and Dave; it is not the factory. The factory's own screens run on your laptop at the control address (the Clients page), and the portals the factory makes are separate sites, one per client. The walk is on your laptop only. Open the fixture portal's learner sign-in page as the learner, complete a lesson, sign out; then open the manager sign-in page as the manager and find that learner on the roster. The addresses and the location of the credentials file are in the chat reply, not on this board, and the credentials themselves are never printed anywhere. Time yourself; that number is Benchmark 2.
needs Alyssa

On: M69 — phantom claims

Alyssa: Are these phantom calls costing us? Why are they happening? How do we stop it? Can we have an agent whose job is to explain things to me as if I were 13, watch for red flags, and brainstorm options?
Claude: Plain version: after a lesson is written, a checker (the grounding judge, a model call) reads every sentence and asks 'is this backed by the client's documents?'. It was also reading table rows and half-finished citations as if they were sentences, so it raised 159 false alarms on the pilot. Each false alarm can trigger a paid repair call, so yes, they cost money — small per call, but every run. Fixed by only sending real claims to the judge: 159 → 3 alarms, measured without a model call. The explainer role exists as of tonight (a role file in the repo): every brief goes through it before it reaches you, with a red-flag scan and options at the end.
fixed

On: Stale task files, phantom sections

Alyssa: What are the stale files, and should this be a concern? What is a phantom file, and are these the gaps costing us extra money?
Claude: Stale task files are to-do cards in the repo that were already finished but never marked done. Normal, no cost; closing them just makes the list truthful. 'Phantom sections' are different: the scanner reads the client's PDFs, and word-salad from bad scans (a page header repeated on every page, for instance) was being treated as a real section heading — so the course got sections that did not exist. That DID cost money, because each phantom section became a thin lesson somebody paid to write. Fixed by teaching the scanner what junk headings look like; 81 of 207 MK11 headings were reclassified.
done

On: audit:content — 65 gaps on MK11

Alyssa: Would it help if I re-upload the MK11 docs, or do you have them? Would more permanent memory help, rather than the context window? Could the chat data be a reference?
Claude: No need to re-upload: the five MK11 documents are in the factory's database, and I exported the whole MK11 course outside the repo as a restorable snapshot. Permanent memory already exists in three layers — the project's BUILD_LOG and THE_PLAN in the repo, a memory folder Claude Code loads every session, and every chat transcript on your Mac (the hours clock reads them). What I did tonight is make the memory more specific: your rules from this brief (Console is the only spend truth; 30–40 lessons; explain plainly; before-and-after record; live activity on the board) are written where every future session reads them first.
done

On: Ledger priced by intent — where is the money going?

Alyssa: So where is the money actually going, and why was this missed by Claude?
Claude: Where it goes, in order of size, from your Console: authoring and assessing courses (the $80 of rework), the nightly eval that failed daily, the Sept 5–7 pilot runs, and small daily judge calls. Why it was missed: the factory had its own notebook (the ledger) and I trusted the notebook over the bank statement. Three things bypassed the notebook — a GitHub workflow with your key, tests that reached the API, and scripts that called the API directly — and clean-ups deleted some entries. Tonight's fixes close two of the three for good: tests cannot reach the paid API at all now (a guard refuses the request before it leaves the machine, proven by a test), and the workflow is disabled at the repository level. The third (scripts) is the next factory-ops task.
fixed

On: Decisions — curriculum-architect

Alyssa: Can you give me a short summary of what decisions were made by the curriculum-architect?
Claude: The architect's rules, in one breath: a course opens with an orientation (why this exists, what it takes, who to ask) — never 'Lesson 1 of Module 1'; lessons are allocated by what a learner needs at that stage, not by how long each PDF is (the factory gave the commission policy 42% of MK11 because it was the longest file; you gave it 21%); document headings are not lesson titles; front matter (cover pages, sign-off forms) is furniture, not lessons; and whether an orientation is present is a judgment call, not a counter — the counter version of that rule was tried three times in one night and removed.
done

On: NEEDS YOU 1 — see the 60 drafts

Alyssa: Can I see the 60 drafted lessons as a Google Doc and review how each is organized so I can give feedback before regeneration?
Claude: Yes. A Google Doc 'MK11 — the 60 drafts, review sheet' with the numbered list (phase, title, minutes) and a mark-up key (MERGE INTO n / KEEP / CUT / MOVE TO PHASE n), and the full text of all 60 on one page linked from it. Your marks become the aim for the next outline.
done

On: NEEDS YOU 2 — approve the prompt change

Alyssa: Do I need to do this manually or do agents handle this for me?
Claude: One word from you here — 'approve rule 18' — and I apply it. Prompt files are on the protected list, so I do not change them without your word; that rule is the reason a prompt cannot drift while you sleep.
needs Alyssa

On: WORRIES — paying for failed runs

Alyssa: I don't think it's fair to pay for failed sessions. Is there any way to perfect things before we deploy them so we are not wasting money on failed tries?
Claude: Mostly yes, and it is the same answer three ways. (1) Check everything that CAN be checked for free before a paid call — the shape ruler, the schema checks, the corpus filter all run at $0 now. (2) Never let a test or a schedule spend money; both are closed tonight. (3) When a run must be paid for, run it on the Batch API at half price and change one thing at a time so a failure teaches something. What cannot be made free: the model has to write the lesson at some point, and a bad lesson costs the same as a good one. The 2-strike rule keeps a bad run from becoming five bad runs.
in progress

On: NEEDS YOU 9 — the leak job on GitHub

Alyssa: How do I fix the leak job, and what does this mean? Was our data leaked? Is our code compromised? How did this happen?
Claude: Nothing leaked and nothing is compromised. 'Leak job' is our name for the test that PROVES no client can ever see another client's data. It has two halves: one shows a cross-client reference is rejected (passes everywhere), the other shows a same-client reference is accepted. That second half picks 'any real row' from the database and, on GitHub's freshly built test database, happened to pick a learner who was already enrolled in the course it tried — a duplicate, so the database said no. On your Mac the rows are in a different order and it passes. It is a fixture collision, the security rule held. The fix is one line in a protected file; the exact change is written up as a proposal in the repo. Say 'apply it' and I will.
needs Alyssa

On: Mid-session: the board showed no session running

Alyssa: I noticed the dashboard shows Claude as live and active for this project. Please count these hours towards the time working on Claude and also update the live dashboard to show activity live while Claude works.
Claude: The board was still showing the afternoon loop as closed while I worked on your notes — a session I had not marked. Fixed at 9:36 PM: marked running, and the pulse now says when it last heard from the session; if that goes past 45 minutes while marked running, it says so instead of pretending. The hours clock reads every Claude Code transcript on this Mac, this conversation included, so these hours are counted on the next rebuild.
fixed

Reports from each desk

VP Review

on trackFable

Fourteen items closed and pushed tonight; two premises of my own corrected on the record — and, asked where the money went, I found it: a nightly cloud eval spending real credit every day and failing every day, invisible to the ledger. Off as of tonight. Tomorrow's ask is still one decision: the $10 regeneration — but now with the real per-course cost known.

Curriculum Architect

blocked or failingOpus

The ruler now refuses 60×2-minute courses — and, run on real data tonight, refuses every MK11 course the factory has ever built. Aiming the prompt (rule 18) is the approval that turns refusal into a good outline.

Gate Engineer

on trackOpus

Every new gate failed first. Today: the judge no longer asks about lessons with nothing to judge (that path had been buying a model call on every local test run), the coverage pin was proven able to fail, and CI's three-week red has its mechanism named and fixed.

Design Director

on trackOpus

Learner dashboard, chips, gate subtitles, the shop and now the manager's roster measured on the brand in light and dark — the pixel, not the class name. The factory (staff) screens still need a staff sign-in to photograph.

Design Critic

idleOpus

Not yet in the loop. First job when wired: the MK11 regeneration's opening screen against the hand-built portal.

Source Auditor

on trackOpus

Tonight the source auditor got its first real instrument: assigned-but-never-cited, triaged and ranked, $0. On MK11 it names 65 things the documents said that the course never used — the harassment policy and the contact directory among them. A citation across a tenant boundary still fails one lesson, not the batch.

Survey Designer

on trackOpus

The client's own notes now reach the outline for every portal type, not only 'custom'. Intake can raise the lesson ceiling with a stated reason; nothing can lower the minutes floor.

Research Analyst

idleSonnet

Idle tonight by design ($0 run). Standing brief: research mode = ingest a researched document as a source, grounding preserved.

What this board cannot see

  • Chats on claude.ai, Cowork and Claude Chrome — no clock on this machine sees them. Hours from them are Alyssa's number.
  • The unspent balance on the API account — only the Anthropic console shows it, so 'not in the ledger' is a ceiling, not a bill.
  • Model calls that failed before returning a token count (147 rows) — billed or not, their cost is unknown to the ledger.
  • Calls made by test and eval runs whose rows were deleted at teardown, and one paid batch orphaned on Aug 20.
  • Costs outside any ledger — subscriptions, the domain, hosting — until Alyssa enters them.
  • The factory's own staff screens, which need a staff sign-in no test can perform: markup-proven, never photographed.
  • Whether a generated course is GOOD. Every gate here is deterministic; taste is the owner's, and the review-hours benchmark has not been run yet.
  • Anything marked 'self-reported' or 'estimate' — the label is there so the number is not mistaken for a measurement.

Recent changes

  • Alyssa released the boards' remaining design lists (1d, 1e, 1f) and ruled 'Fixed sky' for the games. Both design loops are running.
  • Daytime session closed: five of seven items done, one stopped for Alyssa (arcade 1c), $0 spent; brief docs/briefs/2026-09-12-day.md; the simple board is live at /simple.
  • Arcade fix list: five of the critic's points fixed and measured; the fourth review sent it back again, so it stops (two strikes) and the rest is 1c, yours to release.
  • Console re-read after Alyssa's top-up: credit is $50.78 (was $30.79), 30-day total $262.99, today $0.01. A second, plain-language board now lives at /simple — the full board stays.
  • Alyssa approved the outline rule at 30–40 lessons; the outline prompt is at version 12. Next paid step is hers: mark up the side-by-side, then regenerate.
  • Overnight run closed: 15 of 17 items, $0.002 spent, brief written through the CFO and plain-English roles.
  • Failed model calls now keep their price; dev probes refuse to run unledgered; three stale runs closed by id; Alyssa's 33 lessons laid beside the factory's 60 for mark-up.
  • Learners can search their own company documents from the portal and hear a lesson's key points read aloud. The AI-written answer waits on a prompt approval and an embedding pass — both Alyssa's call.
  • Ten measured palette presets on the branding console, researched by two design sub-agents — every one failed the dark-mode button test on the first pass and was corrected by measurement.
  • Managers can post announcements that appear on every learner's dashboard; pinned ones stay on top.
The other 47 changes
  • Learners and managers can message each other inside the portal — proven both ways through the real pages, with audit rows.
  • Managers see time this week and weak spots on the roster, and each learner's standing, time per lesson and arcade bests on their page — the same numbers the learner sees.
  • Leaderboard: arcade bests, best average check score, most study time this week — per-client switch, names as first name and last initial. The nav now marks the current page on every portal screen (it never had).
  • Learners can see where they stand: accuracy per module, graded exactly as their checks were, and time per lesson. The design critic sent the arcade back a second time; it waits for Alyssa.
  • Time in the portal is tracked per lesson and per day, proven by a learner reading a lesson for half a minute.
  • The dashboard draws the journey as a rail through numbered phase nodes, like Alyssa's own MK11 portal. The arcade was sent back by the design critic (its buttons rendered as bare text) and reworked.
  • Chill Study Mode: Mike's six games ported and fed by the client's own lessons and question bank; scores table under RLS; seen light and dark.
  • Alyssa's notes on the morning brief answered on the board; tests can no longer reach the paid API; the nightly eval disabled at the repository level (it had run again this morning); onboarding band set to 30–40 lessons; CFO, plain-English and factory-ops roles added; spend:audit.
  • Manager screens wear the dashboard's header; learner page opens with four tiles. Queue 3 closed.
  • CI unit tests green on GitHub for the first time since Aug 22. The leak job's own, older failure is diagnosed and waiting on Alyssa (protected test).
  • Content Verification Audit renders as a five-part client document; Resources page shows the client's contacts as cards.
  • CI had been red for three weeks. The cause: a test was quietly paying for a model call on Alyssa's API key to pass locally — and had no key on the runner. The judge now skips lessons with nothing to judge; the test can never reach a model.
  • The Content Verification Audit is now a document a client can sign: headline, do-these-first, five parts, sign-off — generated in one command, never carrying an internal id.
  • Coverage regression pin proven: broke the metric on purpose, the test failed; restored, it passed.
  • Resources page: the client's own lines become cards — who to call, help lines, the chain of command — and what the scan could not read is counted, not shown. On MK11: two real contacts, one help line, 23 unreadable lines named as such.
  • Resumed on schedule at 2:03 PM EDT. Eight-hour window; Queue 3; nothing will be spent.
  • Owner's review: the monthly spend and the Numbers tiles were stale, hand-typed values. Now every tile is measured at rebuild and names its source; months follow the Console with the ledger beside; the board stamps its own clock.
  • Paused at the owner's request; scheduled to resume at 2:03 PM EDT with Queue 3.
  • Loop closed: queue empty. No session is running; nothing is being spent.
  • Morning brief written. Night total: 22 queue items, ~45 commits, $0.003 spent on the factory key; three defects of my own found and fixed on the record.
  • Only 28% of batch-routed calls ever used the Batch API. The stage scripts now default to it; immediate delivery is an explicit opt-in. Regeneration will cost roughly half.
  • Where the money went: read from the console by API key — $249 of $259 on the factory key in 30 days; $69 before the ledger existed; a nightly cloud eval paid daily and failed daily (now off); the rest under-recorded. Credits left $35.
  • Content audit: enumerations (the source counts 8 Great Work Habits; extraction recovered 6) and P3-17's 'referenced but never explained' folded in — 25 such gaps on MK11.
  • Ledger defect found and fixed: ~2,000 calls delivered interactively were priced at the 50% batch rate. Board now shows the ledger as written plus the correction, and a section on what it cannot see.
  • Content audit: one command lists what a client's documents contained that the course never used — 65 real gaps on MK11 after triage (harassment policy, the contact directory, the four factors of impulse).
  • Spend: Alyssa's API top-ups entered ($298.20 across nine purchases, Aug 20 – Sep 7). The factory's ledger recorded $80.37 of calls; the gap is unlogged calls plus unspent balance.
  • Outline gate: headings that are really OCR word-salad or a scrambled running title no longer create phantom sections — 81 of 207 MK11 headings reclassified, measured at $0.
  • Grounding: tables and half-citations no longer count as unsupported claims — 159 → 3 artifact assertions on the pilot, measured without a model call.
  • This board: total project spend — model calls measured from the ledger by month and by kind of work; a slot for costs outside the ledger.
  • Reference snapshot: MK11's course exported outside the repo (2,562 rows) with a content-free manifest committed; restore drilled and verified on a fixture. A database reset now costs a command, not $5.
  • This board: the site's galaxy backdrop, and hours on the project measured from git.
  • Clients: a run stuck at 'running' for over a day is now called a stale run, with the fix named. Nine such rows found from earlier sessions.
  • Clients: one bucket per client on the factory — stage, headline, next step, deep links, each course's minutes per lesson.
  • Manager walk: a learner's completion shows on the manager's roster and learner page — walked through the real pages, light and dark.
  • shape:check — the course ruler runs on any built course in one command. All three MK11 courses ever generated refuse it (60×2.1 min, 17×6.8, 27×4.0).
  • The Build Board (this page) — rebuilt by every session, secret-guarded.
  • Pixel proof: chips, gate subtitles and shop CTA measured on brand in light and dark. Shop-CTA premise corrected.
  • M63 closed — roster suite owns its tenant; full suite 1,270 passed, 0 failed.
  • demo:share — one command hands a portal to a reviewer (prints the link, never a password).
  • Client notes reach the outline for every archetype; shape-gate corrected under the security suite.
  • Factory screens hide test and seed tenants by default.
  • The organisation as 8 durable agent roles (VP Review runs on Fable).
  • Practice calls: an 18-turn target with 4 turns of grace, judged before any model call.
  • A tenancy violation fails one lesson, not the whole author batch.
  • Phase timeline renders numbered chips with lock reasons, not a count.
  • Phase gates persist; MK11's six backfilled.
  • THE PLAN v2 is the single source of truth; the shape gate; M60; verified cloud backup in one command.