Portal Factory Build Board

Portal Factory Build Board

The path from tonight's code to a factory that is ready to sell — what is done, what is running, what is next, and what needs a person. For Alyssa and Dave.

updated ET

Reader settings

Text size
A Claude Code session is running Daytime build loop, Sept 11 — Queue 3, item 26 — the CI-only test failure: read the runner's log, name the mechanism
started
Sep 11, 2:03 PM
runs until
Sep 11, 10:03 PM
checks in every
13 min
spend tonight
$0 of $15 cap

Progress to sellable

54%of the path to sellable
6 done 2 in progress 2 waiting on a decision 1 blocked 2 not started
  • B1 reviewed & timed
  • B2 end-to-end walk
  • B3 live domain
  • B4 a new portal
  • B5 first payment

Done counts 1, in progress counts ½. Targets: a demo to sell from by September 24; a factory, not a demo, by October 1. Revenue floor the plan is built to clear: $345,000 a year.

Now · next · needs a person

In progress

  • Queue 3 — six $0 items, started 2:03 PM EDT as scheduled

Next in line

  • Alyssa: walk the fixture portal as learner then manager, and time it
  • Regenerate MK11 Onboarding once approved — through the pipeline's Batch API (about $5–6, not $10) — and compare with shape:check against the export
  • Reconcile 9 stale 'running' generation runs left by earlier sessions
  • Wire the design critic into the review loop

Needs a person

  • AlyssaApprove the outline prompt change (rule 18: lesson-count band, 8–12 minutes per lesson, phase weighting)Until then the ruler refuses but the prompt does not aim.
  • AlyssaSay yes to regenerating MK11 Onboarding (~$10)Unblocks Benchmark 1 — the review-hours number.
  • AlyssaRe-score the pilot's 30 lessons after the segmenter fix (a paid grounding run, a few dollars)Confirms the 159 → 3 artifact drop moves scores only where it should.
  • AlyssaPoint the nameservers at Cloudflare and sign in to FlyUnblocks the whole of Week 3.
  • Alyssa or DaveMove the environment file and the two credential folders into a password managerSecurity hygiene before the first client.

The 13 steps to sellable

  1. Step 1: Zero-cost draft loop

    Done Claude

    Done when A shape proposal is produced and gated with $0 API spend

    Shape gate lands on every outline before a dollar is spent.

  2. Step 2: Fix the ruler

    Done Claude

    Done when A 60×2-minute run refuses; each gate proven to fail first

    Per-archetype caps, 5-minute floor, front matter as furniture, phase balance. Three wrong versions of one rule caught and removed the same night.

  3. Step 3: Persist phase gates; chips instead of counts

    Done Claude

    Done when Gate subtitles visible in a browser

    Seen in a real browser, light and dark, 22:39 Sept 10.

  4. Step 4: M60 + spend-cap metering

    Done Claude

    Done when 4 tests green; coach runs on MK11

    Practice cap now meters only roleplay; 18-turn target with 4 turns of grace.

  5. Step 5: Regenerate MK11 Onboarding against the new ruler

    Waiting on a decision Claude

    Done when ~30 lessons, ≥5 min each, opens with an orientation

    Costs about $5–6 through the Batch API (the scripts had been paying full price). Needs Alyssa's go-ahead — it replaces the current 60 draft lessons, which are exported.

  6. Step 6: You review it — and we time it

    Waiting on a decision Alyssa

    Done when BENCHMARK 1 — the review-hours number exists

    Follows item 5.

  7. Step 7: Agent roles + morning brief

    Done Claude

    Done when 8 role files; a brief you actually read

    8 role files exist; the morning brief for Sept 11 was written and handed over (docs/briefs/2026-09-11.md).

  8. Step 8: Factory UI: one bucket per client

    Done Claude

    Done when You can run a portal without me

    Test tenants hidden by default; /clients shows every client with its stage and the next step, and each client has a bucket page with counts, deep links and each course's minutes per lesson.

  9. Step 9: Design pass: brand on every surface; critic loop

    In progress Claude

    Done when No page ships default-styled

    Learner surfaces measured on-brand pixel-for-pixel. Critic loop not wired.

  10. Step 10: Manager + learner end-to-end walk

    In progress Claude, then Alyssa

    Done when BENCHMARK 2 — a learner completes a lesson; a manager sees it

    Learner walk passed Sept 6. Manager side walked tonight: a completion shows on the roster and the learner page through the real doors. The benchmark itself is Alyssa walking it.

  11. Step 11: Deploy to the real domain

    Blocked Alyssa, then Claude

    Done when BENCHMARK 3 — both hostnames resolve

    Blocked on nameservers → Cloudflare and a Fly login. Both need Alyssa at a keyboard.

  12. Step 12: Generate a NEW portal end-to-end as proof

    Not started Claude

    Done when BENCHMARK 4 — a second client-ready portal from a folder

  13. Step 13: First paying client

    Not started Alyssa

    Done when BENCHMARK 5 — payment

Week by week

Week 1 — Build the ruler Days 1–5

  • Done: Zero-cost draft loop; per-archetype gates; minutes floor; furniture; phase balance
  • Done: Persist phase gates; backfill MK11's six; chips replace counts
  • Done: M60; spend-cap metering
  • Not started: Regenerate MK11 against the ruler (~$10, needs go-ahead)
  • Not started: Alyssa reviews it; we time it — BENCHMARK 1

Week 2 — Make it presentable Days 6–10

  • Done: Factory UI: one bucket per client; hide test tenants
  • Done: 8 agent role files; the morning brief
  • In progress: Design pass: brand tokens on every surface; Resources as cards; critic loop
  • In progress: Manager + learner end-to-end walk — BENCHMARK 2 (both halves reachable; Alyssa's walk pending)

Week 3 — Prove it in production Days 11–15

  • Not started: Nameservers, certificate, deploy (needs Alyssa)
  • Not started: Both hostnames resolve; Cloudflare Access on the Factory — BENCHMARK 3
  • Not started: A brand-new portal end-to-end from a folder — BENCHMARK 4
  • Not started: Buffer: defect tail, module edit form, object-storage backup

Numbers

Tests passing 1,336 0 failing · 2 skipped · last full run source: npm run test log, 2026-09-11T18:20:08.390Z
Security suite 95 / 95 tenant-isolation and leak checks source: npm run test:leak log, 2026-09-11T18:10:01.836Z
Commits, last 24 hours 64 424 on every branch all time · 0 unpushed source: git log
Model spend, last 24 hours $0.00 0 logged calls · loop cap $15.00 · the Console is the bill, this is what the ledger wrote down source: generation_runs, priced by delivery
Real courses that pass the ruler 1 of 6 refused: MK11 · Product Knowledge (Op · MK11 · Product Knowledge · MK11 Onboarding +2 source: shape gate over every non-test course in the database
Runs stuck at 'running' 9 older than a day · never finished · the owner decides whether to mark them abandoned source: pipeline_runs
What MK11 as built cost $19.02 60 lessons · 459 calls · priced by delivery · through the Batch API it would be about half source: generation_runs for that course, re-priced
Spend on courses never published 35% $28.41 of $80.37 in the ledger · the number the ruler exists to bring down source: generation_runs joined to courses.status
Practice call, per turn $0.01 buyer + observer per turn, 138 calls measured · 18-turn target ≈ $0.15 a call source: generation_runs, roleplay stages

Every number here is measured when the board is rebuilt and names its source. Nothing in this section is typed by hand.

Hours on the project

Claude Code on the factory131.2 hmeasured from 14 session transcripts on this Mac since 2026-07-12 · this week 11.6 h
Of that, time around commits86.7 hsince 2026-08-01 · 424 commits in 79 stretches · this week 5.8 h
Claude Code, every project205.9 hall sessions on this Mac, all products, since April
Alyssa, all in300 hself-reported Sept 11: over 300 hours across every chat, Cowork session and Claude Code run — the clocks below cannot see the chats on claude.ai, so this is the whole

Time around commits, by week

07/2711.3 h
08/0310.8 h
08/1721.5 h
08/249.5 h
08/3127.8 h
09/075.8 h

Two clocks, both measured on this Mac, both floors. Sessions: every Claude Code transcript carries timestamps; gaps over 30 minutes end a stretch, and a session counts for the factory when its content is about the factory. Commits: the same idea applied to git, a subset of the first. Neither can see chats on claude.ai, Cowork, reading, Google Docs, or the hand-built July portal — that is why the fourth tile is Alyssa's own number, and why it is the biggest.

Total project spend

Project total spent$263.12$298.20 paid in − $35.08 still on the account · tax included
Paid into the API account$298.209 purchases, August 20 – September 7 · $280.00 of credit + $18.20 tax
Left on the account$35.08console balance, read Sep 10, 11:50 PM ET · 13% of credit remains
Console, last 30 days$258.68portal-factory $249.28 · console (Playground) $9.40 · the factory's own ledger caught $105.45 of it · this week the console says $33.74, the ledger $1.26

By month (the console; the ledger beside it)

Aug$125.56 · ledger $32.80
Sep$123.73 · ledger $47.57

By kind of work

outlining$46.80
quizzes & exams$12.55
writing lessons$12.51
other$5.47
grounding checks$2.09
practice calls$0.95

API account top-ups (Alyssa)

  • August 20$31.95
  • August 20$21.30
  • August 21$53.25
  • August 22$21.30
  • August 25$53.25
  • September 4$21.30
  • September 4$21.30
  • September 5$21.30
  • September 7$53.25

Where the difference went (console minus ledger)

  • Aug 19–21: $68.81 on the factory key before the ledger existed — and the two database resets erased the rows that did exist.
  • A nightly cloud eval ran every day at 09:00 UTC with the real key, on a database that dies with the run, and FAILED every day since Sep 5 — about $1–4 a day, plus five push-triggered runs on Sep 6–7. Switched off Sept 11; re-enable only after a green on-demand run.
  • Sep 5–7 overnight work under-recorded: interactive calls priced at the batch rate (fixed), rows deleted at test teardown, 147 failed calls with no token count.
  • Almost no batch discount was ever earned: $2 off $260 — the pipeline ran interactively nearly every time. The stage scripts now default to the Batch API.

The console is the bill; the ledger is what the factory managed to write down. Money Alyssa put into the API account (credit purchases, tax included), entered by her on Sept 11. Ledger spend is measured every time the board is rebuilt (last measured Sep 11, 2:20 PM ET). Subscriptions, the domain and hosting are not in any ledger the factory can read — they are Alyssa's to enter, and the total above adds them in when she does.

Reports from each desk

VP Review

Fable, on track

Fourteen items closed and pushed tonight; two premises of my own corrected on the record — and, asked where the money went, I found it: a nightly cloud eval spending real credit every day and failing every day, invisible to the ledger. Off as of tonight. Tomorrow's ask is still one decision: the $10 regeneration — but now with the real per-course cost known.

Curriculum Architect

Opus, blocked or failing

The ruler now refuses 60×2-minute courses — and, run on real data tonight, refuses every MK11 course the factory has ever built. Aiming the prompt (rule 18) is the approval that turns refusal into a good outline.

Gate Engineer

Opus, on track

Every new gate failed first, then passed. Tonight: the ruler on real courses, the restore drill (two defects caught), the grounding segmenter (159→3 phantom claims), and the adjacency rule no longer defending sections under OCR word-salad (81 headings reclassified on MK11).

Design Director

Opus, on track

Learner dashboard, chips, gate subtitles, the shop and now the manager's roster measured on the brand in light and dark — the pixel, not the class name. The factory (staff) screens still need a staff sign-in to photograph.

Design Critic

Opus, idle

Not yet in the loop. First job when wired: the MK11 regeneration's opening screen against the hand-built portal.

Source Auditor

Opus, on track

Tonight the source auditor got its first real instrument: assigned-but-never-cited, triaged and ranked, $0. On MK11 it names 65 things the documents said that the course never used — the harassment policy and the contact directory among them. A citation across a tenant boundary still fails one lesson, not the batch.

Survey Designer

Opus, on track

The client's own notes now reach the outline for every portal type, not only 'custom'. Intake can raise the lesson ceiling with a stated reason; nothing can lower the minutes floor.

Research Analyst

Sonnet, idle

Idle tonight by design ($0 run). Standing brief: research mode = ingest a researched document as a source, grounding preserved.

What this board cannot see

  • Chats on claude.ai, Cowork and Claude Chrome — no clock on this machine sees them. Hours from them are Alyssa's number.
  • The unspent balance on the API account — only the Anthropic console shows it, so 'not in the ledger' is a ceiling, not a bill.
  • Model calls that failed before returning a token count (147 rows) — billed or not, their cost is unknown to the ledger.
  • Calls made by test and eval runs whose rows were deleted at teardown, and one paid batch orphaned on Aug 20.
  • Costs outside any ledger — subscriptions, the domain, hosting — until Alyssa enters them.
  • The factory's own staff screens, which need a staff sign-in no test can perform: markup-proven, never photographed.
  • Whether a generated course is GOOD. Every gate here is deterministic; taste is the owner's, and the review-hours benchmark has not been run yet.
  • Anything marked 'self-reported' or 'estimate' — the label is there so the number is not mistaken for a measurement.

Recent changes

  • The Content Verification Audit is now a document a client can sign: headline, do-these-first, five parts, sign-off — generated in one command, never carrying an internal id.
  • Coverage regression pin proven: broke the metric on purpose, the test failed; restored, it passed.
  • Resources page: the client's own lines become cards — who to call, help lines, the chain of command — and what the scan could not read is counted, not shown. On MK11: two real contacts, one help line, 23 unreadable lines named as such.
  • Resumed on schedule at 2:03 PM EDT. Eight-hour window; Queue 3; nothing will be spent.
  • Owner's review: the monthly spend and the Numbers tiles were stale, hand-typed values. Now every tile is measured at rebuild and names its source; months follow the Console with the ledger beside; the board stamps its own clock.
  • Paused at the owner's request; scheduled to resume at 2:03 PM EDT with Queue 3.
  • Loop closed: queue empty. No session is running; nothing is being spent.
  • Morning brief written. Night total: 22 queue items, ~45 commits, $0.003 spent on the factory key; three defects of my own found and fixed on the record.
  • Only 28% of batch-routed calls ever used the Batch API. The stage scripts now default to it; immediate delivery is an explicit opt-in. Regeneration will cost roughly half.
  • Where the money went: read from the console by API key — $249 of $259 on the factory key in 30 days; $69 before the ledger existed; a nightly cloud eval paid daily and failed daily (now off); the rest under-recorded. Credits left $35.
  • Content audit: enumerations (the source counts 8 Great Work Habits; extraction recovered 6) and P3-17's 'referenced but never explained' folded in — 25 such gaps on MK11.
  • Ledger defect found and fixed: ~2,000 calls delivered interactively were priced at the 50% batch rate. Board now shows the ledger as written plus the correction, and a section on what it cannot see.
  • Content audit: one command lists what a client's documents contained that the course never used — 65 real gaps on MK11 after triage (harassment policy, the contact directory, the four factors of impulse).
  • Spend: Alyssa's API top-ups entered ($298.20 across nine purchases, Aug 20 – Sep 7). The factory's ledger recorded $80.37 of calls; the gap is unlogged calls plus unspent balance.
  • Outline gate: headings that are really OCR word-salad or a scrambled running title no longer create phantom sections — 81 of 207 MK11 headings reclassified, measured at $0.
  • Grounding: tables and half-citations no longer count as unsupported claims — 159 → 3 artifact assertions on the pilot, measured without a model call.
  • This board: total project spend — model calls measured from the ledger by month and by kind of work; a slot for costs outside the ledger.
  • Reference snapshot: MK11's course exported outside the repo (2,562 rows) with a content-free manifest committed; restore drilled and verified on a fixture. A database reset now costs a command, not $5.
  • This board: the site's galaxy backdrop, and hours on the project measured from git.
  • Clients: a run stuck at 'running' for over a day is now called a stale run, with the fix named. Nine such rows found from earlier sessions.
  • Clients: one bucket per client on the factory — stage, headline, next step, deep links, each course's minutes per lesson.
  • Manager walk: a learner's completion shows on the manager's roster and learner page — walked through the real pages, light and dark.
  • shape:check — the course ruler runs on any built course in one command. All three MK11 courses ever generated refuse it (60×2.1 min, 17×6.8, 27×4.0).
  • The Build Board (this page) — rebuilt by every session, secret-guarded.
  • Pixel proof: chips, gate subtitles and shop CTA measured on brand in light and dark. Shop-CTA premise corrected.
  • M63 closed — roster suite owns its tenant; full suite 1,270 passed, 0 failed.
  • demo:share — one command hands a portal to a reviewer (prints the link, never a password).
  • Client notes reach the outline for every archetype; shape-gate corrected under the security suite.
  • Factory screens hide test and seed tenants by default.
  • The organisation as 8 durable agent roles (VP Review runs on Fable).
  • Practice calls: an 18-turn target with 4 turns of grace, judged before any model call.
  • A tenancy violation fails one lesson, not the whole author batch.
  • Phase timeline renders numbered chips with lock reasons, not a count.
  • Phase gates persist; MK11's six backfilled.
  • THE PLAN v2 is the single source of truth; the shape gate; M60; verified cloud backup in one command.