Portal Factory Build Board

Portal Factory Build Board

The path from tonight's code to a factory that is ready to sell — what is done, what is running, what is next, and what needs a person. For Alyssa and Dave.

updated ET

Reader settings

Text size
No session running Daytime build loop, Sept 11 — closed — Nothing. Queue 3 closed 2:50 PM EDT; CI unit tests green; one protected-test fix waiting on Alyssa.
started
Sep 11, 2:03 PM
runs until
Dec 31, 7:00 PM
checks in every
13 min
spend tonight
$0 of $15 cap

Progress to sellable

54%of the path to sellable
6 done 2 in progress 2 waiting on a decision 1 blocked 2 not started
  • B1 reviewed & timed
  • B2 end-to-end walk
  • B3 live domain
  • B4 a new portal
  • B5 first payment

Done counts 1, in progress counts ½. Targets: a demo to sell from by September 24; a factory, not a demo, by October 1. Revenue floor the plan is built to clear: $345,000 a year.

Now · next · needs a person

In progress

    Next in line

    • Alyssa: walk the fixture portal as learner then manager, and time it
    • Regenerate MK11 Onboarding once approved — through the pipeline's Batch API (about $5–6, not $10) — and compare with shape:check against the export
    • Reconcile 9 stale 'running' generation runs left by earlier sessions
    • Wire the design critic into the review loop (only Week-2 design item left)

    Needs a person

    • AlyssaApprove the outline prompt change (rule 18: lesson-count band, 8–12 minutes per lesson, phase weighting)Until then the ruler refuses but the prompt does not aim.
    • AlyssaSay yes to regenerating MK11 Onboarding (~$10)Unblocks Benchmark 1 — the review-hours number.
    • AlyssaRe-score the pilot's 30 lessons after the segmenter fix (a paid grounding run, a few dollars)Confirms the 159 → 3 artifact drop moves scores only where it should.
    • AlyssaPoint the nameservers at Cloudflare and sign in to FlyUnblocks the whole of Week 3.
    • Alyssa or DaveMove the environment file and the two credential folders into a password managerSecurity hygiene before the first client.
    • AlyssaOne-line fix in a protected security test so the GitHub leak job goes green: its CONTROL case picks a learner/course pair the seed has already enrolled. Mechanism in BUILD_LOG 15:20 (Sept 11).

    The 13 steps to sellable

    1. Step 1: Zero-cost draft loop

      Done Claude

      Done when A shape proposal is produced and gated with $0 API spend

      Shape gate lands on every outline before a dollar is spent.

    2. Step 2: Fix the ruler

      Done Claude

      Done when A 60×2-minute run refuses; each gate proven to fail first

      Per-archetype caps, 5-minute floor, front matter as furniture, phase balance. Three wrong versions of one rule caught and removed the same night.

    3. Step 3: Persist phase gates; chips instead of counts

      Done Claude

      Done when Gate subtitles visible in a browser

      Seen in a real browser, light and dark, 22:39 Sept 10.

    4. Step 4: M60 + spend-cap metering

      Done Claude

      Done when 4 tests green; coach runs on MK11

      Practice cap now meters only roleplay; 18-turn target with 4 turns of grace.

    5. Step 5: Regenerate MK11 Onboarding against the new ruler

      Waiting on a decision Claude

      Done when ~30 lessons, ≥5 min each, opens with an orientation

      Costs about $5–6 through the Batch API (the scripts had been paying full price). Needs Alyssa's go-ahead — it replaces the current 60 draft lessons, which are exported.

    6. Step 6: You review it — and we time it

      Waiting on a decision Alyssa

      Done when BENCHMARK 1 — the review-hours number exists

      Follows item 5.

    7. Step 7: Agent roles + morning brief

      Done Claude

      Done when 8 role files; a brief you actually read

      8 role files exist; the morning brief for Sept 11 was written and handed over (docs/briefs/2026-09-11.md).

    8. Step 8: Factory UI: one bucket per client

      Done Claude

      Done when You can run a portal without me

      Test tenants hidden by default; /clients shows every client with its stage and the next step, and each client has a bucket page with counts, deep links and each course's minutes per lesson.

    9. Step 9: Design pass: brand on every surface; critic loop

      In progress Claude

      Done when No page ships default-styled

      Learner surfaces measured on-brand pixel-for-pixel. Manager roster and learner page now share the dashboard's header treatment (Sept 11 afternoon, light and dark on the fixture). Critic loop not wired.

    10. Step 10: Manager + learner end-to-end walk

      In progress Claude, then Alyssa

      Done when BENCHMARK 2 — a learner completes a lesson; a manager sees it

      Learner walk passed Sept 6. Manager side walked tonight: a completion shows on the roster and the learner page through the real doors. The benchmark itself is Alyssa walking it.

    11. Step 11: Deploy to the real domain

      Blocked Alyssa, then Claude

      Done when BENCHMARK 3 — both hostnames resolve

      Blocked on nameservers → Cloudflare and a Fly login. Both need Alyssa at a keyboard.

    12. Step 12: Generate a NEW portal end-to-end as proof

      Not started Claude

      Done when BENCHMARK 4 — a second client-ready portal from a folder

    13. Step 13: First paying client

      Not started Alyssa

      Done when BENCHMARK 5 — payment

    Week by week

    Week 1 — Build the ruler Days 1–5

    • Done: Zero-cost draft loop; per-archetype gates; minutes floor; furniture; phase balance
    • Done: Persist phase gates; backfill MK11's six; chips replace counts
    • Done: M60; spend-cap metering
    • Not started: Regenerate MK11 against the ruler (~$10, needs go-ahead)
    • Not started: Alyssa reviews it; we time it — BENCHMARK 1

    Week 2 — Make it presentable Days 6–10

    • Done: Factory UI: one bucket per client; hide test tenants
    • Done: 8 agent role files; the morning brief
    • In progress: Design pass: brand tokens on every surface; Resources as cards; critic loop
    • In progress: Manager + learner end-to-end walk — BENCHMARK 2 (both halves reachable; Alyssa's walk pending)

    Week 3 — Prove it in production Days 11–15

    • Not started: Nameservers, certificate, deploy (needs Alyssa)
    • Not started: Both hostnames resolve; Cloudflare Access on the Factory — BENCHMARK 3
    • Not started: A brand-new portal end-to-end from a folder — BENCHMARK 4
    • Not started: Buffer: defect tail, module edit form, object-storage backup

    Numbers

    Tests passing 1,340 0 failing · 2 skipped · last full run source: npm run test log, 2026-09-11T18:45:22.325Z
    Security suite 95 / 95 tenant-isolation and leak checks source: npm run test:leak log, 2026-09-11T18:45:32.023Z
    Commits, last 24 hours 70 430 on every branch all time · 0 unpushed source: git log
    Model spend, last 24 hours $0.00 0 logged calls · loop cap $15.00 · the Console is the bill, this is what the ledger wrote down source: generation_runs, priced by delivery
    Real courses that pass the ruler 1 of 6 refused: MK11 · Product Knowledge (Op · MK11 · Product Knowledge · MK11 Onboarding +2 source: shape gate over every non-test course in the database
    Runs stuck at 'running' 9 older than a day · never finished · the owner decides whether to mark them abandoned source: pipeline_runs
    What MK11 as built cost $19.02 60 lessons · 459 calls · priced by delivery · through the Batch API it would be about half source: generation_runs for that course, re-priced
    Spend on courses never published 35% $28.41 of $80.37 in the ledger · the number the ruler exists to bring down source: generation_runs joined to courses.status
    Practice call, per turn $0.01 buyer + observer per turn, 138 calls measured · 18-turn target ≈ $0.15 a call source: generation_runs, roleplay stages

    Every number here is measured when the board is rebuilt and names its source. Nothing in this section is typed by hand.

    Hours on the project

    Claude Code on the factory131.7 hmeasured from 14 session transcripts on this Mac since 2026-07-12 · this week 12 h
    Of that, time around commits87.1 hsince 2026-08-01 · 430 commits in 79 stretches · this week 6.2 h
    Claude Code, every project206.3 hall sessions on this Mac, all products, since April
    Alyssa, all in300 hself-reported Sept 11: over 300 hours across every chat, Cowork session and Claude Code run — the clocks below cannot see the chats on claude.ai, so this is the whole

    Time around commits, by week

    07/2711.3 h
    08/0310.8 h
    08/1721.5 h
    08/249.5 h
    08/3127.8 h
    09/076.2 h

    Two clocks, both measured on this Mac, both floors. Sessions: every Claude Code transcript carries timestamps; gaps over 30 minutes end a stretch, and a session counts for the factory when its content is about the factory. Commits: the same idea applied to git, a subset of the first. Neither can see chats on claude.ai, Cowork, reading, Google Docs, or the hand-built July portal — that is why the fourth tile is Alyssa's own number, and why it is the biggest.

    Total project spend

    Project total spent$263.12$298.20 paid in − $35.08 still on the account · tax included
    Paid into the API account$298.209 purchases, August 20 – September 7 · $280.00 of credit + $18.20 tax
    Left on the account$35.08console balance, read Sep 10, 11:50 PM ET · 13% of credit remains
    Console, last 30 days$258.68portal-factory $249.28 · console (Playground) $9.40 · the factory's own ledger caught $105.45 of it · this week the console says $33.74, the ledger $1.26

    By month (the console; the ledger beside it)

    Aug$125.56 · ledger $32.80
    Sep$123.73 · ledger $47.57

    By kind of work

    outlining$46.80
    quizzes & exams$12.55
    writing lessons$12.51
    other$5.47
    grounding checks$2.09
    practice calls$0.95

    API account top-ups (Alyssa)

    • August 20$31.95
    • August 20$21.30
    • August 21$53.25
    • August 22$21.30
    • August 25$53.25
    • September 4$21.30
    • September 4$21.30
    • September 5$21.30
    • September 7$53.25

    Where the difference went (console minus ledger)

    • Aug 19–21: $68.81 on the factory key before the ledger existed — and the two database resets erased the rows that did exist.
    • A nightly cloud eval ran every day at 09:00 UTC with the real key, on a database that dies with the run, and FAILED every day since Sep 5 — about $1–4 a day, plus five push-triggered runs on Sep 6–7. Switched off Sept 11; re-enable only after a green on-demand run.
    • Sep 5–7 overnight work under-recorded: interactive calls priced at the batch rate (fixed), rows deleted at test teardown, 147 failed calls with no token count.
    • Almost no batch discount was ever earned: $2 off $260 — the pipeline ran interactively nearly every time. The stage scripts now default to the Batch API.

    The console is the bill; the ledger is what the factory managed to write down. Money Alyssa put into the API account (credit purchases, tax included), entered by her on Sept 11. Ledger spend is measured every time the board is rebuilt (last measured Sep 11, 2:47 PM ET). Subscriptions, the domain and hosting are not in any ledger the factory can read — they are Alyssa's to enter, and the total above adds them in when she does.

    Reports from each desk

    VP Review

    Fable, on track

    Fourteen items closed and pushed tonight; two premises of my own corrected on the record — and, asked where the money went, I found it: a nightly cloud eval spending real credit every day and failing every day, invisible to the ledger. Off as of tonight. Tomorrow's ask is still one decision: the $10 regeneration — but now with the real per-course cost known.

    Curriculum Architect

    Opus, blocked or failing

    The ruler now refuses 60×2-minute courses — and, run on real data tonight, refuses every MK11 course the factory has ever built. Aiming the prompt (rule 18) is the approval that turns refusal into a good outline.

    Gate Engineer

    Opus, on track

    Every new gate failed first. Today: the judge no longer asks about lessons with nothing to judge (that path had been buying a model call on every local test run), the coverage pin was proven able to fail, and CI's three-week red has its mechanism named and fixed.

    Design Director

    Opus, on track

    Learner dashboard, chips, gate subtitles, the shop and now the manager's roster measured on the brand in light and dark — the pixel, not the class name. The factory (staff) screens still need a staff sign-in to photograph.

    Design Critic

    Opus, idle

    Not yet in the loop. First job when wired: the MK11 regeneration's opening screen against the hand-built portal.

    Source Auditor

    Opus, on track

    Tonight the source auditor got its first real instrument: assigned-but-never-cited, triaged and ranked, $0. On MK11 it names 65 things the documents said that the course never used — the harassment policy and the contact directory among them. A citation across a tenant boundary still fails one lesson, not the batch.

    Survey Designer

    Opus, on track

    The client's own notes now reach the outline for every portal type, not only 'custom'. Intake can raise the lesson ceiling with a stated reason; nothing can lower the minutes floor.

    Research Analyst

    Sonnet, idle

    Idle tonight by design ($0 run). Standing brief: research mode = ingest a researched document as a source, grounding preserved.

    What this board cannot see

    • Chats on claude.ai, Cowork and Claude Chrome — no clock on this machine sees them. Hours from them are Alyssa's number.
    • The unspent balance on the API account — only the Anthropic console shows it, so 'not in the ledger' is a ceiling, not a bill.
    • Model calls that failed before returning a token count (147 rows) — billed or not, their cost is unknown to the ledger.
    • Calls made by test and eval runs whose rows were deleted at teardown, and one paid batch orphaned on Aug 20.
    • Costs outside any ledger — subscriptions, the domain, hosting — until Alyssa enters them.
    • The factory's own staff screens, which need a staff sign-in no test can perform: markup-proven, never photographed.
    • Whether a generated course is GOOD. Every gate here is deterministic; taste is the owner's, and the review-hours benchmark has not been run yet.
    • Anything marked 'self-reported' or 'estimate' — the label is there so the number is not mistaken for a measurement.

    Recent changes

    • Manager screens wear the dashboard's header; learner page opens with four tiles. Queue 3 closed.
    • CI unit tests green on GitHub for the first time since Aug 22. The leak job's own, older failure is diagnosed and waiting on Alyssa (protected test).
    • Content Verification Audit renders as a five-part client document; Resources page shows the client's contacts as cards.
    • CI had been red for three weeks. The cause: a test was quietly paying for a model call on Alyssa's API key to pass locally — and had no key on the runner. The judge now skips lessons with nothing to judge; the test can never reach a model.
    • The Content Verification Audit is now a document a client can sign: headline, do-these-first, five parts, sign-off — generated in one command, never carrying an internal id.
    • Coverage regression pin proven: broke the metric on purpose, the test failed; restored, it passed.
    • Resources page: the client's own lines become cards — who to call, help lines, the chain of command — and what the scan could not read is counted, not shown. On MK11: two real contacts, one help line, 23 unreadable lines named as such.
    • Resumed on schedule at 2:03 PM EDT. Eight-hour window; Queue 3; nothing will be spent.
    • Owner's review: the monthly spend and the Numbers tiles were stale, hand-typed values. Now every tile is measured at rebuild and names its source; months follow the Console with the ledger beside; the board stamps its own clock.
    • Paused at the owner's request; scheduled to resume at 2:03 PM EDT with Queue 3.
    • Loop closed: queue empty. No session is running; nothing is being spent.
    • Morning brief written. Night total: 22 queue items, ~45 commits, $0.003 spent on the factory key; three defects of my own found and fixed on the record.
    • Only 28% of batch-routed calls ever used the Batch API. The stage scripts now default to it; immediate delivery is an explicit opt-in. Regeneration will cost roughly half.
    • Where the money went: read from the console by API key — $249 of $259 on the factory key in 30 days; $69 before the ledger existed; a nightly cloud eval paid daily and failed daily (now off); the rest under-recorded. Credits left $35.
    • Content audit: enumerations (the source counts 8 Great Work Habits; extraction recovered 6) and P3-17's 'referenced but never explained' folded in — 25 such gaps on MK11.
    • Ledger defect found and fixed: ~2,000 calls delivered interactively were priced at the 50% batch rate. Board now shows the ledger as written plus the correction, and a section on what it cannot see.
    • Content audit: one command lists what a client's documents contained that the course never used — 65 real gaps on MK11 after triage (harassment policy, the contact directory, the four factors of impulse).
    • Spend: Alyssa's API top-ups entered ($298.20 across nine purchases, Aug 20 – Sep 7). The factory's ledger recorded $80.37 of calls; the gap is unlogged calls plus unspent balance.
    • Outline gate: headings that are really OCR word-salad or a scrambled running title no longer create phantom sections — 81 of 207 MK11 headings reclassified, measured at $0.
    • Grounding: tables and half-citations no longer count as unsupported claims — 159 → 3 artifact assertions on the pilot, measured without a model call.
    • This board: total project spend — model calls measured from the ledger by month and by kind of work; a slot for costs outside the ledger.
    • Reference snapshot: MK11's course exported outside the repo (2,562 rows) with a content-free manifest committed; restore drilled and verified on a fixture. A database reset now costs a command, not $5.
    • This board: the site's galaxy backdrop, and hours on the project measured from git.
    • Clients: a run stuck at 'running' for over a day is now called a stale run, with the fix named. Nine such rows found from earlier sessions.
    • Clients: one bucket per client on the factory — stage, headline, next step, deep links, each course's minutes per lesson.
    • Manager walk: a learner's completion shows on the manager's roster and learner page — walked through the real pages, light and dark.
    • shape:check — the course ruler runs on any built course in one command. All three MK11 courses ever generated refuse it (60×2.1 min, 17×6.8, 27×4.0).
    • The Build Board (this page) — rebuilt by every session, secret-guarded.
    • Pixel proof: chips, gate subtitles and shop CTA measured on brand in light and dark. Shop-CTA premise corrected.
    • M63 closed — roster suite owns its tenant; full suite 1,270 passed, 0 failed.
    • demo:share — one command hands a portal to a reviewer (prints the link, never a password).
    • Client notes reach the outline for every archetype; shape-gate corrected under the security suite.
    • Factory screens hide test and seed tenants by default.
    • The organisation as 8 durable agent roles (VP Review runs on Fable).
    • Practice calls: an 18-turn target with 4 turns of grace, judged before any model call.
    • A tenancy violation fails one lesson, not the whole author batch.
    • Phase timeline renders numbered chips with lock reasons, not a count.
    • Phase gates persist; MK11's six backfilled.
    • THE PLAN v2 is the single source of truth; the shape gate; M60; verified cloud backup in one command.