Frontier Inference Margins
An honesty-first interactive cost model of frontier-LLM inference serving margins — mechanism-first calculator, typed public-claims registry, and a long-form research report. Every number carries its epistemic status.
- Mechanism-first calculator — a ≈77% central-scenario modeled unit serving margin (not a company gross margin), every assumption adjustable
- 34-claim typed evidence registry with provenance tiers
- Read-only MCP connector at margins-mcp.ashitaorbis.com
- Three test suites and published release gates behind every deploy
Activity Timeline
Read-only token, checksum-verified shfmt, and .pyc cleanup applied to CI. Daily material backlog for prior week persisted.
Root cause: design inverted b*=154/chip vs actual ~16 concurrent-sequence ceiling. Standing recommendation: anchor future designs to 64-sequence batch ceiling. FlexNPU CM384 extracted for cycle-2. ChatGPT Pro polling now fully automated — dispatch, poll, persist.
Operating-point-closure error traced and verified (TPU7 Ironwood: 64 concurrent sequences, 16/chip, 518.86 tok/s/chip). CM384 FlexNPU extracted via parallel research agents. Epoch AI and wafer_ai added to hardware sweep; GPT Pro replica-width consult refuted both prior candidate bounds.
P0/P1 defect classes fixed and re-verified through Rounds 2–3. Script defect recovery advanced from r7 to r12 (r11 released defect-free). IM4 entry harness specified with survivorship regression fixtures.
Root cause of premature completion in disk_recovery path identified, documented, and patched in one commit. Daily automation pipeline persisting 2026-07-19 research findings (Z.ai, DeepSeek/Kimi, Anthropic credit line) to gptpro-reports/.
Ground-truth reconciliation pass underway against effective-MFU scalar model. ChatGPT Pro MCP completion-detection bug filed (commit 53b3dac). Deployment staged pending sign-off.
D2 receipt pack completed with full model/platform/operating-point verification. Cycle-2 re-engineering underway but threshold finalization blocked on owner gate.
Chart billing canonicalized to shared computeMix engine with regression tests. Four dive reports scrubbed of provenance headers, analysis preserved. GPT-Pro daily/weekly reports live via bash-owned fetcher with in-turn polling and 70-min codex fallback.
Page-set values relabeled so they no longer carry a false empirical scent; GPT-Pro outside-review findings remediated and the engine bumped to v2.1.4; quarantined feedback intake (Cloudflare Worker + D1 + Turnstile) embedded on the main page.
Preset dropdowns replaced by a mechanism-first margin-range evidence board with de-named routes and v4 permalinks. Read-only MCP server deployed as a Cloudflare Worker at margins-mcp.ashitaorbis.com behind a fixture-gated honesty envelope.
New traffic-mix axis with a pure codec and atomic permalink replays, specified first by failing-by-design contract fixtures. Nine reception audits run against the live site; all 16 P0 findings remediated and the resolution ledger published.
Honesty-first interactive cost model of frontier-LLM inference serving margins: mechanism-first calculator, typed public-claims registry, and a long-form research report, deployed to Cloudflare Pages.