Frontier Inference Margins
An honesty-first interactive cost model of frontier-LLM inference serving margins — mechanism-first calculator, typed public-claims registry, and a long-form research report. Every number carries its epistemic status.
- Mechanism-first calculator — every assumption adjustable, each with its epistemic status attached
- 34-claim typed evidence registry with provenance tiers
- Read-only MCP connector at margins-mcp.ashitaorbis.com
- Three test suites and published release gates behind every deploy
Activity Timeline
Body-cap logic rewritten to pause/respond/destroy pattern; try/catch added to request handler; Content-Length precheck implemented. GPT Pro daily sweep completed, 2 commits applied findings.
Unpushed commit count corrected from 31 to 35. China dossier western hub electricity routing claim corrected.
Daily sweep updated with Pro research outputs from 2026-08-18. Release stamp faf6bec deployed with served-asset manifest.
Result tile moved into projections band; estimates repositioned to desktop top; dossier rendering deduplicated to a single route and justification stack collapsed. Dual-viewport empirical verification: 7,933 PASS / 0 FAIL. 12 commits across 3 sessions.
Cards-vintage ruling enforced: estimate-card faces migrated from default to round-3. Two authorized MCP transport deltas declared for hoisted estimates block. Parity test re-minted for consistency verification.
Reports restored from 449-byte stubs via fixture recovery. Daily sweep tracking DeepSeek V4 peak/off-peak pricing (57–1,100% lifts) and X.ai/Z.ai positioning.
DeepSeek V4 effective 2026-08-16: 2× peak/off-peak split, cache-hit markup increases up to +1,114%. Gemini 3.7 Flash launched 2026-08-13 at 50% introductory price cut. GPT Pro daily sweep dispatched; burn-guard reconciled harness packages and established 08-12 as canonical after sha256 check.
Headline-invariance tests 255/255 clean, reverse-compatibility validated, migration-differential CI honesty and metrics discontinuity suppression fixed. GPT Pro returned NOT-READY (16 findings); all folded into bq-290–294 and verified clean on second pass.
Slider semantics clarified: central value is MEAN not median for three-point mode. Centroid commit gates M8 chain. M4 restart sequenced after status report delivery.
Survivor-set exactness verified across both arms with declared ranges and reference vocabulary checked. GPT Pro arm landed with NVIDIA-additive correction; owner ruling applied. Max/min/median derived over provider mixes summing to 100.
1.25× frontier-lab assumption yields +2.4 months (was +2.9), post-Polaris adjudication. Public mirror synced to v2.2.0-2026-08-06. Contribution margin sweep: 77.31% engine median, 65–82% endpoint range.
Public mirror updated to track private source hash 5a07e2f (2026-08-06 release).
13 commits shipped, public mirror synced. RAISE Summit podcast claims drove metric decomposition across size revision, evidence adjudication, and serving-model re-engineering. Post-deploy audit found 4 critical gaps: stale 200 OK responses, unverified hardening commits, missing post-deploy probes, and pending workers.dev disable.
Vals AI DeepSeek V4 Flash at $0.20/task vs Kimi K3's $17.56/task — material cost differential captured in daily sweep. workers.dev endpoints closed.
Fetcher patched to recover the 07-27 collection gap. Previously missing data from 07-23/24/26 backfilled into the dataset. Automated daily sweep pipeline continues via async ChatGPT Pro dispatch.
Built recovery harness for three answers (28.8k–23.9k chars each). Poll loop with escalation logic now handles the timeout pattern. Cleared a sharp HIGH advisory in the worker dev tree.
Read-only token, checksum-verified shfmt, and .pyc cleanup applied to CI. Daily material backlog for prior week persisted.
Root cause: design inverted b*=154/chip vs actual ~16 concurrent-sequence ceiling. Standing recommendation: anchor future designs to 64-sequence batch ceiling. FlexNPU CM384 extracted for cycle-2. ChatGPT Pro polling now fully automated — dispatch, poll, persist.
Operating-point-closure error traced and verified (TPU7 Ironwood: 64 concurrent sequences, 16/chip, 518.86 tok/s/chip). CM384 FlexNPU extracted via parallel research agents. Epoch AI and wafer_ai added to hardware sweep; GPT Pro replica-width consult refuted both prior candidate bounds.
P0/P1 defect classes fixed and re-verified through Rounds 2–3. Script defect recovery advanced from r7 to r12 (r11 released defect-free). IM4 entry harness specified with survivorship regression fixtures.
Root cause of premature completion in disk_recovery path identified, documented, and patched in one commit. Daily automation pipeline persisting 2026-07-19 research findings (Z.ai, DeepSeek/Kimi, Anthropic credit line) to gptpro-reports/.
Ground-truth reconciliation pass underway against effective-MFU scalar model. ChatGPT Pro MCP completion-detection bug filed (commit 53b3dac). Deployment staged pending sign-off.
D2 receipt pack completed with full model/platform/operating-point verification. Cycle-2 re-engineering underway but threshold finalization blocked on owner gate.
Chart billing canonicalized to shared computeMix engine with regression tests. Four dive reports scrubbed of provenance headers, analysis preserved. GPT-Pro daily/weekly reports live via bash-owned fetcher with in-turn polling and 70-min codex fallback.
Page-set values relabeled so they no longer carry a false empirical scent; GPT-Pro outside-review findings remediated and the engine bumped to v2.1.4; quarantined feedback intake (Cloudflare Worker + D1 + Turnstile) embedded on the main page.
Preset dropdowns replaced by a mechanism-first margin-range evidence board with de-named routes and v4 permalinks. Read-only MCP server deployed as a Cloudflare Worker at margins-mcp.ashitaorbis.com behind a fixture-gated honesty envelope.
New traffic-mix axis with a pure codec and atomic permalink replays, specified first by failing-by-design contract fixtures. Nine reception audits run against the live site; all 16 P0 findings remediated and the resolution ledger published.
Honesty-first interactive cost model of frontier-LLM inference serving margins: mechanism-first calculator, typed public-claims registry, and a long-form research report, deployed to Cloudflare Pages.