Ashita Orbis

The Knockoff Eval's Null Baseline: Full Numbers

Ashita Orbis|July 25, 2026|3 min read

The parent post closed by naming its own most consequential missing comparison: ensembles against a single GPT-5.5 xhigh call. This page holds the numbers from running it. The design was locked before any result was seen: the identical eight frozen queries, one gpt-5.5 call per query pinned to xhigh reasoning under the same harness the May council members used (workspace access, final-message capture, no personas, no synthesis), judged blind against the four ensemble arms with the original judge prompt, two judges, both positions, and a fresh judgment table with one row per judgment so the May duplicate-write defect could not recur. Model and effort are proven per call from the CLI headers: eight of eight report gpt-5.5 at xhigh.

Results

| pair | wins of 32 | high confidence | GPT-5.5 judge | Opus judge | |---|---|---|---|---| | single call vs Council + Opus synth (1B) | 1B 28 – 4 | 1B 17 – 0 | 1B 15–1 | 1B 13–3 | | single call vs GPT Max + Opus synth (2B) | 2B 25 – 7 | 2B 13 – 1 | 2B 15–1 | 2B 10–6 | | single call vs Council + GPT-5.5 synth (1A) | 1A 19 – 13 | 1A 11 – 1 | 1A 12–4 | single 9–7 | | single call vs GPT Max + GPT-5.5 synth (2A) | 2A 20 – 12 | 2A 7 – 5 | 2A 12–4 | 8–8 split |

Every ensemble arm beats the single call on pooled votes, and the margin tracks the synthesizer. Against the Opus-synthesized arms the single call loses decisively, with a combined high-confidence count of 30 to 1. Against the GPT-5.5-synthesized arms it loses narrowly on pooled votes while the Opus judge scores those two contests as parity or better for the single call. Reported per judge and never merged, following the parent post's own rule; note that the GPT-5.5 judge showed no favor toward its own family's single call, voting against it by wide margins in all four contests.

The judge-vintage anchor

The May judgments cannot be regenerated: the GPT-5.5 judge then ran on a config default that has since moved, and the Opus alias has advanced a version. To license comparing new rows against the May table, the most contested May pair, Council-Opus against GPT-Max-Opus (1B vs 2B), was re-judged with today's pinned judges on all eight queries. May's deduplicated result was a 16–16 tie; the re-judge returned 17–14 with one tie. A shift of a vote and a half on an even 32-vote split reads as small drift, so the null-pair rows are treated as comparable with the May table, with the vintage difference disclosed rather than assumed away.

Reading

The parent post's fear was that a single deeper call might deliver most of the ensemble's value at a quarter of the cost, making the architecture over-engineered. On these eight queries it does not: the flagship configuration (either ensemble with Opus synthesis) beats the single call in 28 of 32 judgments against one panel and 25 of 32 against the other, with the high-confidence judgments nearly unanimous. What the null baseline adds beyond relief is a second appearance of the eval's original headline, in a comparison the original never ran: synthesizer choice dominates. The single call is only narrowly edged by the ensembles that synthesize with GPT-5.5 and is beaten decisively by the same ensembles when only the synthesizer changes to Opus, which is the cleanest evidence yet that the ensemble's value on these queries concentrates in the synthesis seat rather than in the breadth of the panel.

Caveats: one query distribution, thirty-two judgments per pair, no significance testing performed or implied; the synthesizer without any panel (an Opus deep call synthesizing nothing) was never tested, so the synthesis-seat reading is an inference from the synthesizer swap, not a measured single-seat result; judges are one per family and judge-family self-preference remains a live concern the parent post already documents; the single call ran with the same workspace access as the May council members, so this is a fair test of the machinery, not of a bare API call.