Ashita Orbis//understanding ai7 protocols
interactive

Course lesson · Day 20 of 306 figures

Why Benchmark Scores Rot

On 4 September 2026 Artificial Analysis, an independent firm that ranks AI models, dropped a graduate-level science test called GPQA Diamond from its headline index, describing it as an evaluation "that has now been saturated". Six days later Epoch AI, the research group behind the FrontierMath benchmark, announced that every problem in that test's hardest tier had now been solved by AI. A study published in the proceedings of this year's International Conference on Machine Learning (ICML) found, by its own measure, that 29 of 60 widely used tests of language models had largely lost the ability to separate the leading systems. A benchmark score is a claim that a small, fixed set of questions stands in for a skill, and that claim decays in three common ways—contamination, overfitting and saturation—each with its own history, its own signature and its own remedy.

About 26 min read11 min listenPrint edition (PDF)

Published 2026-10-02Sources read through 2026-10-01

Listen · 11 min

The spoken edition of this episode — its own script, read by a synthetic voice, with quoted passages in a second voice. It is not this page read aloud: the written edition you are reading was written separately, and neither needs the other. Download audio

Episode notes

The argument of the spoken edition, how it runs, what to take from it and the sources it read — in the episode’s own words. The full transcript is below; the written edition, with its figures, follows.

The argument

On the fourth of September, Artificial Analysis, an independent company that publishes rankings of AI models, changed what goes into its headline score. It added two tests and took one out. Its changelog gave the reason in a single line.

How it runs

  1. Why it's hard to follow — Two readings of news like this mislead. The first is that saturated means mastered: the machines have learned the subject, so the test has nothing left to ask. The study's authors mean something narrower.
  2. The idea you need — A benchmark score is a claim: that a small, fixed set of questions stands in for a skill. There are three common ways that claim decays, and each breaks a different link. The first is contamination: the test leaks into what the model learned from.
  3. What actually happened — Here's one test's life. G P Q A was published in November twenty twenty-three by researchers led from New York University: graduate-level science questions meant to be Google-proof.
  4. The contrast — So what do you do about a rotting test? Two answers are on offer, and they're aimed at different rots. The first is to keep the questions out of reach.
  5. What to watch — One. Artificial Analysis's index. The company says version five will raise the private share further and that more interim updates are coming.

What to take from it

The idea to keep is that a benchmark score has a shelf life, and three questions tell you how much of it is left. Could the model have seen these questions, or close copies, before the test? Has a whole field been tuning against this one test for years? And can the test still tell the top systems apart, by more than its own margin of error? None of those makes a score worthless. Each tells you how much of its claim is still standing.

To read more: When AI Benchmarks Plateau, by Mubashara Akhtar, Anka Reuel and colleagues, from this year's International Conference on Machine Learning. Its appendix walks through five benchmarks, from saturated to not, in a page.

Sources read for this episode (23)

  1. Artificial Analysis, *Announcing Artificial Analysis Intelligence Index v4.2*, artificialanalysis.ai/articles — 4 September 2026
  2. Artificial Analysis, post on X on Intelligence Index v4.2 — 6 September 2026 (00:28 UTC)
  3. Artificial Analysis, *Intelligence Benchmarking Methodology*, version 4.3.2 — read 1 October 2026
  4. Epoch AI, post on X on FrontierMath Tier 4 — 10 September 2026
  5. Epoch AI, *FrontierMath Tier 4* benchmark page and *Benchmarks* hub — read 1 October 2026
  6. Akhtar, Reuel, Soni, Ahuja et al., *When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation*, arXiv 2602.16763 v4 (ICML 2026) — 18 February 2026; v4 6 August 2026
  7. Rein et al., *GPQA: A Graduate-Level Google-Proof Q&A Benchmark*, arXiv 2311.12022 — 20 November 2023
  8. Brown et al. (OpenAI), *Language Models are Few-Shot Learners*, arXiv 2005.14165 — 28 May 2020
  9. OpenAI, *GPT-4 Technical Report*, arXiv 2303.08774 — March 2023
  10. BIG-bench, *training_on_test_set* README (canary GUID), in the BIG-bench repository on GitHub — read 1 October 2026
  11. Jain et al., *LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code*, arXiv 2403.07974 — 12 March 2024; v2 6 June 2024
  12. Zhang et al. (Scale AI), *A Careful Examination of Large Language Model Performance on Grade School Arithmetic* (GSM1k), arXiv 2405.00332 — 1 May 2024; revised November 2024
  13. International AI Safety Report 2026, arXiv 2602.21012 — February 2026; arXiv 24 February 2026
  14. Recht, Roelofs, Schmidt and Shankar, *Do ImageNet Classifiers Generalize to ImageNet?*, arXiv 1902.10811 (ICML 2019) — 13 February 2019
  15. Roelofs et al., *A Meta-Analysis of Overfitting in Machine Learning*, NeurIPS 2019 — December 2019
  16. Kiela et al., *Dynabench: Rethinking Benchmarking in NLP*, arXiv 2104.14337 — 7 April 2021
  17. Wang et al., arXiv 1804.07461, *GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding* — 20 April 2018
  18. Wang et al., *SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems*, arXiv 1905.00537 — 2 May 2019
  19. Microsoft Research, *Microsoft DeBERTa surpasses human performance on the SuperGLUE benchmark* — 6 January 2021
  20. Gema et al., *Are We Done with MMLU?*, arXiv 2406.04127 (v3) — 6 June 2024; v3 10 January 2025
  21. Phan et al., *Humanity's Last Exam*, arXiv 2501.14249 (quoted from v11) — 24 January 2025; v11 28 July 2026
  22. White et al., *LiveBench: A Challenging, Contamination-Limited LLM Benchmark*, arXiv 2406.19314; LiveBench changelog — 27 June 2024; changelog read 1 October 2026
  23. Microsoft, *SWE-bench-Live* README, in the SWE-bench-Live repository on GitHub — read 1 October 2026
Full transcript — 1,620 words, about 8 min read

The spoken edition, word for word. A quoted passage is its source read aloud — the source’s own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it. Download the transcript (Markdown). Printing this page prints the transcript; the article has its own print edition.

On the fourth of September, Artificial Analysis, an independent company that publishes rankings of AI models, changed what goes into its headline score. It added two tests and took one out. Its changelog gave the reason in a single line.

G P Q A Diamond, an exceptional scientific reasoning evaluation that has now been saturated

— Artificial Analysis, 'Announcing Artificial Analysis Intelligence Index v4.2', 4 September 2026, changelog list under the heading 'Intelligence Index v4.2 changelog'; https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2

As printed in the source: “GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated”

On X, the company added a second reason: the test's multiple-choice format doesn't reflect real tasks in those fields. And a study published at this year's International Conference on Machine Learning looked at sixty widely used tests of language models and found, by its own measure, that twenty-nine had largely lost the power to tell the leading models apart. Today: why a test that once measured something stops measuring what people think it does.

Two readings of news like this mislead.

The first is that saturated means mastered: the machines have learned the subject, so the test has nothing left to ask. The study's authors mean something narrower. The best five models can't be reliably told apart, and they're bunched near the highest score anyone has reached. That needn't be a hundred per cent. One test, LiveBench, came out as very highly saturated with its leaders clustered around seventy-nine per cent, which the authors read as models stalling, not the task being finished. And they add a condition: if a test was valid in the first place, saturation can mean the task is solved. Whether it measured what its name says is a separate question, and a crowded top score doesn't answer it.

The second reading is the cynical one: high scores just mean the models memorised the answers, so benchmark progress is fake. Leakage is real, and we'll get to it. But a test can also stop being useful because the models genuinely got better. Contamination and overfitting are possibilities to investigate, not conclusions you can read off a high score. The useful question isn't whether a score has rotted. It's how, and how much.

A benchmark score is a claim: that a small, fixed set of questions stands in for a skill. There are three common ways that claim decays, and each breaks a different link.

The first is contamination: the test leaks into what the model learned from. A student who has seen the exam paper can score well without knowing the subject. Language models learn from enormous scrapes of the internet, and benchmarks are published on the internet. In May twenty twenty, OpenAI's paper introducing GPT-3 described trying to remove benchmark questions from its training data. Then this.

Unfortunately, a bug in the filtering caused us to ignore some overlaps, and due to the cost of training it was not feasible to retrain the model.

— Tom B. Brown and colleagues (OpenAI), 'Language Models are Few-Shot Learners', arXiv 2005.14165, first posted 28 May 2020 (version 4, 22 July 2020), section 2.2 'Training Dataset', page 9

So they measured instead, rescoring each test on just the questions they were confident the model hadn't seen. Most scores barely moved, and two got an asterisk. Their own conclusion was careful: either their check had overestimated contamination, or contamination had little effect. In twenty twenty-three, OpenAI's GPT-4 report noted that parts of a test suite called BIG-bench had been mixed into its training data by accident, even though BIG-bench carries a marker string meant to help keep it out. A marker is a label, not a lock.

The second is overfitting: fitting the particular sample instead of the general pattern. A model with enough freedom will find patterns in any sample, including ones that are just accidents of that sample. It learns the noise along with the signal.

With benchmarks, the worry is a whole field doing this slowly, keeping whichever ideas score best on the same public test, year after year. In twenty nineteen, Benjamin Recht and three colleagues at Berkeley tested that worry. They rebuilt the ImageNet image-recognition test from scratch, following the original recipe, and ran the existing models on it. Accuracy fell by eleven to fourteen percentage points. But the models kept almost exactly the same order, and

accuracy gains on the original test sets translate to larger gains on the new test sets.

— Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt and Vaishaal Shankar, 'Do ImageNet Classifiers Generalize to ImageNet?', arXiv 1902.10811, first posted 13 February 2019 (version 2, 12 June 2019; ICML 2019), abstract

Their results suggested the drop came from the new images being slightly harder, not from years of tuning against the old ones. A study of a hundred and twenty competitions on the Kaggle platform, the same year, found little evidence of substantial overfitting either.

The third is saturation: the test runs out of room. In twenty twenty-one, Douwe Kiela of Facebook AI Research and colleagues put it this way.

While it used to take decades for machine learning models to surpass estimates of human performance on benchmark tasks, that milestone is now routinely reached within just a few years for newer datasets

— Douwe Kiela and colleagues, 'Dynabench: Rethinking Benchmarking in NLP', arXiv 2104.14337, 7 April 2021 (NAACL 2021), section 1 'Introduction', page 1

One language test, GLUE, saturated within a year, they wrote. A saturated test hasn't necessarily gone wrong. It has stopped separating the systems people want to compare.

Here's one test's life. G P Q A was published in November twenty twenty-three by researchers led from New York University: graduate-level science questions meant to be Google-proof. Its best-checked set, the Diamond questions, kept questions that experts got right and most skilled non-experts, searching the web, got wrong. GPT-4 scored thirty-nine per cent. In the study's table, first posted in February this year, the five best models scored between eighty-three and eighty-eight. This September it left the index. The knowledge test M M L U had run out of room earlier: Humanity's Last Exam was released in January twenty twenty-five because, its authors said, models now scored over ninety per cent on tests like it.

The hardest maths is following. On the tenth of September, the research group Epoch AI posted that every problem in the top tier of its FrontierMath benchmark had now been solved by AI. Epoch counts a problem as solved once any model has solved it in any attempt. Two things belong beside that. Epoch says OpenAI funded the benchmark and has exclusive access to some of its problems, and the model that solved the last one was OpenAI's. And in June, Epoch addressed errors in forty-two per cent of its problems. Near the top, a test's own mistakes start to blur what a higher score means.

The study's wider results carry a caution. Of its sixty benchmarks, twenty-nine showed high or very high saturation. Among those under two years old, about forty-three per cent were saturated; among those over five, about fifty-five. The authors call that trend modest and not statistically significant, so it's a direction, not a law. Bigger test sets went with less saturation, and so, with age muddying the comparison, did questions written by experts.

So what do you do about a rotting test? Two answers are on offer, and they're aimed at different rots.

The first is to keep the questions out of reach. On the fourth of September, Artificial Analysis put forty per cent of its index's weight on private, held-out tests, double the share before, which it says reduces the ability of labs to game evaluations. By my count of its current table, it's now forty-five. LiveBench was designed the other way, replacing a slice of its questions every month so they postdate a model's training.

The second answer comes from the same study, and it's less comforting about the first. Its four private benchmarks saturated much like the fifty-six public ones.

Hiding test data does not appear to prevent saturation once benchmarks are widely adopted

— Mubashara Akhtar, Anka Reuel and colleagues, 'When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation', arXiv 2602.16763, version 4, 6 August 2026 (ICML 2026), section 4.1, paragraph 'Accessibility and task design', page 6

Four is a small sample, so that's a hint, not a verdict. But the authors also note that even tests built to keep adding new questions can saturate when they can't measure finely enough to separate the leaders, and LiveBench's own changelog describes an update meant to fix saturation and contamination only to some extent. The authors back refreshing too, but their remedies lead with resolution: bigger and harder tests, uncertainty reported beside the score, and rules, written in advance, for revising or retiring a test once it stops separating the leaders.

I think both answers are right, about mostly different diseases. Secrecy and freshness are aimed at contamination, though freshness can slow saturation too. Size, difficulty and honest error bars are what let a test keep separating the leaders. And a private test has a cost of its own: nobody outside can check it. Fresh doesn't mean fair, and private doesn't mean valid.

One. Artificial Analysis's index. The company says version five will raise the private share further and that more interim updates are coming. Through October, watch its changelog: does the private share keep rising past forty-five per cent, which test is retired next, and does it say why?

Two. Whether the live tests stay live. Microsoft's SWE-bench Live promises fifty new Python software issues a month. Check in early November whether October's arrived. LiveBench's public changelog has had no entry since the eighth of January.

Three. Epoch's newer test, FrontierMath Erdős: sixty-eight unsolved problems, launched on the third of September. OpenAI's GPT-6 Astra solved two. Watch how fast that number moves. That's a new test's clock starting.

The idea to keep is that a benchmark score has a shelf life, and three questions tell you how much of it is left. Could the model have seen these questions, or close copies, before the test? Has a whole field been tuning against this one test for years? And can the test still tell the top systems apart, by more than its own margin of error? None of those makes a score worthless. Each tells you how much of its claim is still standing.

To read more: When AI Benchmarks Plateau, by Mubashara Akhtar, Anka Reuel and colleagues, from this year's International Conference on Machine Learning. Its appendix walks through five benchmarks, from saturated to not, in a page.

Sources (23)

  1. Artificial Analysis, Announcing Artificial Analysis Intelligence Index v4.2, artificialanalysis.ai/articles — 4 September 2026
  2. Artificial Analysis, post on X on Intelligence Index v4.2 — 6 September 2026 (00:28 UTC)
  3. Artificial Analysis, Intelligence Benchmarking Methodology, version 4.3.2 — read 1 October 2026
  4. Epoch AI, post on X on FrontierMath Tier 4 — 10 September 2026
  5. Epoch AI, FrontierMath Tier 4 benchmark page and Benchmarks hub — read 1 October 2026
  6. Akhtar, Reuel, Soni, Ahuja et al., When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation, arXiv 2602.16763 v4 (ICML 2026) — 18 February 2026; v4 6 August 2026
  7. Rein et al., GPQA: A Graduate-Level Google-Proof Q&A Benchmark, arXiv 2311.12022 — 20 November 2023
  8. Brown et al. (OpenAI), Language Models are Few-Shot Learners, arXiv 2005.14165 — 28 May 2020
  9. OpenAI, GPT-4 Technical Report, arXiv 2303.08774 — March 2023
  10. BIG-bench, training_on_test_set README (canary GUID), in the BIG-bench repository on GitHub — read 1 October 2026
  11. Jain et al., LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, arXiv 2403.07974 — 12 March 2024; v2 6 June 2024
  12. Zhang et al. (Scale AI), A Careful Examination of Large Language Model Performance on Grade School Arithmetic (GSM1k), arXiv 2405.00332 — 1 May 2024; revised November 2024
  13. International AI Safety Report 2026, arXiv 2602.21012 — February 2026; arXiv 24 February 2026
  14. Recht, Roelofs, Schmidt and Shankar, Do ImageNet Classifiers Generalize to ImageNet?, arXiv 1902.10811 (ICML 2019) — 13 February 2019
  15. Roelofs et al., A Meta-Analysis of Overfitting in Machine Learning, NeurIPS 2019 — December 2019
  16. Kiela et al., Dynabench: Rethinking Benchmarking in NLP, arXiv 2104.14337 — 7 April 2021
  17. Wang et al., arXiv 1804.07461, GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding — 20 April 2018
  18. Wang et al., SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems, arXiv 1905.00537 — 2 May 2019
  19. Microsoft Research, Microsoft DeBERTa surpasses human performance on the SuperGLUE benchmark — 6 January 2021
  20. Gema et al., Are We Done with MMLU?, arXiv 2406.04127 (v3) — 6 June 2024; v3 10 January 2025
  21. Phan et al., Humanity's Last Exam, arXiv 2501.14249 (quoted from v11) — 24 January 2025; v11 28 July 2026
  22. White et al., LiveBench: A Challenging, Contamination-Limited LLM Benchmark, arXiv 2406.19314; LiveBench changelog — 27 June 2024; changelog read 1 October 2026
  23. Microsoft, SWE-bench-Live README, in the SWE-bench-Live repository on GitHub — read 1 October 2026

1. A test retired, a tier finished

Artificial Analysis publishes an "Intelligence Index" that folds a model's results on a basket of tests into a single number. Version 4.2, announced on 4 September, added two tests and removed one. The changelog gave the reason in a line:

GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated

A post on the firm's account on X on 6 September gave a second reason beside the first: the test "provides little signal for frontier models given its saturation, and its multiple-choice format doesn't reflect real tasks in those domains". The firm still runs GPQA Diamond on new models and reports the result on its own; it has left the index, not the firm's test suite.

GPQA was published in November 2023 as "A Graduate-Level Google-Proof Q&A Benchmark" by eight researchers at New York University, two of whom also list affiliations with Cohere and Anthropic. Its Diamond subset holds 198 questions in biology, physics and chemistry, kept, with narrow exceptions, because both expert validators answered them correctly while most skilled non-experts, searching the web without restriction, did not. In the original paper GPT-4 answered 38.8% of the Diamond questions correctly. In the ICML study's table of leading scores, already present in its first version of February 2026, the five best models scored between 82.9% and 87.7%.

Figure 1. GPQA Diamond: from a hard test to a retired one (accuracy, %)
Skilled non-experts with web access (2023)21.9%GPT-4, original paper (November 2023)38.8%Expert validators (2023)81.2%Fifth-best model in the study's table82.9%Best model in the study's table87.7%
Sources: Rein et al., 'GPQA: A Graduate-Level Google-Proof Q&A Benchmark', arXiv 2311.12022, 20 November 2023, Table 5 (GPT-4 few-shot chain-of-thought; validator accuracies, which the authors mark as skewed by selection effects because the Diamond set was chosen using those validators' answers); Akhtar et al., 'When AI Benchmarks Plateau', arXiv 2602.16763 v4, 6 August 2026, Table 6 (top five models on the leaderboard the study used; the same rows appear, as Table 5, in the paper's first version of 18 February 2026, and the study prints no snapshot date). Artificial Analysis removed GPQA Diamond from its Intelligence Index on 4 September 2026. Questions have four answer options, so guessing scores about 25%.
Table view
Figure 1. GPQA Diamond: from a hard test to a retired one (accuracy, %)
Who answered, and whenShare of the 198 Diamond questions answered correctly (%)
Skilled non-experts with web access (2023)21.9%
GPT-4, original paper (November 2023)38.8%
Expert validators (2023)81.2%
Fifth-best model in the study's table82.9%
Best model in the study's table87.7%

Epoch's announcement, posted on 10 September, read: "Every FrontierMath Tier 4 problem has now been solved by AI, with GPT-6 Astra solving the last problem standing." Epoch counts a problem as solved once any model has solved it in any attempt (its October 2025 update reported "the total number ever solved"), which is not the same as one model scoring full marks. Two facts about the test belong beside the headline. Epoch's own page states that "FrontierMath was developed with funding from OpenAI, who has exclusive access to a subset of the benchmark"; in October 2025 Epoch put that subset at 28 of the 48 Tier 4 problems it then evaluated, holding back the other 20; the June revision later removed seven Tier 4 problems. The model that solved the last problem is OpenAI's. And on 12 June 2026, the page records, Epoch "released a major update, addressing errors in 42% of problems" across the benchmark's tiers.

The ICML paper, "When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation", is the work of 37 researchers led by Mubashara Akhtar of ETH Zurich and Anka Reuel of Stanford, assembled through a coalition called EvalEval. First posted in February and revised three times, most recently on 6 August, it examined 60 text-based benchmarks: those used in at least five of 61 model cards and technical reports published between January 2022 and November 2025 by developers including OpenAI, Anthropic, Google, Meta and Alibaba, supplemented by heavily cited benchmark papers and by hand-picked additions to cover particular designs. Its abstract reports that "nearly half of our benchmarks exhibit saturation, with rates increasing with age".

2. Two readings that mislead

2.1 "Saturated means solved"

The natural reading of "saturated" is that the machines have learned the subject and the test has nothing left to ask. The study's definition is narrower:

A benchmark is saturated if the evaluated models cannot be reliably distinguished by their performance scores and any further improvements are not statistically distinguishable under the evaluation protocol.

In practice the authors take the five best scores on a benchmark's leaderboard, compare their spread with the uncertainty that a test of that size carries, and convert the result into an index between 0 and 1; an index of 0.7 or more counts as high saturation. The ceiling against which crowding is judged is "the highest observed model performance", not 100% and not a human score, and the authors note that "human-level performance does not imply saturation", because models can remain statistically distinguishable after passing a human baseline.

The consequence is that saturation can arrive well short of perfection. Of five benchmarks worked through in the paper's appendix, LiveBench, a test designed to resist contamination by changing its questions every month, scored 0.99, the most saturated of the five, with its leaders clustered within about a point of each other at around 79%. The authors read that as "model-level stagnation rather than task completion". Humanity's Last Exam, a deliberately harder test released in January 2025, scored 0.22.

Figure 2. How crowded the top of the leaderboard was: saturation index for five benchmarks
LiveBench (1,000)1.0MATH-500 (500)0.9LiveCodeBench (1,000)0.8TruthfulQA (817)0.6Humanity's Last Exam (2,500)0.2
Source: Akhtar et al., 'When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation', arXiv 2602.16763 v4, 6 August 2026 (ICML 2026), Appendix E, Table 7. The index compares the spread of the five best leaderboard scores with the statistical uncertainty of a test that size; 0.7 and above is 'high', 0.9 and above 'very high'. Spread of the top five, in percentage points: LiveBench 1.09 (at about 79%), MATH-500 1.0 (98.2 to 99.2), LiveCodeBench 3.9, TruthfulQA 6.7, Humanity's Last Exam 11.4. The study prints no snapshot date; this table was added in a revision after its first version of 18 February 2026.
Table view
Figure 2. How crowded the top of the leaderboard was: saturation index for five benchmarks
Benchmark (test questions)Saturation index, 0 = top models well separated, 1 = indistinguishable
LiveBench (1,000)1.0
MATH-500 (500)0.9
LiveCodeBench (1,000)0.8
TruthfulQA (817)0.6
Humanity's Last Exam (2,500)0.2

The authors attach a condition of their own. Saturation, they write, is "not always negative": "if the benchmark was valid", saturation "means that a task can be considered 'solved'". The conditional matters. A score supports a claim about a skill only to the extent that the test measures that skill, and a crowded leaderboard adds no evidence either way on that question. A saturated test of the wrong thing is a test of the wrong thing that everyone now passes. Saturation is also relative to a purpose: a test that no longer separates the strongest systems can still separate weaker ones, or catch a model that has regressed.

The headline count itself rests on a choice of yardstick. The paper's sensitivity analysis shows that the ordering of benchmarks from least to most saturated changes little when its two tuning parameters are varied (rank correlations of 0.88 to 0.92), but that only between 18% and 48% of benchmarks stay in exactly the same one of its five bands; the authors add that "most changes occur between neighbouring bins rather than large shifts". Together with a sample chosen for wide use rather than at random, that makes "twenty-nine of sixty" a reading of about half, on this study's yardstick, rather than a census of the field.

2.2 "So the scores are fake"

The opposite reading treats every high score as memorisation. The measured record does not support it as a rule. When OpenAI checked GPT-3 against its own training data in 2020, it found that potential contamination was "often high (with a quarter of benchmarks scoring over 50%)", yet "in most cases performance changes only negligibly" when the overlapping questions were removed. When Scale AI wrote a fresh set of grade-school mathematics problems in 2024 to test for memorisation of an older set, it found real drops for some model families and "minimal signs of overfitting" for the models at the frontier. When researchers at Berkeley rebuilt the ImageNet test from scratch in 2019, the leading models lost accuracy but kept their order. A score that has rotted has not necessarily been faked; the useful question is which part of the claim still stands.

3. The idea: three ways a score decays

A benchmark score asserts that performance on a particular set of items stands in for a wider skill. The assertion has three weak points. The items may have leaked into what the model learned from. The field may have tuned its choices so closely to those particular items that the score no longer speaks for the skill. Or the items may have run out of power to tell the strongest systems apart. The first is contamination, the second overfitting, the third saturation.

Figure 3. Three points at which a benchmark score stops supporting the claim made for it
A skill people care aboute.g. graduate-level science reasoningA fixed, published set of questionsthe test stands in for the skillModel training and developmentcontamination: the questions leak into trainingdata; overfitting: years of choices tuned to thisone testScores on the testsaturation: the best scores bunch within thetest's margin of errorA claim about the skillholds only as far as each link above holdssampled intomeetsproducesread as
Schematic, not measured data. Each kind of decay breaks a different link, which is why each needs a different check: a fresh matched test for contamination, a re-collected test or a final private ranking for overfitting, and the spread of top scores against their uncertainty for saturation.
Table view
Figure 3. Three points at which a benchmark score stops supporting the claim made for it — stages
#StageNote
1A skill people care aboute.g. graduate-level science reasoning
2A fixed, published set of questionsthe test stands in for the skill
3Model training and developmentcontamination: the questions leak into training data; overfitting: years of choices tuned to this one test
4Scores on the testsaturation: the best scores bunch within the test's margin of error
5A claim about the skillholds only as far as each link above holds
Figure 3. Three points at which a benchmark score stops supporting the claim made for it — connections
FromToLabel
A skill people care aboutA fixed, published set of questionssampled into
A fixed, published set of questionsModel training and developmentmeets
Model training and developmentScores on the testproduces
Scores on the testA claim about the skillread as

3.1 Contamination: the exam paper seen in advance

A student who has seen the exam paper can score well without knowing the subject. Language models learn from enormous scrapes of the internet, and benchmarks are published on the internet. The problem was recognised early. OpenAI's paper introducing GPT-3, posted on 28 May 2020, described an attempt to remove every benchmark's test questions from the training data, and then this:

Unfortunately, a bug in the filtering caused us to ignore some overlaps, and due to the cost of training it was not feasible to retrain the model.

Unable to retrain, the authors measured the damage instead, rescoring each benchmark on a "clean" subset of questions with no long overlap with the training data. Most scores barely moved; two results (PIQA and Winograd) were marked with asterisks, and several language-modelling benchmarks were dropped from the paper because almost all of their text appeared in the training set. The conclusion was carefully hedged: "either our conservative method substantially overestimated contamination or that contamination has little effect on performance".

Good intentions on the benchmark side have not been enough either. BIG-bench, a large collaborative collection of test tasks, embeds a "canary" string whose purpose, its documentation says, is "to allow researchers to better filter BIG-bench tasks out of the training data for large language models". OpenAI's GPT-4 report of March 2023 states that "during our contamination check we discovered that portions of BIG-bench [48] were inadvertently mixed into the training set, and we excluded it from our reported results". A canary is a label and a detector, not a lock.

The cleanest way to see contamination is to date the questions. LiveCodeBench, introduced by researchers at Berkeley, MIT and Cornell in March 2024 and revised that June, collects programming-contest problems with their publication dates. Its authors observed "a stark drop" in the performance of a DeepSeek coding model on problems published after August 2023, just before the model's release, which they read as suggesting "that the earlier problems might indeed be contaminated". The drop appeared mainly on problems from one platform, LeetCode; on other platforms performance was "relatively smooth across the months".

Scale AI's GSM1k study, posted in May 2024 and revised in November, took the matched-fresh-test route. Its authors commissioned 1,205 new grade-school problems in the style of the widely used GSM8K test and found "accuracy drops of up to 8%"—eight percentage points in its appendix tables—with "several families of models showing evidence of systematic overfitting across almost all model sizes". (The first version of the paper said up to 13%; the figure was revised.) Models more likely to reproduce GSM8K's text verbatim tended to show bigger gaps, a correlation the authors read as suggesting "that some models may have partially memorized GSM8k". They also found that "many models, especially those on the frontier, show minimal signs of overfitting".

Figure 4. The same kind of question, old and new: accuracy on GSM8K and on the freshly written GSM1k (%)
Yi-6B-Chat: GSM8K43.7%Yi-6B-Chat: GSM1k35.7%Meta-Llama-3-8B-Instruct: GSM8K75.2%Meta-Llama-3-8B-Instruct: GSM1k69%gpt-4o: GSM8K93.1%gpt-4o: GSM1k92.9%gpt-4: GSM8K91.1%gpt-4: GSM1k92.3%
Source: Zhang et al. (Scale AI), 'A Careful Examination of Large Language Model Performance on Grade School Arithmetic', arXiv 2405.00332 v4, 22 November 2024, Appendix F, standard-prompt results (API models queried between 16 April and 10 July 2024; identifiers as the authors report them). Darker bars: GSM8K, the public test; lighter bars: GSM1k, 1,205 newly written problems matched in style and difficulty, which the authors describe as 'only highly similar, but not identically distributed'. The largest gap in the table is Yi-6B-Chat's 8.0 points; gpt-4 scored higher on the new set. A gap invites investigation; it does not by itself separate memorisation from other differences between the sets.
Table view
Figure 4. The same kind of question, old and new: accuracy on GSM8K and on the freshly written GSM1k (%)
Model and testAccuracy (%)
Yi-6B-Chat: GSM8K43.7%
Yi-6B-Chat: GSM1k35.7%
Meta-Llama-3-8B-Instruct: GSM8K75.2%
Meta-Llama-3-8B-Instruct: GSM1k69%
gpt-4o: GSM8K93.1%
gpt-4o: GSM1k92.9%
gpt-4: GSM8K91.1%
gpt-4: GSM1k92.3%

The International AI Safety Report, published in February 2026 under the chairmanship of Yoshua Bengio, put the general problem plainly: "many models may have been trained using data from these same benchmarks – a problem called 'data contamination', which most developers do not currently track or disclose".

3.2 Overfitting: fitting the sample, not the pattern

A model with enough freedom will find patterns in any sample, including patterns that are accidents of that sample. The defence is to judge the model on data it was not fitted to—a held-out test set. Benchmarks are held-out test sets shared by a whole field, which creates a subtler version of the problem. If many researchers try many ideas and keep the ones that score best on the same public test, the test is no longer truly held out: the field as a whole has been fitted to it. Researchers call this adaptive overfitting, "overfitting caused by test set reuse".

The fear is reasonable; the evidence on its size is more reassuring than the fear. In 2019 Benjamin Recht and three colleagues at the University of California, Berkeley rebuilt the test set of ImageNet, the image-recognition benchmark on which a decade of progress had been reported, by following the original collection procedure as closely as possible. Accuracy fell by 11 to 14 percentage points, a loss the authors equated to "approximately five years of progress". Yet the models kept almost exactly the same order, and

accuracy gains on the original test sets translate to larger gains on the new test sets.

Their results "suggest that the accuracy drops are not caused by adaptivity", but by the models' inability to handle slightly harder images than those in the original test. The authors add that "it is unclear what properties of our new images cause the accuracy drops". A companion study of 120 competitions on Kaggle, a platform where teams submit repeatedly against a public leaderboard before a final private ranking, found "somewhat surprisingly, little evidence of substantial overfitting". The fall on a fresh test was real; the evidence pointed away from years of tuning to the old test as its cause.

Whether that reassurance carries over to language models is open. The ImageNet and Kaggle studies examined test sets that were reused but did not leak into training data. Public language-model benchmarks face both problems at once, and GSM1k's family-specific gaps suggest that the combination can matter.

3.3 Saturation: a test that runs out of room

Every test has a ceiling. In 2021 Douwe Kiela of Facebook AI Research and colleagues described how quickly benchmarks now reach theirs:

While it used to take decades for machine learning models to surpass estimates of human performance on benchmark tasks, that milestone is now routinely reached within just a few years for newer datasets

Their figure traced six benchmarks from first result to the human estimate. Handwritten-digit recognition (MNIST) and conversational speech transcription (Switchboard) took about two decades; ImageNet took six years; the reading-comprehension test SQuAD took two; its successor, and the language-understanding suite GLUE, took about one.

Figure 5. Years from a benchmark's first plotted result to its first result at or above the human estimate
MNIST, digits (1998)20yrsSwitchboard, speech (1998)19yrsImageNet, images (2009)6yrsSQuAD 1.1, reading (2016)2yrsSQuAD 2.0, reading (2018)1yrsGLUE, language (2018)1yrs
Source: Kiela et al., 'Dynabench: Rethinking Benchmarking in NLP', arXiv 2104.14337, 7 April 2021, Figure 1 ('Benchmark saturation over time for popular benchmarks, normalized with initial performance at minus one and human performance at zero'). The paper prints the figure but not its data; the plotted points were recovered from the PDF's vector drawing and are accurate to about a year. MNIST and Switchboard begin at the plot's left edge, so their spans are lower bounds; MNIST sat just below the human mark from 2013 and crossed it at its next plotted point, 2018. 'Human performance' is each benchmark's own estimate, which Kiela et al. call 'the narrow criteria used to define human performance'.
Table view
Figure 5. Years from a benchmark's first plotted result to its first result at or above the human estimate
Benchmark (first plotted year)Years to reach the human estimate (approx.)
MNIST, digits (1998)20yrs
Switchboard, speech (1998)19yrs
ImageNet, images (2009)6yrs
SQuAD 1.1, reading (2016)2yrs
SQuAD 2.0, reading (2018)1yrs
GLUE, language (2018)1yrs

GLUE shows the cycle in its own documents. Its launch paper, posted in April 2018, concluded that "solving GLUE is beyond the capabilities of current models and methods". A year later the paper introducing its successor, SuperGLUE, reported that "performance on the benchmark has recently surpassed the level of non-expert humans, suggesting limited headroom for further research". On 6 January 2021 Microsoft announced that a single DeBERTa model had passed SuperGLUE's human baseline "for the first time in terms of macro-average score (89.9 versus 89.8)", and added that "the model is by no means reaching the human-level intelligence of NLU".

Two further features make saturation harder to read than a full score suggests. The first is that the ceiling is often set by the test's own errors. A re-annotation of MMLU, a widely used multiple-choice knowledge test, first posted in June 2024, estimates in its current version (January 2025) that "6.49% of MMLU questions contain errors", and found errors in 57% of the questions it analysed in the virology section; at the other end of the difficulty range, FrontierMath's June revision addressed errors in 42% of its problems. Near the top, a higher score can mean agreeing with a wrong answer key, and a correct answer can be marked wrong, so errors blur the ceiling rather than fixing it at a number. The second is that saturation drives replacement. Humanity's Last Exam, released in January 2025, opens by observing that "LLMs now achieve over 90% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities", and offers 2,500 expert-written questions "designed to be the final closed-ended academic benchmark of its kind".

4. What happened, 2018–2026

Date Event Decay it illustrates
April 2018 GLUE launched; "beyond the capabilities of current models and methods" saturation
February 2019 ImageNet re-collected: accuracy down 11–14 points, order preserved overfitting (largely not found)
May 2019 SuperGLUE launched because GLUE's human baseline had been passed saturation
May 2020 GPT-3's filtering bug; most clean-subset scores barely move contamination
January 2021 DeBERTa passes SuperGLUE's human baseline, 89.9 to 89.8 saturation
March 2023 GPT-4 report: parts of BIG-bench "inadvertently mixed into the training set" contamination
November 2023 GPQA published; GPT-4 scores 38.8% on its Diamond set —
March 2024 (revised June) LiveCodeBench dates its problems; drops after models' cutoffs contamination
May 2024 (revised November) GSM1k: drops of up to 13% in its first version, up to 8% after revision; minimal at the frontier contamination and overfitting
January 2025 Humanity's Last Exam, built because MMLU passed 90% saturation
February–August 2026 "When AI Benchmarks Plateau": 29 of 60 benchmarks highly saturated saturation
June 2026 FrontierMath v2 addresses errors in 42% of problems the ceiling set by errors
4 September 2026 Artificial Analysis removes GPQA Diamond from its index saturation
10 September 2026 Epoch: every FrontierMath Tier 4 problem solved by some AI saturation

The ICML study's wider results fit the pattern, with a caution attached. Of its 60 benchmarks, 29 showed high or very high saturation, 14 of them very high. Among benchmarks released within the previous 24 months, 42.9% were saturated; among those more than 60 months old, 54.5%. The authors describe that trend as "modest and not statistically significant at conventional thresholds", though in their joint statistical model benchmark age and test-set size "show the most consistent effects". Larger test sets went with less saturation, and expert-written benchmarks were less saturated than crowdsourced ones of a similar age, though the authors note that curation categories also differ in age; two expert-curated tests, ARC-AGI and BIG-Bench Hard, "remain unsaturated despite prolonged exposure". How often a benchmark was cited did not predict saturation once its age was taken into account. The paper's own causal statement is hedged: its results are "consistent with" repeated optimisation compressing the differences between frontier models, "although our analysis does not directly identify the causal mechanism".

Figure 6. Sixty benchmarks, one yardstick: the ICML 2026 saturation study in four numbers
29 of 60
Benchmarks with high or very high saturation
14 of them very high (index 0.9 or more)
42.9%
Share saturated, released within 24 months
against 54.5% for those over 60 months old; trend not statistically significant
4 of 60
Private test sets in the sample
saturated much like the 56 public ones
18–48%
Benchmarks keeping their band when the method's settings change
the ordering is far more stable (rank correlation 0.88–0.92)
Source: Akhtar et al., 'When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation', arXiv 2602.16763 v4, 6 August 2026, published at ICML 2026: section 4 (counts and age bins), section 4.1 (public versus private), Table 1 (sensitivity to the parameters k and alpha). The sample is purposive, not random: benchmarks used in at least five of 61 model reports from January 2022 to November 2025, heavily cited benchmark papers, and additions chosen to cover particular designs. All four numbers are from version 4; the first version of 18 February 2026 counted 8 private benchmarks and had no sensitivity table. The study prints no snapshot date.
Table view
Figure 6. Sixty benchmarks, one yardstick: the ICML 2026 saturation study in four numbers
MeasureValue
Benchmarks with high or very high saturation29 of 60
Share saturated, released within 24 months42.9%
Private test sets in the sample4 of 60
Benchmarks keeping their band when the method's settings change18–48%

5. Two postures: hide and refresh, or resolve and retire

The responses on offer divide into two, and they are aimed at different kinds of decay.

The first is to keep the questions out of reach of training, either by hiding them or by replacing them. Artificial Analysis's version 4.2 raised the share of its index weighted on private, held-out tests to 40%, "double the figure from v4.1", and said that "This reduces the ability for labs to game evaluations"; the firm says the share "will increase further in Index v5". Its current methodology table (version 4.3.2) marks tests carrying 45% of the weight as private, by a sum of the weights it lists; the page itself prints no total. LiveBench, introduced in June 2024, replaces about a sixth of its questions in each update, so that it is "fully refreshed roughly every 6 months", and withholds each month's new questions for a month. The current arXiv version of its paper (April 2025) calls it "Contamination-Limited", although the citation in the project's own README still reads "Contamination-Free"; its changelog entry of 25 November 2025 describes an update made "to resolve (to some extent) saturation and contamination of the benchmark". Microsoft's SWE-bench-Live, a test built from real software-repository issues, has promised since September 2025 that "each month, we will add 50 newly verified, high-quality issues to the dataset test split", while freezing two smaller splits so that leaderboard comparisons stay fair.

The second posture comes from the ICML study, and it is less comforting about the first. The four private benchmarks in its sample behaved like the 56 public ones:

Hiding test data does not appear to prevent saturation once benchmarks are widely adopted

Four is a small group, and the comparison is suggestive rather than conclusive. The worked examples in its appendix, chosen to span the scale rather than to rank designs, include a related caution: LiveCodeBench, the authors write, "demonstrates that also dynamically constructed benchmarks can saturate when evaluation resolution is limited". Their recommendations lead with resolution: larger test sets, harder items or scores broken down by sub-skill, so that "score differences between models exceed expected evaluation uncertainty"; confidence intervals and the spread of top scores reported beside the peak; and explicit criteria, written in at design time, for revising or retiring a test once the best systems can no longer be told apart. They endorse refreshing as well, but as one tool among several.

Private tests also carry a cost the public kind does not: outsiders cannot check them. FrontierMath is largely private, yet its funder holds a subset; Artificial Analysis's held-out tests cannot be inspected by the public. Moving questions out of view moves trust from the test to whoever holds it.

The criterion that fits the evidence is that the postures are complements, each aimed mainly at a different decay. Secrecy and freshness are aimed at contamination, and the study suggests that refreshing can also slow the convergence that comes with exposure. Size, difficulty and honest error bars are what let a test keep separating the strongest systems. Re-collected tests and final private rankings detect overfitting. A benchmark programme that does only one of these is defending against one failure while the other two proceed.

6. What to watch

Artificial Analysis's next index version, from October 2026. Its methodology page stood at version 4.3.2 on 1 October; the firm has promised further "incremental releases" and a version 5 in which the private share "will increase further". The checkable questions are whether the private share, 45% of the weight by the current table, keeps rising, whether another constituent is retired as saturated, and whether each retirement comes with its reason. Humanity's Last Exam, the least saturated of the study's worked examples, still carries 10% of the index's weight.

FrontierMath Erdős, Epoch's newer test, through October. Launched on 3 September with 68 unsolved "Erdős problems"—open questions associated with the mathematician Paul Erdős—whose solutions must be written as machine-checkable proofs in the Lean language, it began with GPT-6 Astra solving 2 of 68; Epoch says no earlier model solved any. The rate at which that number moves is the start of a new test's clock.

Whether the refreshing designs refresh, by early November. SWE-bench-Live's stated policy, since September 2025, is 50 new verified Python issues a month; its most recent dataset news, on 21 August 2026, reported 1,077 task instances in its multi-language edition. LiveBench's public changelog has no entry after 8 January 2026, although its code repository is still being updated. A "live" benchmark that stops adding questions becomes a static one, and decays like one.

7. The idea to keep

A benchmark score has a shelf life. Three questions establish how much of its claim is still standing. Could the model have seen these questions, or close copies of them, before it was tested? Has a whole field been tuning its choices against this one test for years? And can the test still tell the best systems apart by more than its own margin of error? None of the three makes a score worthless; each locates the part of it that has decayed. A thorough current treatment is Akhtar, Reuel and colleagues' "When AI Benchmarks Plateau" (ICML 2026), whose appendix walks through five benchmarks, from saturated to not, in a page; for the opposite caution, Recht and colleagues' "Do ImageNet Classifiers Generalize to ImageNet?" (2019) remains a clear demonstration that a falling score is not always a fraud uncovered.

Next lesson — Day 21: When the Metric Becomes the Game

Sources

Source Date
Artificial Analysis, Announcing Artificial Analysis Intelligence Index v4.2, artificialanalysis.ai/articles 4 September 2026
Artificial Analysis, post on X on Intelligence Index v4.2 6 September 2026 (00:28 UTC)
Artificial Analysis, Intelligence Benchmarking Methodology, version 4.3.2 read 1 October 2026
Epoch AI, post on X on FrontierMath Tier 4 10 September 2026
Epoch AI, FrontierMath Tier 4 benchmark page and Benchmarks hub read 1 October 2026
Akhtar, Reuel, Soni, Ahuja et al., When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation, arXiv 2602.16763 v4 (ICML 2026) 18 February 2026; v4 6 August 2026
Rein et al., GPQA: A Graduate-Level Google-Proof Q&A Benchmark, arXiv 2311.12022 20 November 2023
Brown et al. (OpenAI), Language Models are Few-Shot Learners, arXiv 2005.14165 28 May 2020
OpenAI, GPT-4 Technical Report, arXiv 2303.08774 March 2023
BIG-bench, training_on_test_set README (canary GUID), in the BIG-bench repository on GitHub read 1 October 2026
Jain et al., LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, arXiv 2403.07974 12 March 2024; v2 6 June 2024
Zhang et al. (Scale AI), A Careful Examination of Large Language Model Performance on Grade School Arithmetic (GSM1k), arXiv 2405.00332 1 May 2024; revised November 2024
International AI Safety Report 2026, arXiv 2602.21012 February 2026; arXiv 24 February 2026
Recht, Roelofs, Schmidt and Shankar, Do ImageNet Classifiers Generalize to ImageNet?, arXiv 1902.10811 (ICML 2019) 13 February 2019
Roelofs et al., A Meta-Analysis of Overfitting in Machine Learning, NeurIPS 2019 December 2019
Kiela et al., Dynabench: Rethinking Benchmarking in NLP, arXiv 2104.14337 7 April 2021
Wang et al., arXiv 1804.07461, GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding 20 April 2018
Wang et al., SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems, arXiv 1905.00537 2 May 2019
Microsoft Research, Microsoft DeBERTa surpasses human performance on the SuperGLUE benchmark 6 January 2021
Gema et al., Are We Done with MMLU?, arXiv 2406.04127 (v3) 6 June 2024; v3 10 January 2025
Phan et al., Humanity's Last Exam, arXiv 2501.14249 (quoted from v11) 24 January 2025; v11 28 July 2026
White et al., LiveBench: A Challenging, Contamination-Limited LLM Benchmark, arXiv 2406.19314; LiveBench changelog 27 June 2024; changelog read 1 October 2026
Microsoft, SWE-bench-Live README, in the SWE-bench-Live repository on GitHub read 1 October 2026

Day 17 is written and not yet available here.