The Limits of Benchmarks as Value Signals

·Research
A light monochrome line chart titled "The frontier is converging": eight schematic AI-model score lines rise out of a wide 2023 spread and bunch tightly together just under a dashed benchmark-saturation ceiling by 2025, with a callout reading "Top-2 model gap 4.9% to 0.7%" sourced to the 2025 Stanford HAI AI Index.

Introduction

Every few months a new model tops the leaderboard, and every few months the headline matters a little less. The scores are real and they are still climbing, but the two things they are supposed to connect - a higher benchmark and a better outcome for the business or the person using the tool - have quietly come apart. The frontier models have converged to within a fraction of a point of each other; the benchmarks that separate them saturate within months; and the enterprises deploying them are learning, expensively, that the model was never the thing that decided whether the project worked. The value of AI does not live in the leaderboard rank. It lives in the application built around the model: the harness, the data, the workflow, and the evaluation that turn a capability into delivered work.

  • The convergence:the gap between the best model and the tenth-best has nearly vanished, so “buy the top model” is no longer a source of advantage.
  • The measurement: benchmarks saturate fast, are gameable, and frequently fail to measure the thing a buyer actually cares about.
  • The ROI gap: when enterprise AI fails, independent studies locate the cause in workflow and data integration, not in the intelligence of the model.
  • The fix: hold the model roughly fixed and win on the system around it - the same lesson that turns a chat window into a coding agent.

The frontier has converged

Start with the leaderboard on its own terms, because the most striking thing about it in 2025 is how little it now separates the leaders. Stanford’s AI Index, the field’s most-cited annual accounting of progress, reports that on the Chatbot Arena leaderboard the gap between the top two models shrank from 4.9% in 2023 to just 0.7% in 2024, and the gap between the first and tenth-ranked models fell from 11.9% to 5.4% in a single year. The open-weight models closed in just as fast: the best open model trailed the best closed model by 8.04% in January 2024 and by only 1.70% by February 2025. When the first five names on the board are separated by rounding error, and a freely downloadable model is within two points of the most expensive API, the practical question for a buyer changes. You are no longer choosing between a capable model and an incapable one. You are choosing between models that are all, for almost any business task, capable enough.

This is what erodes the marginal value of a benchmark point. A score that used to sort the usable models from the unusable ones now sorts the excellent from the slightly-more-excellent - a distinction that a procurement team can feel on a demo but rarely on a profit-and-loss statement. The differentiation did not disappear. It moved. It moved off the leaderboard and into everything that surrounds the model, which is exactly the territory a score cannot see.

Benchmarks saturate faster than they can be built

The convergence has a companion problem: the tests themselves wear out. The AI Index is blunt that the classic suite has been used up, naming the saturation of traditional benchmarks like MMLU, GSM8K, and HumanEval as the reason researchers keep having to invent new ones. Models now exceed 90% on MMLU, which is why the authors of Humanity’s Last Exam, a 2,500-question expert benchmark, begin their paperby noting that “LLMs now achieve over 90% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities.” The benchmarks introduced in 2023 to be hard - MMMU, GPQA, and SWE-bench - did not stay hard for long: within a year scores rose by 18.8 and 48.9 percentage points on the first two, and on SWE-bench, a test of real GitHub bug-fixing, systems went from solving 4.4% of the problems to 71.7% - a jump of roughly sixty-seven points in twelve months.

A test that goes from discriminating to saturated in a year carries less and less information at the top. To keep measuring the frontier at all, the community now has to build exams designed to be nearly unsolvable: Humanity’s Last Exam, where the best system scored just 8.80% at the 2025 AI Index cutoff, and FrontierMath, where state-of-the-art models solved under 2% of problems at its late-2024 release. Those floor scores are snapshots, not verdicts - reasoning models have since climbed well above them, and that is the point of the next section. What the floors show is not that the models are weak but that a saturated benchmark is a fast-decaying signal: whatever a leaderboard tells you about a model this quarter, it will tell you less next quarter, and it never told you much about your data.

What a high score misses

Even before a benchmark saturates, there is a deeper question of whether the number means what it appears to mean. A 2025 systematic review put this to the test at scale: twenty-nine reviewers examined 445 LLM benchmarks from the leading NLP and machine-learning venues and found patterns across what they measure, how they task the model, and how they score it that, in the authors’ words, “undermine the validity of the resulting claims.” A benchmark is a proxy, and a proxy is only as good as its construction; a great many are measuring something other than the capability printed on the label.

The most-quoted AI benchmark headline of the era makes the concrete case. When GPT-4 launched, its reported 90th-percentile score on the Uniform Bar Exam became shorthand for “smarter than most lawyers.” A peer-reviewed reanalysis in the journal Artificial Intelligence and Law took the same raw performance and re-scored it against meaningful comparison groups. Against everyone who sat the exam the model landed below the 69th percentile, and on the open-ended essays - the part that most resembles legal work - it fell to roughly the 48th percentile; measured only against those who actually passed, it sat in the 15th percentile on the essays. The capability was the same in every case. Only the framing changed. A percentile is a marketing decision as much as a measurement, and a buyer who reads “top 10%” as “fit for my work” is reading a great deal into a single number.

And where a benchmark becomes valuable enough to win, it becomes valuable enough to game - Goodhart’s law, on schedule. The Chatbot Arena leaderboard is the one businesses cite most, and a 2025 study from Cohere Labs, The Leaderboard Illusion, documented how it can be worked: the authors identified 27 private model variantsthat Meta tested in the run-up to Llama-4, showed that a handful of large providers received an outsized share of the arena’s user data, and found that this access alone could produce relative performance gains of up to 112% on the arena’s own distribution - “overfitting to Arena-specific dynamics rather than general model quality.” The arena’s own maintainers, in a separate analysis, showed the softer version of the same problem: answer length and formatting, not correctness, drive a large share of the human preference votes, and once you statistically control for style the ranking visibly reorders. A model can climb the board by writing longer, better-formatted answers without getting any more reliable at the task in front of it.

The ROI gap is a workflow gap, not a model gap

If the leaderboard were a good proxy for value, better models would be producing better business results. They are not, and the studies that measure enterprise outcomes are unusually specific about why. MIT’s Project NANDA, in its 2025 report The GenAI Divide, found that despite $30-40 billion in enterprise investment, 95% of organizations were getting zero return, with only about 5% of integrated pilots capturing real value. The report’s own diagnosis is the load-bearing sentence for this whole argument: the divide “does not seem to be driven by model quality or regulation, but seems to be determined by approach.” Its lead author put the mechanism plainly - generic tools “stall in enterprise use since they don’t learn from or adapt to workflows,” a learning gap rather than an intelligence gap.

That is not one contrarian study. Boston Consulting Group, surveying a thousand executives across fifty-nine countries, found that 74% of companies had yet to show tangible valuefrom AI and that only 26% had built the capabilities to move beyond proofs of concept, with a bare 4% at the frontier. Every one of those companies can call the same frontier models; the 74/26 split is decided by organizational capability, not by which API they hold. Gartner, forecasting from the analyst’s chair, predicted that at least 30% of generative-AI projects would be abandoned after proof of concept by the end of 2025, and the causes it named - poor data quality, inadequate risk controls, escalating costs, unclear business value - are every one an application-layer problem. Model capability is conspicuously absent from the list. Three independent methodologies, one field study, one thousand-executive survey, and one analyst forecast, arrive at the same address.

Value lives in the harness

The clearest way to see that the value is in the application is to hold the model fixed and change only the system around it. Andrew Ng ran exactly that comparison and reported it in The Batch: on the HumanEval coding benchmark, GPT-3.5 answered 48.1% of problems correctly in a single zero-shot pass and GPT-4 answered 67.0%. But wrapping the weakermodel, GPT-3.5, in an iterative agent loop - letting it plan, call tools, run its own code, and revise - took it up to 95.1%. The upgrade from GPT-3.5 to GPT-4 bought about nineteen points; the harness around GPT-3.5 bought forty-seven. As Ng put it, the model upgrade was “dwarfed by incorporating an iterative agent workflow.”

This is not a prompt-engineering trick; it is a research result about where performance comes from. A team from Princeton and Stanford built SWE-agent to study exactly that question, designing a purpose-built “agent-computer interface” - the commands a model uses to open files, edit code, navigate a repository, and run tests - and holding the underlying model constant. The interface alone lifted the system to 12.5% pass@1 on SWE-bench and 87.7% on HumanEvalFix, “far exceeding the previous state-of-the-art achieved with non-interactive” language models. The paper’s subject is not a smarter model. It is the harness. The people building at the frontier say the same thing in the open: Ng has predicted that agent workflows will drive more near-term progress than the next generation of foundation models, and Andrej Karpathy, describing the shape of a modern AI product, locates the work in orchestrating multiple model calls, heavy context management, and an “autonomy slider” with fast human verification - product engineering, not raw model IQ.

The tool you may be reading this in is the proof by demonstration. A coding agent like Claude Code is, at its center, the same foundation model available in a chat window. What makes it useful is nearly all scaffolding: the ability to read and search a real codebase, run commands and see their output, execute tests and act on failures, hold context across many steps, ask permission before touching anything irreversible, and loop until the work is actually done. Swap the model underneath for the next checkpoint and the tool gets incrementally better; remove the harness and you are back to a chat window that can describe a fix but cannot make one. The capability is necessary. The application is what converts it into work.

What the leaderboard measuresWhat decides business value
Unit of comparisonA model, in isolation, on a fixed test setA system, on your data, in your workflow
How it improvesNext checkpoint, next training runRetrieval, tools, evaluation, and integration
Shelf life of the signalMonths - the benchmark saturatesCompounds as the data and harness mature
How much a top rank helpsMarginal - the field has convergedDecisive - most competitors never build it
What it can be gamed byData access, answer style, contaminationOnly by actually working on the task

Where benchmarks still matter

An argument worth making has to concede what is true on the other side, and there is a real other side. Benchmarks are not decoupled from capability; they track it, imperfectly, and the capability is genuinely rising. METR, measuring the length of task a model can complete on its own, found that this horizon has been doubling roughly every seven months for six years - a property of the base model, not of any harness, and the precondition that makes long-horizon agents viable at all. Some capability jumps have delivered value no amount of scaffolding could have manufactured: the o1-style reasoning advance improved over its predecessor by 43 percentage points on the AIME 2024 math exam and, in Ng’s own newsletter, is credited with immediate downstream gains “in math and coding performance, more accurate answers, more capable robots, and rapid progress in AI agents.” Long context is the same story: Gemini 1.5 Pro’s near-perfect recall above 99.7% out to a million tokens across text, video, and audio (holding 99.2% out to ten million tokens on text) opened up whole-codebase and book-length applications that a shorter-context model simply cannot do, no matter how it is wrapped.

So the precise claim is not that benchmarks are meaningless or that model progress has stopped. It is that a benchmark still validly signals when a new capability frontier has arrived - SWE-bench going from 4.4% to 71.7% in a year is the sound of coding agents becoming possible - while the marginalbusiness value of each additional point past “capable enough” is falling fast. Capability is the precondition. It is necessary, and it is not sufficient, and it is not the part most buyers are short of.

Build the system, don’t chase the leaderboard

Put the pieces together and the strategy writes itself. The models have converged, so picking the highest-scoring one buys little. The benchmarks that rank them saturate quickly, miss what they claim to measure, and can be gamed. And when AI fails to pay off inside a real organization, the cause is almost never the model - it is the absence of the data, integration, and workflow the model needs to be useful. The leverage has moved from model selection to system building, and the evidence that the system is where value lives is now hard to miss: the same model, given a harness, does work it could not do without one.

This is the discipline Ardenus brings to the operators who run the physical economy - the field-service businesses most likely to be sold a chatbot and least equipped to make one work. Rather than bolt a leaderboard-topping model onto a fragmented data estate and hope, Ardenus builds the intelligence layer around it: unifying the data a business already generates, governing it, modeling it into an ontology, retrieving the right slice per question, and acting through governed workflows. The model matters, but only as one component of a system engineered to turn capability into outcomes. We make the companion argument - that the model is rarely the thing that fails - in our research on the data foundation beneath enterprise AI, and you can see the platform itself on the technology page. Pick a capable-enough model. Then win on the system you build around it.

Sources and methodology

  1. 2025 AI Index Report, Technical Performance chapter, Stanford Institute for Human-Centered AI (2025) - model convergence (top-2 Chatbot Arena Elo gap 4.9% to 0.7%; open-weight vs closed-weight 8.04% to 1.70%), benchmark saturation, the MMMU/GPQA/SWE-bench single-year gains, and the Humanity’s Last Exam 8.80% figure.
  2. Humanity’s Last Exam (Phan et al., Center for AI Safety and Scale AI, 2025) and FrontierMath (Glazer et al., Epoch AI, 2024) - purpose-built harder benchmarks; the under-2% and 8.80% floor scores are dated to their 2024/2025 release.
  3. Measuring what Matters: Construct Validity in Large Language Model Benchmarks (Bean et al., NeurIPS 2025 Datasets & Benchmarks Track) - the 29-reviewer systematic review of 445 benchmarks.
  4. Re-evaluating GPT-4’s Bar Exam Performance (Eric Martinez, Artificial Intelligence and Law, 2024) - the 90th-percentile claim re-scored to below the 69th percentile overall and ~48th on the essays.
  5. The Leaderboard Illusion (Singh et al., Cohere Labs, 2025) - the 27 private Meta variants and up to 112% relative gains from arena-data access; and Does Style Matter? Disentangling Style and Substance in Chatbot Arena (Li, Angelopoulos & Chiang, LMSYS, 2024).
  6. The GenAI Divide: State of AI in Business 2025, MIT Project NANDA (2025), as reported by Fortune - the 95% of organizations with zero return and the “determined by approach” diagnosis.
  7. AI Adoption in 2024: 74% of Companies Struggle to Achieve and Scale Value, Boston Consulting Group (2024), and Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept (2024).
  8. Agentic Design Patterns Part 1 (Andrew Ng, The Batch, 2024) - GPT-3.5 in an agent loop at 95.1% on HumanEval versus the 48.1% to 67.0% raw model upgrade; and The Batch, Issue 241 (2024).
  9. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (Yang et al., NeurIPS 2024) - the interface, not the model, driving 12.5% pass@1 on SWE-bench; and Andrej Karpathy, Software Is Changing (Again) (Y Combinator, 2025).
  10. Measuring AI Ability to Complete Long Tasks (METR, 2025); The Batch on reasoning models (DeepLearning.AI, 2025); and the Gemini 1.5 Technical Report (Google DeepMind, 2024) - the counter-case that base-model capability still raises the ceiling.