The Wire — Signal
August 5, 2026
Latest edition All past editions
1 item stale
This edition is more than 48 hours old — labeled here rather than hidden, per this site's real-data-or-no-data rule.
A benchmark score is not a fact about a model
Why it matters: the two most-cited independent leaderboards run the same 89 tasks through the same agent harness and disagree by more than four points — several times the gap between the models they are ranking.
Terminal-Bench 2.1 has become the default citation for agentic coding ability, and two independent evaluators publish scores for it. Vals AI puts GPT-5.6 Sol on top at 85.77%, with Claude Opus 5 at 84.64%. Artificial Analysis also puts GPT-5.6 Sol on top — at 89.5%, with Claude Opus 5 at 89.1%. Same 89 tasks, same Terminus 2 harness, both pass@1, both run independently, both documented openly.
Do the arithmetic on the disagreement rather than the ranking. Sol moves 3.73 points between the two boards; Opus 5 moves 4.46. The gap between the two models is 1.13 points on one board and 0.40 on the other. The spread between measurements is three to eleven times the difference being measured. Any sentence of the form "X beats Y on Terminal-Bench" is reporting something smaller than the noise of the method that produced it.
That spread is not sloppiness, and both pages say plainly what they did. Vals runs the benchmark remotely on Daytona under a time limit. Artificial Analysis runs Terminus 2 in an e2b sandbox and averages pass@1 over three repeats per task. Artificial Analysis also puts the reasoning configuration in the model name itself — "GPT-5.6 Sol (xhigh)", "Claude Opus 5 (Adaptive Reasoning, Max Effort)" — where Vals states no effort setting at all. Those are legitimate, disclosed choices. They are also the entire disagreement.
One disclosure is sharper than the rest. Vals notes that the Opus 5 run used Claude Opus 4.8 as a fallback for refusals, that nine passing task results across three runs were affected, and that counting those as failures moves the score from 84.64% to 81.27%. That single accounting decision is worth 3.37 points — larger than the model gap on either leaderboard. Vals publishes it and ships a toggle to view it both ways, which is the right behaviour. The number that travels is still 89.1, or 84.64, or 81.27, depending on which page someone had open.
None of this makes either leaderboard bad. Both are more careful than the vendor charts they displaced, and in both cases the provenance is right there on the page. The failure is downstream, and it is a stripping problem: the method does not survive the number leaving the page it was measured on. "Opus 5 scores 89.1% on Terminal-Bench" silently asserts a sandbox, a repeat count, a reasoning-effort setting and a refusal-accounting policy, while containing none of them.
The discipline that follows is not distrust — it is carrying the method with the number: which board, which harness, which configuration, which date. That is why this site's own pricing comparator prints a verification date beside every figure. And there is a practical test for anyone actually picking a model on these numbers: if the decision flips on four points, you are not choosing between models, you are choosing between two evaluation setups. Run the benchmark on the harness you intend to ship on. That result is the only one that is a fact about your system.
Read the original A benchmark score is not a fact about a model