Within Explosion Evidence
What AI Tests Miss About Intelligence
AI benchmarks show rapid capability gains, but they do not fully reveal whether systems can achieve reliable open-ended intelligence.
On this page
- What current evaluations measure
- Where benchmark gains fall short
- The search for better capability signals
Page outline Jump by section
Introduction
AI benchmarks have become one of the main ways the world measures progress towards more capable artificial intelligence. Results on tests such as mathematics exams, coding challenges, scientific question sets and reasoning evaluations show that leading systems have improved rapidly. But benchmark scores are not the same thing as intelligence itself. They measure performance on selected tasks under controlled conditions, while real intelligence also involves reliability, adaptability, judgement, learning from experience and the ability to pursue goals in messy environments.
This measurement gap matters for the intelligence explosion debate. If AI systems are becoming better at solving problems, they could eventually accelerate science, engineering and civilisation-wide innovation. Yet the evidence for such a transition depends partly on whether benchmark improvements represent deeper, general capabilities or only narrow gains on increasingly specialised tests. The central question is therefore not simply “are AI scores rising?” but “what kind of intelligence are those scores actually measuring?”[International AI Safety Report]internationalaisafetyreport.orginternational ai safety report 2025International AI Safety ReportInternational AI Safety Report 2025 | International AI Safety ReportJanuary 29, 2025…
What current evaluations measure
Modern AI benchmarks are designed to turn difficult questions about capability into measurable comparisons. They allow researchers to track whether a new model is better than earlier systems and whether progress is occurring in areas relevant to future AI usefulness.
From exam performance to expert tasks
Many widely discussed benchmarks resemble human examinations. The Massive Multitask Language Understanding (MMLU) benchmark tests knowledge and reasoning across dozens of academic and professional subjects. Other evaluations, such as graduate-level science question sets and advanced mathematics tests, aim to measure whether AI systems can handle increasingly difficult intellectual work. These benchmarks have shown clear progress: newer general-purpose models have reached much higher performance levels than earlier systems across many academic and technical domains.[arXiv]arxiv.orgInternational AI Safety Report 2025: First Key Update: Capabilities and Risk ImplicationsOctober 15, 2025…
Coding benchmarks provide another important signal because software development is closely connected to the possibility of AI-assisted research and self-improvement. Evaluations such as SWE-bench test whether AI systems can resolve real software issues drawn from actual code repositories rather than simply answer programming questions. These tests are more demanding because success requires understanding a codebase, making changes and checking whether those changes work.
However, even stronger benchmarks remain snapshots. A system that performs well on a defined test may still struggle when the task changes slightly, when information is incomplete or when the objective is unclear. The International AI Safety Report has highlighted that current systems can achieve impressive results on expert-level benchmarks while still showing uneven reliability in practical settings.[arXiv]arxiv.orgInternational AI Safety Report 2025: First Key Update: Capabilities and Risk ImplicationsOctober 15, 2025…
Measuring longer chains of work
A newer approach focuses less on isolated questions and more on whether AI agents can complete extended tasks. The AI evaluation organisation METR has developed “time horizon” measurements, which estimate the length of software and related tasks that AI agents can complete with a given probability of success. Instead of asking whether a model can answer a difficult question, this approach asks whether it can carry out a multi-step piece of work over a longer period.[Metr]metr.orgTask-Completion Time Horizons of Frontier AI ModelsTask-Completion Time Horizons of Frontier AI Models - METRMay 8, 2026…
This type of evaluation is particularly relevant to intelligence explosion arguments. A system that can reliably complete longer research or engineering tasks could contribute more meaningfully to scientific discovery and technological development. But even these measurements have boundaries: METR’s evaluations are concentrated mainly on software engineering, machine learning and cybersecurity tasks, rather than the full range of human intellectual activity.[Metr]metr.orgTask-Completion Time Horizons of Frontier AI ModelsTask-Completion Time Horizons of Frontier AI Models - METRMay 8, 2026…
Where benchmark gains fall short
The central problem is that intelligence is broader than test performance. Benchmarks capture useful pieces of capability, but they rarely measure the full combination of abilities needed for flexible, real-world intelligence.
High scores do not guarantee dependable behaviour
A benchmark usually provides a clean problem with a known answer. Real environments are different. Scientists, engineers, doctors and policymakers often work with incomplete information, changing objectives and unexpected obstacles. A system may produce a correct answer in a test setting while still making unreliable decisions when deployed.
This distinction is especially important for advanced AI systems because errors can become harder to detect as systems become more capable. A model that fails obviously on simple tasks is easy to evaluate. A system that performs well most of the time but occasionally makes subtle mistakes in complex situations creates a more difficult measurement challenge. The International AI Safety Report notes that reliability problems, including fabricated information and incorrect outputs, can persist even as general-purpose AI systems become more capable.[International AI Safety Report]internationalaisafetyreport.orginternational ai safety report 2025International AI Safety ReportInternational AI Safety Report 2025 | International AI Safety ReportJanuary 29, 2025…
Benchmark saturation can hide remaining weaknesses
As AI systems improve, some older benchmarks become less informative. When a model reaches very high scores on a test, that evaluation no longer reveals much about the remaining differences between systems. Researchers then create harder benchmarks, but those too can eventually become outdated.
This creates a moving target. A rising benchmark score can mean that AI is genuinely becoming more intelligent, but it can also reflect that researchers have become better at designing tests that expose existing strengths. The challenge is separating real advances in general capability from progress that is specific to a particular evaluation format.[vox.com]vox.comIt's getting harder to measure just how good AI is gettingIn 2024, it was observed that existing AI systems are powerful enough to significantly change our world. OpenAI's latest large language m…
Passing tests is different from understanding tasks
Another limitation is that benchmarks often measure whether an AI system can produce successful outputs, not whether it has human-like understanding. A model may solve a problem through statistical patterns, learned strategies or unexpected shortcuts rather than through reasoning that transfers broadly.
This does not mean benchmark results are meaningless. Human intelligence itself can be measured through tests, and strong performance on difficult tasks is evidence of important abilities. The issue is that no single benchmark captures intelligence as a whole. Progress requires combining many signals: task performance, robustness, autonomy, learning ability, scientific contribution and real-world reliability.
The search for better capability signals
As AI systems move closer to more general forms of problem-solving, researchers are developing evaluations designed to measure capabilities that traditional benchmarks miss.
From answers to actions
One major shift is towards agent evaluations. Instead of asking whether a model can produce the right response, these tests examine whether an AI system can plan, use tools, correct mistakes and complete objectives over multiple steps.
METR’s task-completion measurements represent this direction. Their work attempts to quantify how complex a task an AI agent can handle before reliability drops significantly. Such measures may be more informative for understanding whether AI could become a serious contributor to research and development, because scientific progress depends on completing long sequences of connected tasks rather than answering isolated questions.[Metr]metr.orgTask-Completion Time Horizons of Frontier AI ModelsTask-Completion Time Horizons of Frontier AI Models - METRMay 8, 2026…
Measuring reliability, not only ability
Future evaluations will likely need to place greater emphasis on consistency. A system that can occasionally solve extremely difficult problems but cannot reliably know when it is wrong may be less useful than a slightly less capable system that behaves predictably.
Important questions include:
- Can the system recognise uncertainty?
- Can it recover from mistakes?
- Can it adapt when conditions change?
- Can humans understand why it reached a conclusion?
- Does performance transfer beyond the environment where it was tested?
These questions are central to judging whether AI progress could support a broader intelligence expansion. For an AI-enabled human bloom, the important milestone is not simply higher scores, but dependable systems that can safely extend human scientific, economic and creative capabilities.
Why the measurement gap matters for the intelligence explosion debate
Benchmark results provide evidence that AI capabilities are increasing, but they do not by themselves prove that an intelligence explosion is underway. The strongest evidence for rapid progress comes from a combination of factors: improving benchmark performance, longer task completion, better scientific and engineering assistance, and falling costs for capable systems.[Metr]metr.orgTask-Completion Time Horizons of Frontier AI ModelsTask-Completion Time Horizons of Frontier AI Models - METRMay 8, 2026…
The unresolved question is whether these improvements will continue to compound into systems that can dramatically accelerate their own development and human innovation. If benchmarks are capturing only narrow improvements, expectations about future acceleration may be too optimistic. If they are revealing genuine growth in flexible problem-solving ability, they may be early signs of a much larger shift in humanity’s ability to create knowledge.
The measurement gap therefore sits at the centre of the uncertainty around advanced AI. Better benchmarks will not settle every debate about superintelligence, but they will help distinguish between two very different futures: one where AI becomes a collection of increasingly powerful tools, and another where it becomes a general-purpose engine for accelerating discovery, abundance and human flourishing.
Amazon book picks
Further Reading
Books and field guides related to What AI Tests Miss About Intelligence. Use these as the next step if you want deeper reading beyond the article.
The Alignment Problem: Machine Learning and Human Values
Finalist for the Los Angeles Times Book Prize A jaw-dropping exploration of everything that goes wrong when we build AI systems and the m...
Artificial Intelligence: A Guide for Thinking Humans
“After reading Mitchell’s guide, you’ll know what you don’t know and what other people don’t know, even though they claim to know it. And...
Rebooting AI: Building Artificial Intelligence We Can Trust
Two leaders in the field offer a compelling analysis of the current state of the art and reveal the steps we must take to achieve a robus...
Weapons of Math Destruction
'A manual for the 21st-century citizen... accessible, refreshingly critical, relevant and urgent' - Financial Times 'Fascinating and deep...
eBay marketplace picks
Marketplace Samples
Live-tested eBay searches with available results related to this page.
Selected fromcomputer science poster oneBay.co.uk.
Endnotes
1.
Source: arxiv.org
Link:https://arxiv.org/abs/2510.13653
Source snippet
International AI Safety Report 2025: First Key Update: Capabilities and Risk ImplicationsOctober 15, 2025...
Published: October 15, 2025
2.
Source: metr.org
Title: Task-Completion Time Horizons of [Frontier]({{ ‘compute-controls/’ | relative_url }}) AI Models
Link:https://metr.org/time-horizons/
Source snippet
Task-Completion Time Horizons of Frontier AI Models - METRMay 8, 2026...
Published: May 8, 2026
3.
Source: vox.com
Title: It’s getting harder to measure just how good AI is getting
Link:https://www.vox.com/future-perfect/394336/artificial-intelligence-openai-o3-benchmarks-agi
Source snippet
In 2024, it was observed that existing AI systems are powerful enough to significantly change our world. OpenAI's latest large language m...
4.
Source: metr.org
Title: Impact of modelling assumptions on time horizon results
Link:https://metr.org/notes/2026-03-20-impact-of-modelling-assumptions-on-time-horizon-results/
5.
Source: metr.org
Title: How Does Time Horizon Vary Across Domains?
Link:https://metr.org/blog/2025-07-14-how-does-time-horizon-vary-across-domains/
6.
Source: metr.org
Link:https://metr.org/index.html
7.
Source: youtube.com
Title: AI Safety Benchmarks Do Not Benchmark Safety
Link:https://www.youtube.com/watch?v=HaKi5uwX6p0
Source snippet
International AI Safety Report: First Key Update, 2025...
8.
Source: youtube.com
Link:https://www.youtube.com/watch?v=JxgMMwspdHs
9.
Source: internationalaisafetyreport.org
Title: international ai safety report 2025
Link:https://internationalaisafetyreport.org/publication/international-ai-safety-report-2025
Source snippet
International AI Safety ReportInternational AI Safety Report 2025 | International AI Safety ReportJanuary 29, 2025...
Published: January 29, 2025
10.
Source: OpenAI
Title: separating signal from noise coding evaluations
Link:https://openai.com/index/separating-signal-from-noise-coding-evaluations/
Source snippet
comSeparating signal from noise in coding evaluations | OpenAIJuly 8, 2026 — OpenAI July 8, 2026 ResearchPublication SEPARATING SIGNAL FR...
Published: July 8, 2026
11.
Source: OpenAI
Title: why we no longer evaluate swe bench verified
Link:https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/
12.
Source: internationalaisafetyreport.org
Title: Publications | International AI Safety Report
Link:https://internationalaisafetyreport.org/publications
13.
Source: internationalaisafetyreport.org
Title: International AI Safety Report
Link:https://internationalaisafetyreport.org/
14.
Source: internationalaisafetyreport.org
Title: international ai safety report 2026
Link:https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026
15.
Source: internationalaisafetyreport.org
Link:https://internationalaisafetyreport.org/publication/second-key-update-technical-safeguards-and-risk-management
16.
Source: internationalaisafetyreport.org
Link:https://internationalaisafetyreport.org/publication/first-key-update-capabilities-and-risk-implications
17.
Source: GOV.UK
Title: international ai safety report 2025
Link:https://www.gov.uk/government/publications/international-ai-safety-report-2025/international-ai-safety-report-2025
18.
Source: internationalaisafetyreport.org
Link:https://internationalaisafetyreport.org/about
19.
Source: rai.ac.uk
Title: international ai safety report january 2025
Link:https://rai.ac.uk/reports/international-ai-safety-report-january-2025/
Published: january 2025
Additional References
20.
Source: research.mental-momentum.ai
Title: Key takeaways * Current AI benchmarks primarily measure crystallized intelligen
Link:https://research.mental-momentum.ai/r/capabilities-limitations-ai-benchmarks-kl0kpk
Source snippet
and Limitations of AI BenchmarksJune 14, 2026 — Updated 2026-06-14 What do AI benchmarks like MMLU and ARC actually measure — and what th...
Published: June 14, 2026
21.
Source: youtube.com
Title: Why AI Agent Benchmarks Are Breaking — and How to Evaluate What Matters
Link:https://www.youtube.com/watch?v=9Xd-P19BdQ8
Source snippet
How We Test AI: Benchmark Datasets Explained (MMLU, GSM8K & More)...
22.
Source: youtube.com
Link:https://www.youtube.com/watch?v=mrHxGC0eW_Q
Source snippet
AI Safety Benchmarks Do Not Benchmark Safety - Orestis Papakyriakopoulos...
23.
Source: youtube.com
Title: How We Test AI: Benchmark Datasets Explained (MMLU, GSM8K & More)
Link:https://www.youtube.com/watch?v=7t1RdmiW3fc
Source snippet
State of AI Benchmarks 2026: What the Tests Actually Tell Us...
24.
Source: report-ai.org
Title: Technical Per
Link:https://report-ai.org/indexes/technical-benchmarks/ai-models-benchmarks-statistics-2026/
Source snippet
AI Model Benchmarks 2026: SWE-bench, MMLU & Frontier Models - The AI IndexJune 5, 2026 — AI MODEL BENCHMARKS 2026: SWE-BENCH, MMLU & FRON...
Published: June 5, 2026
25.
Source: evals.[alignment]({{ ‘alignment/’ | relative_url }}). org
Title: * April
Link:https://evals.alignment.org/time-horizons/
Source snippet
alignment.orgTask-Completion Time Horizons of Frontier AI Models - METRMay 8, 2026 — UPDATES * May 8th, 2026: Added Claude Mythos Preview...
Published: May 8, 2026
26.
Source: ai-insight.org
Link:https://www.ai-insight.org/reports/benchmark-methodology-2026
27.
Source: doi.org
Link:https://doi.org/10.1007/s10462-026-11571-0
28.
Source: waggertron.github.io
Link:https://waggertron.github.io/tech-learning/topics/ai/benchmarks/
29.
Source: metavert.io
Link:https://metavert.io/metr-benchmarking



