Within Alignment Risks
How AI Safety Tests Find Hidden Risks
Safety evaluations aim to reveal dangerous capabilities and failures before increasingly capable AI systems are deployed widely.
On this page
- Beyond traditional performance benchmarks
- Testing autonomy and dangerous capabilities
- Limits of evaluating future systems
Page outline Jump by section
Introduction
AI safety evaluations are designed to answer a practical question before powerful systems are widely deployed: what can this model actually do, and what could go wrong if those abilities are placed in real-world settings? Traditional AI benchmarks measure skills such as language understanding, coding ability or reasoning accuracy, but advanced systems require additional tests for dangerous capabilities, unreliable behaviour and unexpected autonomy.[GOV.UK]GOV.UKinternational ai safety report 2025Withdrawn] International AI Safety Report 2025 - GOV.UKFebruary 18, 2025…
These evaluations matter because an AI-enabled human bloom depends not only on increasing capability but on ensuring that greater intelligence remains compatible with human goals. A system that accelerates medical research, scientific discovery or education could expand human flourishing, but only if society can identify and manage serious failure modes before those systems become deeply embedded in critical infrastructure and decision-making. Safety evaluations are therefore a bridge between optimism about advanced AI and the practical work needed to make that future safer.[GOV.UK]GOV.UKinternational ai safety report 2025Withdrawn] International AI Safety Report 2025 - GOV.UKFebruary 18, 2025…
Beyond traditional performance benchmarks
For many years, AI progress was measured mainly by capability benchmarks: can a model answer questions, translate text, write code or recognise patterns? These tests remain useful, but they do not necessarily reveal whether a model can behave dangerously when given more freedom, tools or real-world objectives.
Safety evaluations ask different questions. Can a system deceive evaluators? Can it pursue a goal in an unintended way? Can it misuse access to software tools, sensitive information or external resources? Can it operate effectively enough that human oversight becomes difficult? The focus shifts from “how intelligent is the model?” to “how does that intelligence behave under realistic conditions?”[arXiv]arxiv.orgarXiv Evaluating Frontier Models for Dangerous CapabilitiesarXiv Evaluating Frontier Models for Dangerous Capabilities
One important development has been the rise of dangerous capability evaluations. These tests attempt to measure whether frontier models show abilities relevant to serious risks, including cyber operations, persuasion, deception, self-reasoning and forms of autonomous action. A research programme evaluating Gemini models developed tests across areas such as cybersecurity, persuasion, self-proliferation and reasoning about their own capabilities. The researchers did not find strong evidence of extreme dangerous abilities in the models tested, but argued that evaluation methods needed to improve as systems became more capable.[arXiv]arxiv.orgarXiv Evaluating Frontier Models for Dangerous CapabilitiesarXiv Evaluating Frontier Models for Dangerous Capabilities
This reflects a wider challenge: safety testing is not simply about finding existing failures. It is about creating warning systems for capabilities that may become important only when models cross new thresholds. A model that performs poorly on a narrow laboratory task today may become a significant concern after improvements in memory, tool use, autonomy or access to external systems.
Testing autonomy and dangerous capabilities
Red-teaming models before deployment
A central method in AI safety evaluation is red-teaming: deliberately trying to make a system fail. Instead of asking only whether a model behaves well in ordinary use, researchers search for situations where it produces harmful, misleading or unexpected outcomes.
Red teams may test whether a model can be persuaded to bypass safeguards, whether it follows harmful instructions, whether it gives unreliable advice in high-stakes settings, or whether it behaves differently when placed under pressure. The goal is not to prove that a model is unsafe, but to discover weaknesses while there is still time to address them.[arXiv]arxiv.orgarXiv Holistic Safety and Responsibility Evaluations of Advanced AI ModelsarXiv Holistic Safety and Responsibility Evaluations of Advanced AI Models
Major AI developers increasingly publish system cards describing these evaluations. These documents typically combine capability tests, safety assessments and deployment decisions, allowing researchers and the public to understand why a model was released with particular safeguards. Anthropic’s system cards, for example, document capability evaluations and safety assessments for Claude models as part of its responsible deployment process.[anthropic.com]anthropic.comModel system cards \ AnthropicModel system cards \ Anthropic
Measuring autonomous behaviour
A major shift in safety evaluation is the move from testing isolated answers to testing agents: systems that can plan, use tools, complete tasks over time and interact with digital environments.
This creates new risks because a model may not need to be generally superhuman to cause problems. A system that can write and execute code, search information, manage accounts or coordinate actions may create risks through persistence and access rather than through a single exceptionally intelligent response.
The Model Evaluation & Threat Research organisation (METR) has developed evaluations that measure how long and how independently frontier AI systems can complete complex tasks. Its “time horizon” approach measures the length of tasks AI agents can successfully perform, providing one way to track increasing autonomy rather than only raw benchmark scores.[Evals]evals.alignment.orgEvals METREvals METR
Autonomy evaluations are especially relevant to the long-term AI bloom vision because many optimistic scenarios depend on AI systems carrying out extended scientific, engineering and organisational work. The same capability that could accelerate drug discovery or climate technology could create greater risks if systems pursue poorly specified goals or operate without adequate supervision.
Testing for misuse and security risks
Safety evaluations also examine how advanced models could change the ability of people to cause harm. This includes questions such as whether AI meaningfully lowers the expertise required for cyberattacks, biological misuse or large-scale manipulation.
Researchers increasingly distinguish between a model’s internal capability and its impact when combined with human users, tools and access. A moderately capable model with widespread availability may create different risks from a highly capable but tightly controlled system. Some researchers have argued that evaluations should measure “harmful capability uplift” — the additional ability a model gives users to carry out harmful actions compared with existing tools.[arXiv]arxiv.orgEvaluating Human-AI Safety: A Framework for Measuring Harmful Capability UpliftMarch 6, 2026…
This is important for governance because safety cannot be judged only by asking whether a model itself has dangerous intentions. The wider question is whether deployment changes what individuals, organisations or governments can do.
Limits of evaluating future systems
Safety evaluations are becoming more sophisticated, but they remain an imperfect measurement tool. The biggest difficulty is that tests are always based on current understanding of risk. Future AI systems may develop capabilities, strategies or forms of autonomy that existing evaluations were not designed to detect.
A second challenge is that models can behave differently in controlled tests and real environments. Laboratory evaluations often involve carefully designed prompts, limited access and short time periods. Real deployment may involve millions of users, unexpected incentives, complex social situations and integration with other systems.
There is also a risk of “benchmark gaming”: improving performance on known tests without genuinely improving safety. As evaluations become widely used, developers and researchers must ensure that tests measure meaningful behaviour rather than becoming targets that systems are simply optimised to pass.
International reviews of AI safety research have highlighted evaluation as a critical part of risk management, while also noting the difficulty of assessing rapidly advancing general-purpose systems. The International AI Safety Report emphasised the need for stronger methods to identify and assess risks from advanced AI, including risks from misuse, malfunctions and broader systemic effects.[GOV.UK]GOV.UKinternational ai safety report 2025Withdrawn] International AI Safety Report 2025 - GOV.UKFebruary 18, 2025…
Why better evaluations matter for an AI-enabled future
Safety evaluations do not attempt to predict the entire future of AI. Their purpose is more practical: to create evidence that helps societies decide when and how powerful systems can be trusted.
For an AI bloom scenario, this matters because many of the largest potential benefits involve giving AI systems greater responsibility. Scientific acceleration may require AI researchers that can run experiments and propose new ideas. Healthcare advances may require systems that assist with complex medical decisions. Education and productivity gains may depend on AI agents that act with increasing independence.
The same expansion of capability creates the need for stronger evidence about reliability, control and alignment. Evaluations are therefore not a barrier to technological progress; they are part of the infrastructure that could make ambitious progress safer. A future where AI contributes to longer healthy lives, abundant knowledge and broader human opportunity depends not only on building more capable systems, but on learning how to test them before their power exceeds our ability to manage their failures.[GOV.UK]GOV.UKinternational ai safety report 2025Withdrawn] International AI Safety Report 2025 - GOV.UKFebruary 18, 2025…
Amazon book picks
Further Reading
Books and field guides related to How AI Safety Tests Find Hidden Risks. Use these as the next step if you want deeper reading beyond the article.
The Alignment Problem: Machine Learning and Human Values
Finalist for the Los Angeles Times Book Prize A jaw-dropping exploration of everything that goes wrong when we build AI systems and the m...
Human Compatible: Artificial Intelligence and the Problem of...
A leading artificial intelligence researcher lays out a new approach to AI that will enable us to coexist successfully with increasingly...
Artificial Intelligence: A Guide for Thinking Humans
“After reading Mitchell’s guide, you’ll know what you don’t know and what other people don’t know, even though they claim to know it. And...
Rebooting AI: Building Artificial Intelligence We Can Trust
Two leaders in the field offer a compelling analysis of the current state of the art and reveal the steps we must take to achieve a robus...
eBay marketplace picks
Marketplace Samples
Live-tested eBay searches with available results related to this page.
Selected fromrobotics kit oneBay.co.uk.
Endnotes
1.
Source: GOV.UK
Title: international ai safety report 2025
Link:https://www.gov.uk/government/publications/international-ai-safety-report-2025/international-ai-safety-report-2025
Source snippet
[Withdrawn] International AI Safety Report 2025 - GOV.UKFebruary 18, 2025...
Published: February 18, 2025
2.
Source: arxiv.org
Title: arXiv Evaluating Frontier Models for Dangerous Capabilities
Link:https://arxiv.org/abs/2403.13793
3.
Source: arxiv.org
Title: arXiv Holistic Safety and Responsibility Evaluations of Advanced AI Models
Link:https://arxiv.org/abs/2404.14068
4.
Source: anthropic.com
Title: Model system cards \ Anthropic
Link:https://www.anthropic.com/system-cards
5.
Source: evals.alignment.org
Title: Evals METR
Link:https://evals.alignment.org/
6.
Source: metr.org
Link:https://metr.org/index.html
7.
Source: arxiv.org
Link:https://arxiv.org/abs/2603.26676
Source snippet
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability UpliftMarch 6, 2026...
Published: March 6, 2026
8.
Source: GOV.UK
Title: International scientific report on the safety of advanced AI: interim report
Link:https://www.gov.uk/government/publications/international-scientific-report-on-the-safety-of-advanced-ai/international-scientific-report-on-the-safety-of-advanced-ai-interim-report
9.
Source: industry.gov.au
Title: These are AI systems that can do many different
Link:https://www.industry.gov.au/publications/international-ai-safety-report-2025
Source snippet
International AI safety report 2025 | Department of Industry Science and ResourcesJune 1, 2026 — PUBLISHER * AI Safety Institute INTRODUC...
Published: June 1, 2026
10.
Source: anthropic.com
Title: ’s Transparency Hub \ Anthropic
Link:https://www.anthropic.com/transparency/model-report
11.
Source: anthropic.com
Title: ’s Transparency Hub
Link:https://www.anthropic.com/transparency?destination=%2Fleadership-organisations%2Fpower-not-knowing%3F_ref%3Dfinder%26aad%3Dbahjij57inr5cguioijjb3vyc2uilcj1cmwioijodhrwoi8vaw5zzwfklmvkdsisimlkijo2mdy3oduxm30gogzfva–4fcfb8808f6cfb242d04e6c7d8c30ce481285d72_reffinder%26adid%3Dmba-landingpage%26campaignid%3Dggl_search_apac%26campaignname%3Dapac-sgen_ggl-brandgen-dp-mba_mt-phrase%26device%3Dc%26f0%3Dskill_s7571%26f1%3Dskill_s7581%26field_format_value%3D2%26field_programme_name%3Dmba%26field_programme_type0%3D161%26field_programme_type1%3D161%26field_sch_country_1%3Dall%26field_sch_region%3Dall%26field_scholarship_type%3Dinsead_scholarship%26gad_source%3D1%26programme_code%3Dmba%26session%3Dflex%26siteid%3Ddp_ggl%26term%3Dinseadmba_p&e45d281a_page=1
12.
Source: anthropic.com
Title: ’s Transparency Hub \ Anthropic
Link:https://www.anthropic.com/transparency?e45d281a_page=5
13.
Source: anthropic.com
Link:https://www.anthropic.com/research/bloom?mibextid=UyTHkb
14.
Source: anthropic.com
Title: Cyber evaluations of Claude 4 \ Anthropic
Link:https://www.anthropic.com/research/claude-4-cyber
15.
Source: evaluations.metr.org
Title: claude 3 7 report
Link:https://evaluations.metr.org/claude-3-7-report/
16.
Source: evals.alignment.org
Title: claude 3 7 report
Link:https://evals.alignment.org/evaluations/claude-3-7-report/
17.
Source: GOV.UK
Link:https://www.gov.uk/government/publications/international-scientific-report-on-the-safety-of-advanced-ai
18.
Source: evaluations.metr.org
Link:https://evaluations.metr.org/
19.
Source: evals.alignment.org
Title: 2023 03 18 update on recent evals
Link:https://evals.alignment.org/blog/2023-03-18-update-on-recent-evals/
20.
Source: evals.alignment.org
Title: risk assessment
Link:https://evals.alignment.org/risk-assessment/
21.
Source: evals.alignment.org
Link:https://evals.alignment.org/about
22.
Source: evals.alignment.org
Link:https://evals.alignment.org/research/
23.
Source: evals.alignment.org
Title: measuring autonomous ai capabilities
Link:https://evals.alignment.org/measuring-autonomous-ai-capabilities/
24.
Source: aisecurityandsafety.org
Title: Evals (AI Evaluations) — [AI Governance]({{ ‘ai-governance/’ | relative_url }}) Definition & Guide | AI Safety Directory
Link:https://aisecurityandsafety.org/en/glossary/evals/
25.
Source: internationalaisafetyreport.org
Title: Publications | International AI Safety Report
Link:https://internationalaisafetyreport.org/publications
26.
Source: internationalaisafetyreport.org
Title: International AI Safety Report
Link:https://internationalaisafetyreport.org/
27.
Source: internationalaisafetyreport.org
Title: international ai safety report 2025
Link:https://internationalaisafetyreport.org/publication/international-ai-safety-report-2025
28.
Source: internationalaisafetyreport.org
Link:https://internationalaisafetyreport.org/about
Additional References
29.
Source: aisecurityandsafety.org
Title: As models grow more capa
Link:https://aisecurityandsafety.org/it/guides/ai-model-evaluation/
Source snippet
AI Model Evaluation: Safety Benchmarks, Red Teaming & Testing (2026) | AI Safety DirectoryMarch 27, 2026 — AI MODEL EVALUATION: SAFETY BE...
Published: March 27, 2026
30.
Source: youtube.com
Title: Cheating Behaviour in Frontier Model Evaluations
Link:https://www.youtube.com/watch?v=Bq6V6ltTTPk
Source snippet
Google DeepMind CEO: We're Not Ready For The Scariest Part That Is Yet To Come...
31.
Source: youtube.com
Link:https://www.youtube.com/watch?v=ZTmRT2Hg1oM
Source snippet
Cheating Behaviour in Frontier Model Evaluations...
32.
Source: internationalaisafetyreport.org
Link:https://internationalaisafetyreport.org/publication/second-key-update-technical-safeguards-and-risk-management
33.
Source: deepmind.google
Link:https://deepmind.google/research/publications/78149/
34.
Source: youtube.com
Link:https://www.youtube.com/watch?v=M5Ho6AA7rSw
Source snippet
Evaluating Malicious AI capabilities...
35.
Source: internationalaisafetyreport.org
Link:https://internationalaisafetyreport.org/publication/first-key-update-capabilities-and-risk-implications
36.
Source: youtube.com
Title: Evaluating Malicious AI capabilities
Link:https://www.youtube.com/watch?v=2veiKzqTPRw
Source snippet
DeepMind frontier safety | Mary Phuong | EAG London: 2024...
37.
Source: daringfireball.net
Title: Daring Fireball: Anthropic’s ‘System Card’ for Claude 4 (Opus and Sonnet)
Link:https://daringfireball.net/linked/2025/05/23/anthropic-claude-4-system-card
38.
Source: youtube.com
Title: Google Deep Mind CEO: We’re Not Ready For The Scariest Part That Is Yet To Come
Link:https://www.youtube.com/watch?v=SW3HzRV2Ztg



