Within Safety Frameworks

Can warning signs catch dangerous AI early?

Google DeepMind's Critical Capability Levels show how labs are trying to detect dangerous abilities before they are obvious in deployed systems.

On this page

  • What Critical Capability Levels are meant to measure
  • Why buffers matter before capabilities fully emerge
  • Where voluntary warning systems may fall short
Preview for Can warning signs catch dangerous AI early?

Introduction

One of the most concrete mechanisms emerging in frontier AI safety is the use of Critical Capability Levels (CCLs) by[Google DeepMind]deepmind.googleGoogle Deep Mind Introducing the Frontier Safety Framework — Google Deep MindGoogle DeepMindIntroducing the Frontier Safety Framework — Google DeepMindMay 17, 2024…Published: May 17, 2024 labs like Google DeepMind to act as early‑warning safety gates for increasingly capable systems. CCLs are predefined thresholds of capability that, according to DeepMind’s Frontier Safety Framework (FSF), mark when a model could — without strong safeguards in place — pose a heightened risk of severe harm. These thresholds serve as internal tripwires: models are evaluated periodically against them, and if an evaluation signals that a model is nearing a CCL, stronger security controls, deployment restrictions, or development pauses are triggered. This approach aims to provide a structured, anticipatory way for developers to “catch” dangerous capabilities early, long before they are publicly deployed or integrated into critical infrastructure.[Google DeepMind]deepmind.googleGoogle Deep Mind Introducing the Frontier Safety Framework — Google Deep MindGoogle DeepMindIntroducing the Frontier Safety Framework — Google DeepMindMay 17, 2024…Published: May 17, 2024

Deep Mind CCLs illustration 1 In the context of Frontier AI safety frameworks for autonomy and control, DeepMind’s CCLs represent an attempt to connect the technical progress of AI capabilities with risk governance actions — creating measurable, graded thresholds that guide decision‑making during development. From the perspective of the AI Bloom frame, these early‑warning gates are a response to the challenge of harnessing potentially world‑transforming intelligence while retaining human control and oversight, ensuring that benefits do not outpace society’s ability to govern emerging risks.

What Critical Capability Levels Are and How They Work

At its core, a Critical Capability Level (CCL) is a threshold of capability identified by frontier AI developers as indicative of a class of risks that warrant strong safety measures if approached or crossed. DeepMind’s FSF defines CCLs by analysing “paths through which a model could cause severe harm” in certain high‑risk domains and then identifying the minimal set of capabilities a model would need to traverse those paths.[Google Cloud Storage]storage.googleapis.comGoogle Cloud Storage Frontier Safety Framework YesterdayGoogle Cloud Storage Frontier Safety Framework Yesterday

Two defining features of CCLs as early‑warning safety gates are:

  • Domain‑specific thresholds: CCLs are not generic “power scores” but are scoped to specific domains of risk — such as autonomy, biosecurity, cyberattacks, machine‑learning research acceleration, harmful manipulation or misalignment with operator intentions. Each domain has its own CCLs that map onto the kinds of capability developments that matter for that domain’s risks.[Google Cloud Storage]storage.googleapis.comGoogle Cloud Storage Frontier Safety Framework YesterdayGoogle Cloud Storage Frontier Safety Framework Yesterday
  • Early warning evaluations: DeepMind’s framework calls for periodic evaluations of frontier models against these CCLs using suites of tests designed to signal whether a model is approaching a threshold. These are referred to as “early warning evaluations” and are intended to give teams notice before a full‑blown CCL is reached. The frequency and timing of these evaluations — for example tied to increases in compute or changes in fine‑tuning — are structured so that there is a buffer zone between detection and threshold crossing.[AI Safety Atlas]ai-safety-atlas.comAI Safety Atlas Evaluation FrameworksAI Safety AtlasEvaluation Frameworks - Chapter 5 - AI Safety Atlas…

This design aims to prevent “end‑of‑pipeline” surprises by embedding checks into the development lifecycle itself. In effect, CCLs act like guardrails: if a model starts to look as though it could meaningfully implement certain autonomous behaviours, generate effective cyber‑attack code, or substantially accelerate AI research without controls, the evaluation flags this in time to adjust the development plan or impose mitigations.

Why Buffers Matter Before Capabilities Fully Emerge

The core idea of using CCLs as early warnings relies on safety buffers — zones of monitoring and mitigation that kick in before a model actually becomes capable of harm. DeepMind’s own implementation ties evaluations to measurable progress indicators such as effective compute and training milestones, so that CCL approaches can be detected early and acted upon.[AI Safety Atlas]ai-safety-atlas.comAI Safety Atlas Evaluation FrameworksAI Safety AtlasEvaluation Frameworks - Chapter 5 - AI Safety Atlas…

The importance of these buffers can be understood in three ways:

1. Proportional Mitigation: If models are only judged after they clearly exhibit dangerous capabilities, it becomes much harder to mitigate those risks without rolling back development or introducing disruptive controls. Early warnings buy time to apply escalating security and deployment restrictions in proportion to the risk.

2. Continuous Risk Insight: Rather than a single pass/fail evaluation at the end of development, periodic evaluations create a trajectory of capability growth. This means teams can see how close a model is to a given CCL and adjust resources and safeguards accordingly.

3. Guarding Against Emergence: A key concern in frontier AI safety is that complex behaviours can appear unpredictably as models scale (so‑called “emergent capabilities”). Early warning evaluations are one of the few practical ways labs have today to detect emergent paths to problematic capabilities before they fully materialise.

In these ways, CCL‑based gating recognises that we may not be able to predict every harmful outcome in advance, but we can monitor specific capability vectors and impose controls before conditions worsen.

Deep Mind CCLs illustration 2

Where Voluntary Warning Systems May Fall Short

While DeepMind’s CCLs represent a more systematic approach than ad‑hoc safety judgements, they also have limitations worth understanding in the broader context of frontier AI governance.

Uncertainty in Capability‑to‑Risk Mapping: Defining what minimal combination of capabilities constitutes a real risk is inherently speculative. Critics point out that there may be multiple different capability configurations that could result in similar harms, some of which might not be captured by the predefined CCL tests. DeepMind’s framework itself admits that CCLs evolve as understanding improves. A narrow focus on certain domains could miss other harm vectors not yet well modelled.[Machine Intelligence Research Institute]intelligence.orgMachine Intelligence Research InstituteExisting Safety Frameworks Imply Unreasonable Confidence - Machine Intelligence Research Institute…

Predictive Limits: Early warning evaluations assume that models exhibiting partial or near‑threshold behaviours are a reliable indicator of future risk. But because modern AI systems can surprise researchers with unexpected combinations of skills, CCLs may under‑ or over‑estimate real risk in some cases.

Voluntary Adoption and Comparability: As a voluntary, lab‑internal policy, CCLs lack external enforcement and standardisation. Different organisations may define thresholds differently or evaluate at divergent intervals, making collective risk governance harder. Although DeepMind has shared its framework publicly in hopes of standardising practice, adoption outside a few labs remains uneven.[AI Security & Safety Directory]aisecurityandsafety.orggoogle deepmind frontier safety frameworkAI Security & Safety DirectoryGoogle DeepMind Frontier Safety Framework (International, 2026): | AI Safety DirectoryMarch 10, 2026…Published: March 10, 2026

Focus on Capability Rather Than Motivation or Deployment Context: CCLs primarily evaluate what a model could do, not how it will be used or how intentions behind deployment could shape outcomes. This means that even models below certain CCLs could be misused in harmful ways if combined with human intent or operational settings not covered by the classifications.

Despite these caveats, CCLs are significant because they formalise for the first time how a leading frontier lab tries to translate technical capability growth into structured safety decisions, tying model evaluations directly to governance actions rather than relying solely on judgement calls at the end of development.

CCLs and the Broader AI Bloom Question

From the perspective of humanity’s long‑term future, early warning safety gates like CCLs are a response to the tension between transformative potential and systemic risk. Advanced AI holds enormous promise — accelerating science, solving complex global problems and enhancing human capabilities — but that promise is tied to systems whose capabilities could outstrip our ability to control them.

CCLs attempt to bridge that gap by embedding anticipatory checks and mitigations into the development lifecycle of frontier models. They do not guarantee safety, nor are they a complete solution to alignment challenges, but they represent an operational attempt to keep powerful capabilities on a trajectory where benefits can be realised while risks remain manageable. In doing so, they align with a central question of the AI Bloom project: how to reap the transformative benefits of AI while preserving human direction, agency and governance over powerful technologies.

Deep Mind CCLs illustration 3

Amazon book picks

Further Reading

Books and field guides related to Can warning signs catch dangerous AI early?. Use these as the next step if you want deeper reading beyond the article.

eBay marketplace picks

Marketplace Samples

Example marketplace items related to this page. Use the search link to explore similar finds on eBay.

UsingUSA

Endnotes

1. Source: deepmind.google
Title: Google Deep Mind Introducing the Frontier Safety Framework — Google Deep Mind
Link:https://deepmind.google/discover/blog/introducing-the-frontier-safety-framework/

Source snippet

Google DeepMindIntroducing the Frontier Safety Framework — Google DeepMindMay 17, 2024...

Published: May 17, 2024

2. Source: ai-safety-atlas.com
Title: AI Safety Atlas Evaluation Frameworks
Link:https://ai-safety-atlas.com/chapters/v1/evaluations/evaluation-frameworks/

Source snippet

AI Safety AtlasEvaluation Frameworks - Chapter 5 - AI Safety Atlas...

3. Source: intelligence.org
Link:https://intelligence.org/2025/04/09/existing-safety-frameworks-imply-unreasonable-confidence/

Source snippet

Machine Intelligence Research InstituteExisting Safety Frameworks Imply Unreasonable Confidence - Machine Intelligence Research Institute...

4. Source: deepmind.google
Title: Google Deep Mind strengthens the Frontier Safety Framework — Google Deep Mind
Link:https://deepmind.google/discover/blog/strengthening-our-frontier-safety-framework/

Source snippet

Google DeepMind strengthens the Frontier Safety Framework — Google DeepMindSeptember 22, 2025 — September 22, 2025 Responsibility & Safet...

Published: September 22, 2025

5. Source: deepmind.google
Title: Updating the Frontier Safety Framework — Google Deep Mind
Link:https://deepmind.google/discover/blog/updating-the-frontier-safety-framework/

Source snippet

Updating the Frontier Safety Framework — Google DeepMindFebruary 4, 2025 — February 4, 2025 Responsibility & Safety UPDATING THE FRONTIER...

Published: February 4, 2025

6. Source: deepmind.google
Title: evaluating frontier models for dangerous capabilities
Link:https://deepmind.google/research/publications/evaluating-frontier-models-for-dangerous-capabilities/

Source snippet

Google DeepMindMarch 21, 2024 — March 21, 2024 EVALUATING FRONTIER MODELS FOR DANGEROUS CAPABILITIES View publication Download ABSTRACT T...

Published: March 21, 2024

7. Source: storage.googleapis.com
Title: Google Cloud Storage Frontier Safety Framework Yesterday
Link:https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/introducing-the-frontier-safety-framework/fsf-technical-report.pdf

8. Source: storage.googleapis.com
Title: frontier safety framework 3 1
Link:https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3-1.pdf

Source snippet

Google Cloud StorageFrontier Safety Framework 3.1May 16, 2026...

Published: May 16, 2026

9. Source: aisecurityandsafety.org
Title: google deepmind frontier safety framework
Link:https://aisecurityandsafety.org/en/frameworks/google-deepmind-frontier-safety-framework/

Source snippet

AI Security & Safety DirectoryGoogle DeepMind Frontier Safety Framework (International, 2026): | AI Safety DirectoryMarch 10, 2026...

Published: March 10, 2026

10. Source: aisecurityandsafety.org
Title: google deepmind frontier safety framework
Link:https://aisecurityandsafety.org/frameworks/google-deepmind-frontier-safety-framework/

Source snippet

(International, 2026): | AI Safety DirectoryMarch 10, 2026 — GOOGLE DEEPMIND FRONTIER SAFETY FRAMEWORK best practice active Last updated...

Published: March 10, 2026

Additional References

11. Source: metr.org
Link:https://metr.org/common-elements

Source snippet

Common Elements of Frontier AI Safety Policies - METRDecember 16, 2025 — CAPABILITY THRESHOLDS Descriptions of AI capability levels which...

Published: December 16, 2025

12. Source: arstechnica.com
Title: Deep Mind AI safety report explores the perils of “misaligned” AI
Link:https://arstechnica.com/google/2025/09/deepmind-ai-safety-report-explores-the-perils-of-misaligned-ai/

Source snippet

DEEPMIND AI SAFETY REPORT EXPLORES THE PERILS OF “MISALIGNED” AI DeepMind releases version 3.0 of its AI Frontier Safety Framework with n...

13. Source: scribd.com
Link:https://www.scribd.com/document/868635487/Fsf-Technical-Report

Source snippet

CRITICAL CAPABILITY LEVELS: The Framework is built around capability thresholds called “Critical Capability Levels.” These are capability...

14. Source: lesswrong.com
Title: Deep Mind: Frontier Safety Framework — Less Wrong
Link:https://www.lesswrong.com/posts/AFQt6uByLYNrNgyBb/deepmind-frontier-safety-framework

Source snippet

DeepMind: Frontier Safety Framework — LessWrongMay 17, 2024 — DeepMind: Frontier Safety Framework 3 min read • Excerpt AIFrontpage 64 DEE...

Published: May 17, 2024

15. Source: greaterwrong.com
Title: Deep Mind: Frontier Safety Framework
Link:https://www.greaterwrong.com/posts/AFQt6uByLYNrNgyBb/deepmind-frontier-safety-framework

Source snippet

DeepMind: Frontier Safety Framework - LessWrong 2.0 viewerMay 17, 2024 — EXCERPT > Today, we are introducing our Frontier Safety Framewor...

Published: May 17, 2024

16. Source: youtube.com
Link:https://www.youtube.com/watch?v=nbnTvNEumZI

Source snippet

It Begins: AI Models Have Started Forming Alliances Against Us...

17. Source: youtube.com
Title: A Crash Course on AI Standards with Google Deep Mind’s Owen Larter
Link:https://www.youtube.com/watch?v=ToYhh_jU9n0

Source snippet

Tutorial on AI Alignment (part 1 of 2): Safety Vulnerabilities of Current Frontier Models...

18. Source: youtube.com
Title: It Begins: AI Models Have Started Forming Alliances Against Us
Link:https://www.youtube.com/watch?v=8tu4Ws8O5_4

Source snippet

DeepMind frontier safety | Mary Phuong | EAG London: 2024...

19. Source: youtube.com
Title: The Risks That Really Worry Deep Mind — And How They Test
Link:https://www.youtube.com/watch?v=6e_LgAu_QIw

Source snippet

A Crash Course on AI Standards with Google DeepMind's Owen Larter...

20. Source: agora.eto.tech
Link:https://agora.eto.tech/instrument/2040

Source snippet

DeepMind Frontier Safety Framework Version 2.0 – ETO AGORAFebruary 4, 2025 — GOOGLE DEEPMIND FRONTIER SAFETY FRAMEWORK VERSION 2.0 Propos...

Published: February 4, 2025

Topic Tree

Follow this branch

Parent topic

Safety Frameworks How Labs Decide When AI Becomes Too Risky

Related pages 2