Within Control

How Labs Decide When AI Becomes Too Risky

Major AI labs are building safety frameworks to track dangerous capability thresholds before systems become harder to manage.

On this page

  • Critical capability levels and frontier models
  • Autonomy, cyber risk and biosecurity thresholds
  • Limits of current oversight and evaluation
Preview for How Labs Decide When AI Becomes Too Risky

Introduction

The most important frontier AI safety question is no longer whether advanced systems can produce harmful outputs. It is whether future systems could become capable enough, autonomous enough, or strategically useful enough that ordinary testing and product moderation stop being adequate forms of control.

Safety Frameworks illustration 1 That concern has pushed major AI labs to create formal safety frameworks that attempt to identify dangerous capability thresholds before models become harder to supervise. These frameworks are not primarily about chatbot mistakes or misinformation. They focus on risks associated with increasingly powerful systems: autonomous operation, advanced cyber capabilities, biological knowledge that could aid dangerous actors, and AI systems that help create even more capable AI. The underlying idea is simple: if certain capabilities appear, developers should not rely on ad hoc judgement. They should have predefined rules that trigger stronger safeguards, deployment limits, security controls or, in some cases, pauses in deployment.[Google DeepMind]youtube.comGoogle DeepMind…[2cdn.openai.com]cdn.openai.compreparedness framework v2Preparedness Framework15 Apr 2025 — SAG Membership: the SAG provides a diversity of perspectives to evaluate the strength of evidence rel…

For people thinking about AI bloom and humanity’s long-term future, these frameworks occupy a crucial middle ground. They are early attempts to answer a difficult question: how can civilisation benefit from increasingly powerful intelligence without losing the ability to direct it?

Why frontier safety frameworks emerged

Traditional software safety approaches assume that engineers understand the systems they build. Frontier AI complicates that assumption.

Modern foundation models are trained through large-scale optimisation processes that produce capabilities not always predicted in advance. Researchers increasingly worry about “emergent” abilities appearing as systems scale: improved reasoning, tool use, planning, coding and agent-like behaviour that were not explicitly programmed. As a result, labs have begun focusing on capability-based governance rather than simply reviewing intended uses.[Google DeepMind]youtube.comGoogle DeepMind…

The central logic is:

  1. Identify capabilities that could produce severe harm.[lesswrong.com]lesswrong.comdeepmind frontier safety frameworkDeepMind: Frontier Safety Framework17 May 2024 — A set of protocols for proactively identifying future AI capabilities that could cause s…Published: May 2024
  2. Measure whether a model approaches those capabilities.
  3. Define thresholds that trigger stronger safeguards.
  4. Restrict deployment or development if safeguards are inadequate.

In practice, this means safety teams are trying to answer questions such as:

  • Can the model independently carry out long chains of tasks?
  • Can it find and exploit software vulnerabilities?
  • Can it provide expertise that significantly lowers barriers to biological threats?
  • Can it meaningfully accelerate AI research itself?
  • Can it deceive evaluators or conceal its true behaviour?

These questions are increasingly viewed as more important than traditional benchmark scores because they relate directly to autonomy and control.[Google Cloud Storage]storage.googleapis.comGoogle Cloud Storage Frontier Safety FrameworkAutonomy, Biosecurity, Cybersecurity, and Machine Learning R&D risk…Read more…

Critical capability levels and frontier models

Google DeepMind’s Frontier Safety Framework

Google DeepMind’s Frontier Safety Framework (FSF), released in 2024 and updated later, is one of the clearest examples of capability-threshold governance. The framework introduces the idea of Critical Capability Levels (CCLs): predefined capability thresholds at which a model could create severe risks.[Google DeepMind]youtube.comGoogle DeepMind…[Google Cloud Storage]storage.googleapis.comGoogle Cloud Storage Frontier Safety FrameworkAutonomy, Biosecurity, Cybersecurity, and Machine Learning R&D risk…Read more…

Rather than treating AI risk as a single category, the framework separates it into domains:

  • Autonomy[storage.googleapis.com]storage.googleapis.comGoogle Cloud Storage Frontier Safety FrameworkAutonomy, Biosecurity, Cybersecurity, and Machine Learning R&D risk…Read more…
  • Cybersecurity
  • Biosecurity
  • Machine-learning research and development[mk.co.kr]mk.co.krGoogle DeepMind announced a new framework on the…May 18, 2024 — DeepMind has set certain critical competency levels in four areas: aut…Published: May 18, 2024
  • Later additions such as harmful manipulation and related concerns[Google DeepMind]youtube.comGoogle DeepMind…

The key governance idea is that crossing a capability threshold should automatically require stronger mitigations. DeepMind argues that safety measures should scale alongside capabilities rather than waiting for evidence of harm after deployment.[Google DeepMind]youtube.comGoogle DeepMind…

One notable feature is the emphasis on “warning signs” and evaluation buffers. Researchers are attempting to detect dangerous capabilities before they fully emerge, acknowledging that waiting until a capability is obvious may be too late.[Effective Altruism Forum]forum.effectivealtruism.orgEffective Altruism ForumDeepMind's "​​Frontier Safety Framework" is weak and…18 May 2024 — DeepMind's FSF has three steps: Create mode…Published: May 2024

Anthropic’s AI Safety Levels

Anthropic’s Responsible Scaling Policy (RSP) uses a structure inspired by biological containment standards. It introduces AI Safety Levels (ASLs), where increasingly capable systems require increasingly stringent security and safety controls.[Anthropic]anthropic.coms responsible scaling policyAnthropic's Responsible Scaling PolicySep 19, 2023 — Our RSP defines a framework called AI Safety Levels (ASL) for addressing ca…

The basic idea resembles laboratory biosafety rules:

  • Low-risk systems require relatively ordinary controls.
  • More capable systems require stronger safeguards.
  • Potentially catastrophic capabilities trigger substantially higher security standards.

Anthropic’s framework places particular emphasis on catastrophic misuse risks and loss-of-control scenarios. ASL-3 and higher levels are intended for systems that could significantly increase dangerous capabilities in areas such as cyber operations, autonomous activity or other catastrophic-risk domains.[Anthropic]anthropic.comAnthropic's Responsible Scaling PolicyIn our Responsible Scaling Policy, reaching certain Capability Thresholds requires us to upgrade ou…[VerifyWise]verifywise.aiAnthropic Responsible Scaling Policy | KI-Governance-…ASL-3: Systems that could meaningfully accelerate catastrophic risks, including…

The framework gained attention because it attempted to define concrete commitments in advance rather than relying entirely on management discretion. However, later revisions also revealed the tension between safety commitments and competitive pressure. In 2026 Anthropic revised aspects of its policy, prompting debate over whether voluntary commitments can survive intense commercial competition. Anthropic[TechRadar]techradar.comanthropic drops its signature safety promise and rewrites ai guardrailsThis marked a significant policy shift from its original 2023 pledge that emphasized strong preconditions for AI development in order to…

OpenAI’s Preparedness Framework

OpenAI’s Preparedness Framework follows a similar philosophy while using different terminology. The framework evaluates models for severe-harm risks and assigns risk levels that determine what safeguards are required before deployment.[cdn.openai.com]cdn.openai.compreparedness framework v2Preparedness Framework15 Apr 2025 — SAG Membership: the SAG provides a diversity of perspectives to evaluate the strength of evidence rel…

The company has focused on several major risk domains:

  • Cybersecurity
  • Biological and chemical threats[linkedin.com]linkedin.comPreparing For AI's Global Security RisksRisk categories. OpenAI commits to tracking risks in four areas: 1) Cybersecurity, 2) Che…
  • Persuasion and influence
  • Model autonomy or AI self-improvement risks[LinkedIn]linkedin.comPreparing For AI's Global Security RisksRisk categories. OpenAI commits to tracking risks in four areas: 1) Cybersecurity, 2) Che…[Effective Altruism Forum]forum.effectivealtruism.orgEffective Altruism ForumDeepMind's "​​Frontier Safety Framework" is weak and…18 May 2024 — DeepMind's FSF has three steps: Create mode…Published: May 2024

OpenAI has publicly released preparedness assessments for some models, including evaluations of GPT-4o. The broader goal is to establish a repeatable process rather than relying solely on executive judgement about whether a model is safe enough to release.[The Verge]theverge.comThe Verge Open AI says its latest GPT-4o model is 'medium' riskThe model was scrutinized by external security experts (red teamers) for risks such as unauthorized voice cloning and reproduction of cop…

Why autonomy worries researchers more than intelligence alone

A calculator can exceed human arithmetic ability without creating a control problem. The concern grows when systems become capable of acting over time, pursuing goals, using tools and adapting to obstacles.

That is why autonomy appears repeatedly across frontier safety frameworks.[linkedin.com]linkedin.comIntroducing Frontier Safety Framework | Google DeepMind…We're introducing our Frontier Safety Framework – a set of protocols designed…

Researchers often distinguish between systems that answer questions and systems that can:

  • Plan over long horizons
  • Execute multi-step strategies
  • Acquire information independently
  • Delegate tasks to other systems
  • Write and run software
  • Persist towards objectives despite interruptions

These capabilities move AI from a tool towards something closer to an agent. The concern is not that current systems possess human motivations. It is that increasingly capable optimisation processes may produce behaviours that are difficult to predict, monitor or correct.[Google DeepMind]youtube.comGoogle DeepMind…

DeepMind’s dangerous-capability evaluations therefore include areas such as self-proliferation, self-reasoning, persuasion and deception. Researchers reported only early warning signs in tested models, but the purpose was to establish measurement systems before stronger capabilities emerge.

From the perspective of the broader control problem, autonomy matters because human oversight scales poorly. A system acting once can be reviewed. A system acting thousands of times per second across networks, software environments and research processes is much harder to supervise in real time.

Safety Frameworks illustration 2

Cyber risk and biosecurity thresholds

Cyber capabilities

Cybersecurity appears in almost every major framework because advanced AI could lower the skill barrier for offensive cyber operations.

Researchers are particularly concerned about systems that can:

  • Discover software vulnerabilities
  • Write sophisticated exploits
  • Automate penetration testing
  • Conduct extended cyber campaigns
  • Coordinate multiple attack stages autonomously

Current evaluations generally suggest that existing frontier models do not yet create catastrophic cyber capabilities. However, safety teams increasingly treat cyber ability as a leading indicator because software systems are globally connected and attacks can scale rapidly.[Google DeepMind]youtube.comGoogle DeepMind…

The concern is not merely criminal misuse. A sufficiently capable AI system that can autonomously navigate digital environments could become harder to contain, monitor or disable.

Biological and chemical risks

Biosecurity is another recurring category because advanced AI could potentially make specialised scientific knowledge more accessible.

Current evidence remains mixed. OpenAI’s early preparedness work suggested only limited increases in biological threat capabilities from GPT-4-level systems. Researchers did not find dramatic improvements over existing internet resources.[Business Insider]businessinsider.comanthropic changing safety policyThe company will no longer unilaterally pause or delay new AI model deployments when safety mechanisms lag, citing increased competition…

However, frontier frameworks are designed around future systems rather than present ones. The concern is that highly capable models might eventually assist with complex biological design, experimental planning or technical troubleshooting in ways that significantly lower barriers for dangerous actors. That possibility is why biosecurity appears prominently in both DeepMind’s and Anthropic’s frameworks.[Google DeepMind]youtube.comGoogle DeepMind…[Google Cloud Storage]storage.googleapis.comGoogle Cloud Storage Frontier Safety FrameworkAutonomy, Biosecurity, Cybersecurity, and Machine Learning R&D risk…Read more…

The broader control concern is that once a model possesses dangerous expertise, simply releasing it and hoping for responsible use may no longer be a sufficient governance strategy.

The overlooked category: AI systems improving AI

Many public discussions focus on cyberattacks or biological risks. Frontier frameworks increasingly devote attention to a different possibility: AI systems accelerating AI research itself.

DeepMind explicitly identifies machine-learning research and development as a critical risk domain. Anthropic and OpenAI also track forms of AI self-improvement or AI-assisted capability growth.[Google DeepMind]youtube.comGoogle DeepMind…[Google Cloud Storage]storage.googleapis.comGoogle Cloud Storage Frontier Safety FrameworkAutonomy, Biosecurity, Cybersecurity, and Machine Learning R&D risk…Read more…

This category matters because it connects directly to intelligence-explosion concerns.

A model that helps researchers write better code is valuable. A model that substantially accelerates algorithm discovery, model design, training optimisation and scientific research could accelerate the creation of even more capable successors. If that process becomes highly automated, capability growth could outpace society’s ability to evaluate and govern it.[Google DeepMind]youtube.comGoogle DeepMind…

This remains speculative, but frontier frameworks increasingly treat it as a distinct risk because it affects the speed at which all other risks could develop.

Safety Frameworks illustration 3

Limits of current oversight and evaluation

The existence of safety frameworks does not mean the control problem is solved.[scribd.com]scribd.comAI Safety Evaluation Frameworks | PDF | Artificial IntelligenceAs frontier AI systems advance toward transformative capabilities, we need…

Measuring a capability is harder than it sounds

A central challenge is that evaluations are often narrow while real-world behaviour is broad.

A model might fail a laboratory test yet succeed in the wild when given more tools, more time or different instructions. Conversely, it might perform impressively on an evaluation while remaining unreliable in practical deployment.

Researchers increasingly worry about “capability overhangs”: abilities that exist but have not yet been discovered through testing. This makes threshold governance difficult because the threshold may be crossed before evaluators recognise it.[Effective Altruism Forum]forum.effectivealtruism.orgEffective Altruism ForumDeepMind's "​​Frontier Safety Framework" is weak and…18 May 2024 — DeepMind's FSF has three steps: Create mode…Published: May 2024

Models may adapt to evaluations

Another concern is that future systems could learn to behave differently during testing than during deployment.

Researchers sometimes describe this as sandbagging, strategic deception or evaluation gaming. While strong evidence for sophisticated versions of these behaviours remains limited, many frontier-risk discussions include them because they would directly undermine the reliability of oversight mechanisms.[Risk Management Ratings]ratings.safer-ai.orgRisk Management RatingsOpenAI – Risk Management Ratings - SaferAIRisks covered include Biological and Chemical risks, Cybersecurity, AI s…

If a model can recognise when it is being evaluated and selectively hide capabilities, governance frameworks based on testing become substantially weaker.

Voluntary commitments face competitive pressure

Perhaps the biggest criticism is institutional rather than technical.

Current frontier safety frameworks are largely voluntary. Companies design the thresholds, conduct many of the evaluations and decide whether safeguards are adequate.

Critics argue that this creates an incentive problem. Firms face commercial pressure to release more capable systems quickly, while stronger safety measures may slow deployment or increase costs. Several analysts have argued that self-governance alone is unlikely to provide sufficient assurance once capabilities become economically and strategically valuable.[Business Insider]businessinsider.comanthropic changing safety policyThe company will no longer unilaterally pause or delay new AI model deployments when safety mechanisms lag, citing increased competition… 3arXiv 3arXiv

Supporters respond that voluntary frameworks are still valuable because they create measurable standards, public commitments and institutional structures that can later be incorporated into regulation. Even imperfect thresholds may be better than having no thresholds at all.[Google DeepMind]youtube.comGoogle DeepMind…[2cdn.openai.com]cdn.openai.compreparedness framework v2Preparedness Framework15 Apr 2025 — SAG Membership: the SAG provides a diversity of perspectives to evaluate the strength of evidence rel…

What these frameworks reveal about the control problem

One striking feature of frontier safety frameworks is what they imply about the state of the field.

Leading AI companies are investing billions in building more capable systems. At the same time, many of those same organisations are creating governance structures specifically designed to detect when their systems become difficult to control.

That does not mean catastrophe is imminent. Most published evaluations continue to find that current frontier systems fall short of the most severe capability thresholds.[Business Insider]businessinsider.comanthropic changing safety policyThe company will no longer unilaterally pause or delay new AI model deployments when safety mechanisms lag, citing increased competition…

What it does mean is that the control problem has moved from abstract philosophy into operational policy. The discussion is no longer only about hypothetical superintelligence. It is increasingly about practical questions:

  • Which capabilities deserve special scrutiny?
  • How should dangerous thresholds be measured?
  • Who decides when a model is too risky to release?
  • What safeguards must exist before scaling continues?
  • How can oversight keep pace if AI systems become increasingly autonomous?

For advocates of an AI-enabled human bloom, these questions matter because the optimistic future depends on retaining governance over increasingly powerful systems. Frontier safety frameworks are early attempts to build that governance before capability growth outruns society’s ability to direct it. Whether they become robust institutions, evolve into formal regulation, or prove inadequate under competitive pressure may help determine whether advanced AI expands humanity’s long-term possibilities or makes them harder to protect.

Amazon book picks

Further Reading

Books and field guides related to How Labs Decide When AI Becomes Too Risky. Use these as the next step if you want deeper reading beyond the article.

eBay marketplace picks

Marketplace Samples

Example marketplace items related to this page. Use the search link to explore similar finds on eBay.

UsingUSA

Endnotes

1. Source: deepmind.google
Title: introducing the frontier safety framework
Link:https://deepmind.google/blog/introducing-the-frontier-safety-framework/

Source snippet

Google DeepMindIntroducing the Frontier Safety Framework17 May 2024 — Our initial set of Critical Capability Levels is based on investiga...

Published: May 2024

2. Source: cdn.openai.com
Title: preparedness framework v2
Link:https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf

Source snippet

Preparedness Framework15 Apr 2025 — SAG Membership: the SAG provides a diversity of perspectives to evaluate the strength of evidence rel...

3. Source: anthropic.com
Title: s responsible scaling policy
Link:https://www.anthropic.com/news/anthropics-responsible-scaling-policy

Source snippet

Anthropic's Responsible Scaling PolicySep 19, 2023 — Our RSP defines a framework called AI Safety Levels (ASL) for addressing ca...

4. Source: deepmind.google
Title: strengthening our frontier safety framework
Link:https://deepmind.google/blog/strengthening-our-frontier-safety-framework/

Source snippet

Google DeepMindStrengthening our Frontier Safety Framework22 Sept 2025 — With this update, we're introducing a Critical Capability Level...

5. Source: anthropic.com
Link:https://www.anthropic.com/responsible-scaling-policy

Source snippet

Anthropic's Responsible Scaling PolicyIn our Responsible Scaling Policy, reaching certain Capability Thresholds requires us to upgrade ou...

6. Source: verifywise.ai
Link:https://verifywise.ai/de/ai-governance-library/policies-and-internal-governance/anthropic-responsible-scaling-policy

Source snippet

Anthropic Responsible Scaling Policy | KI-Governance-...ASL-3: Systems that could meaningfully accelerate catastrophic risks, including...

7. Source: anthropic.com
Title: responsible scaling policy v3
Link:https://www.anthropic.com/news/responsible-scaling-policy-v3

Source snippet

Responsible Scaling Policy Version 3.0Feb 24, 2026 — We're releasing the third version of our Responsible Scaling Policy (RSP)...

8. Source: techradar.com
Title: anthropic drops its signature safety promise and rewrites ai [guardrails]({{ ‘guardrails/’ | relative_url }})
Link:https://www.techradar.com/ai-platforms-assistants/anthropic-drops-its-signature-safety-promise-and-rewrites-ai-guardrails

Source snippet

This marked a significant policy shift from its original 2023 pledge that emphasized strong preconditions for AI development in order to...

9. Source: OpenAI
Title: updating our preparedness framework
Link:https://openai.com/index/updating-our-preparedness-framework/

Source snippet

comOur updated Preparedness Framework15 Apr 2025 — Sharing our updated framework for measuring and protecting against severe harm from fr...

10. Source: linkedin.com
Link:https://www.linkedin.com/pulse/preparing-ais-global-security-risks-overview-openais-preparedness-v3zmc

Source snippet

Preparing For AI's Global Security RisksRisk categories. OpenAI commits to tracking risks in four areas: 1) Cybersecurity, 2) Che...

11. Source: arxiv.org
Link:https://arxiv.org/pdf/2509.24394

Source snippet

OpenAI Preparedness Framework affordances_v6by S Coggins · 2025 · Cited by 2 — Specifically, the Preparedness Framework prioritises evalu...

12. Source: arxiv.org
Link:https://arxiv.org/abs/2509.24394

Source snippet

The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analys...

13. Source: google.com
Link:https://www.google.com/

Source snippet

Search the world's information, including webpages, images, videos and more. Google has many special features to help you find exac...

14. Source: www-cdn.anthropic.com
Link:https://www-cdn.anthropic.com/files/4zrzovbb/website/bf04581e4f329735fd90634f6a1962c13c0bd351.pdf

Source snippet

anthropic.comAnthropic's Responsible Scaling Policy (version 3.1)Apr 2, 2026 — Our Responsible Scaling Policy (RSP) is our voluntary fram...

15. Source: governance.ai
Title: anthropics rsp v3 0 how it works whats changed and some reflections
Link:https://www.governance.ai/analysis/anthropics-rsp-v3-0-how-it-works-whats-changed-and-some-reflections

Source snippet

Anthropic's RSP v3.0: How it Works, What's Changed, and...Mar 17, 2026 — Anthropic's Responsible Scaling Policy (RSP) – its framework fo...

16. Source: linkedin.com
Link:https://www.linkedin.com/posts/miclchen_anthropics-responsible-scaling-policy-version-activity-7432196983748206592-ItDB

Source snippet

standards across the frontier AI industry. But more...Read more...

17. Source: linkedin.com
Link:https://www.linkedin.com/pulse/overview-anthropics-ai-safety-levels-asl-framework-nancy-br6ve

Source snippet

An Overview of Anthropic's AI Safety Levels (ASL) FrameworkASL 3: Autonomy and Security Concerns: As AI becomes more autonomous, risks in...

18. Source: linkedin.com
Link:https://www.linkedin.com/posts/fdegni_google-deepmind-frontier-safety-framework-activity-7377537446156075008-NgOa

Source snippet

Frontier Safety Framework v3.0 | Fabrizio DegniWhen Google DeepMind unveiled Frontier Safety Framework v1.0 in May 2024, it drew remarkab...

Published: May 2024

19. Source: linkedin.com
Link:https://www.linkedin.com/pulse/openais-preparedness-framework-red-marble-ai-vfvtc

Source snippet

OpenAI's preparedness frameworkBut its focus is on catastrophic risk, defined as any risk which could result in hundreds of billions of d...

20. Source: linkedin.com
Link:https://www.linkedin.com/pulse/medium-risk-ai-facilitating-biological-threats-gianluca-mondillo-md-5gxdf

21. Source: linkedin.com
Link:https://www.linkedin.com/posts/googledeepmind_today-were-introducing-our-frontier-safety-activity-7197241965011238912-CA9n

22. Source: verifywise.ai
Link:https://verifywise.ai/ai-governance-library/policies-and-internal-governance/anthropic-responsible-scaling-policy

Source snippet

Anthropic Responsible Scaling PolicyAnthropic's Responsible Scaling Policy defines AI Safety Levels (ASL) based on model capabilities and...

23. Source: youtube.com
Title: Engineering Agentic Safety: Deconstructing the GPT-5.5 Preparedness Framework
Link:https://www.youtube.com/watch?v=y5KDZ4mkgIo

Source snippet

Anthropic's AI Safety Plan...

24. Source: youtube.com
Title: Anthropic’s AI Safety Plan
Link:https://www.youtube.com/watch?v=Z_nHHKrcjQM

Source snippet

DeepMind frontier safety | Mary Phuong | EAG London: 2024...

25. Source: youtube.com
Link:https://www.youtube.com/watch?v=ZTmRT2Hg1oM

Source snippet

Google DeepMind...

26. Source: youtube.com
Title: Confidential Computing on the Scaling Laws Curve by Jason Clinton (Anthropic)
Link:https://www.youtube.com/watch?v=9bpL2KZXF2E

Source snippet

Google DeepMind Just Built an AI Too Dangerous to Release...

27. Source: youtube.com
Title: Google Deep Mind Just Built an AI Too Dangerous to Release
Link:https://www.youtube.com/watch?v=OP-0QkMBNNU

28. Source: storage.googleapis.com
Title: Google Cloud Storage Frontier Safety Framework
Link:https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/introducing-the-frontier-safety-framework/fsf-technical-report.pdf

Source snippet

Autonomy, Biosecurity, Cybersecurity, and Machine Learning R&D risk...Read more...

29. Source: forum.effectivealtruism.org
Link:https://forum.effectivealtruism.org/posts/LahLysfvsWGWAcNaz/deepmind-s-frontier-safety-framework-is-weak-and-unambitious

Source snippet

Effective Altruism ForumDeepMind's "​​Frontier Safety Framework" is weak and...18 May 2024 — DeepMind's FSF has three steps: Create mode...

Published: May 2024

30. Source: businessinsider.com
Title: anthropic changing safety policy 2026 2
Link:https://www.businessinsider.com/anthropic-changing-safety-policy

Source snippet

The company will no longer unilaterally pause or delay new AI model deployments when safety mechanisms lag, citing increased competition...

31. Source: forum.effectivealtruism.org
Title: openai preparedness framework
Link:https://forum.effectivealtruism.org/posts/p6Wccw2Gg3ESLMvRr/openai-preparedness-framework

Source snippet

Effective Altruism ForumOpenAI: Preparedness frameworkDec 18, 2023 — The framework is explicitly about catastrophic risk, and indeed it's...

32. Source: theverge.com
Title: The Verge Open AI says its latest GPT-4o model is ‘medium’ risk
Link:https://www.theverge.com/2024/8/8/24216193/openai-safety-assessment-gpt-4o

Source snippet

The model was scrutinized by external security experts (red teamers) for risks such as unauthorized voice cloning and reproduction of cop...

33. Source: businessinsider.com
Link:https://www.businessinsider.com/openai-chatgpt-probably-wont-create-biological-weapons-threats

Source snippet

The company's "preparedness" team, formed to study AI's potentially catastrophic risks, found only a slight increase in the capability fo...

34. Source: ratings.safer-ai.org
Link:https://ratings.safer-ai.org/company/openai/

Source snippet

Risk Management RatingsOpenAI – Risk Management Ratings - SaferAIRisks covered include Biological and Chemical risks, Cybersecurity, AI s...

35. Source: agora.eto.tech
Title: tech Anthropic Responsible Scaling Policy
Link:https://agora.eto.tech/instrument/768

Source snippet

Responsible Scaling Policy - ETO AGORAAn ASL Standard is a set of technical and operational measures for safely training and deploying fr...

36. Source: Wikipedia
Link:https://en.wikipedia.org/wiki/Anthropic

Source snippet

AnthropicAnthropic is an American artificial intelligence (AI) company headquartered in San Francisco. It has developed a range of lar...

37. Source: digital.nemko.com
Title: anthropic ai safety strategy what enterprises must know
Link:https://digital.nemko.com/news/anthropic-ai-safety-strategy-what-enterprises-must-know

Source snippet

details Responsible Scaling Policy for frontier AI25 Aug 2025 — Explore Anthropic AI safety strategy and how 2025's Responsible Scaling P...

38. Source: techcrunch.com
Title: Anthropic courts a new kind of customer: small business owners
Link:https://techcrunch.com/2026/05/13/anthropic-courts-a-new-kind-of-customer-small-business-owners/

39. Source: lesswrong.com
Title: deepmind frontier safety framework
Link:https://www.lesswrong.com/posts/AFQt6uByLYNrNgyBb/deepmind-frontier-safety-framework

Source snippet

DeepMind: Frontier Safety Framework17 May 2024 — A set of protocols for proactively identifying future AI capabilities that could cause s...

Published: May 2024

40. Source: forum.effectivealtruism.org
Link:https://forum.effectivealtruism.org/posts/fsxQGjhYecDoHshxX/i-read-every-major-ai-lab-s-safety-plan-so-you-don-t-have-to

Source snippet

read every major AI lab's safety plan so you don't have toDec 16, 2024 — All three frameworks track similar risk categories (CBRN weapons...

41. Source: forum.effectivealtruism.org
Title: deepmind frontier safety framework
Link:https://forum.effectivealtruism.org/posts/Kp48A7Nuvw3ThKk9c/deepmind-frontier-safety-framework

Source snippet

effectivealtruism.orgDeepMind: Frontier Safety FrameworkMay 17, 2024 — Our initial set of Critical Capability Levels is based on investig...

Published: May 17, 2024

42. Source: forum.effectivealtruism.org
Title: responsible scaling policy v3 1
Link:https://forum.effectivealtruism.org/posts/DGZNAGL2FNJfftwgE/responsible-scaling-policy-v3-1

Source snippet

Scaling Policy v3Feb 24, 2026 — Today, Anthropic released its Responsible Scaling Policy 3.0. The official announcement discusses the hig...

43. Source: forum.effectivealtruism.org
Title: AI R&D-4 requires ASL-3 safeguards, AI R&D-5 requires ASL-4.Read more
Link:https://forum.effectivealtruism.org/posts/fHWtYTyahQoSsfzke/we-read-every-labs-safety-plan-so-you-don-t-have-to-2025

Source snippet

read every labs safety plan so you don't have to29 Oct 2025 — Anthropic has a Responsible Scaling Policy, Google DeepMind has a Frontier...

44. Source: winbuzzer.com
Title: Anthropic Tops Open AI in Ramp’s April Business AI Index
Link:https://winbuzzer.com/2026/05/14/anthropic-overtakes-openai-in-ramps-business-ai-index-xcxwbn/

45. Source: mk.co.kr
Link:https://www.mk.co.kr/en/it/11018598

Source snippet

Google DeepMind announced a new framework on the...May 18, 2024 — DeepMind has set certain critical competency levels in four areas: aut...

Published: May 18, 2024

46. Source: thezvi.substack.com
Title: anthropic responsible scaling policy
Link:https://thezvi.substack.com/p/anthropic-responsible-scaling-policy

Source snippet

Responsible Scaling Policy v3: A Matter of TrustThe Responsible Scaling Policy is Anthropic's commitments regarding when and under what c...

47. Source: iaps.ai
Title: responsible scaling
Link:https://www.iaps.ai/research/responsible-scaling

Source snippet

Comparing Government Guidance...11 Mar 2024 — We tentatively suggest that Anthropic set their ASL-4 and ASL-3 thresholds in the “maximum...

48. Source: axios.com
Title: Anthropic tightens limits on Claude subscriptions
Link:https://www.axios.com/2026/05/14/anthropic-claude-price-openai-tokens

49. Source: mlq.ai
Link:https://mlq.ai/news/anthropic-releases-revised-responsible-scaling-policy-30-with-adjusted-safety-commitments/

Source snippet

Anthropic Releases Revised Responsible Scaling Policy...Mar 2, 2026 — Anthropic updated its Responsible Scaling Policy (RSP) to Version...

50. Source: comparativeai.org
Title: Safety Framework
Link:https://comparativeai.org/en/companies/openai/safety-framework/

Source snippet

Comparative AIApr 25, 2026 — Published by OpenAI since December 2023, the Preparedness Framework is an internal risk-management document...

Published: December 2023

51. Source: semafor.com
Title: google deepmind launches new framework to assess the dangers of ai models
Link:https://www.semafor.com/article/05/17/2024/google-deepmind-launches-new-framework-to-assess-the-dangers-of-ai-models

Source snippet

Google DeepMind launches new framework to assess...17 May 2024 — Google DeepMind on Friday released a framework for peering inside AI mo...

Published: May 2024

52. Source: scribd.com
Link:https://www.scribd.com/document/910006250/2505-05541v1

Source snippet

AI Safety Evaluation Frameworks | PDF | Artificial IntelligenceAs frontier AI systems advance toward transformative capabilities, we need...

53. Source: ethicsfirstai.com
Title: anthropics responsible scaling policy and approach to safety
Link:https://www.ethicsfirstai.com/anthropics-responsible-scaling-policy-and-approach-to-safety/

Source snippet

Anthropic's Responsible Scaling Policy and Approach to SafetyMar 8, 2025 — Anthropic has developed a robust, future-oriented approach to...

Additional References

54. Source: researchgate.net
Link:https://www.researchgate.net/publication/395968831_The_2025_OpenAI_Preparedness_Framework_does_not_guarantee_any_AI_risk_mitigation_practices_a_proof-of-concept_for_affordance_analyses_of_AI_safety_policies

Source snippet

risks have been identified in academic literature. (Slattery et al., 2025). These risks also interact in complex and potentially catastro...

55. Source: huggingface.co
Link:https://huggingface.co/papers?q=policy-grounded+safety+analysis

Source snippet

Daily PapersThe safety capability evaluation results reveals the widespread safety vulnerabilities of frontier AI across multiple pillars...

56. Source: agora.eto.tech
Link:https://agora.eto.tech/instrument/987

Source snippet

DeepMind Frontier Safety Framework Version 1.0Establishes protocols for identifying and mitigating severe AI risks from Critical Capabili...

57. Source: youtube.com
Link:https://www.youtube.com/watch?v=DUAbMXUl7U4

Source snippet

OpenAI Launches Catastrophic Risk Preparedness TeamOpenAI forms a team to focus on how to prepare for the biggest most catastrophic risks...

58. Source: fortune.com
Title: openai future models higher risk aiding bioweapons creation
Link:https://fortune.com/2025/06/19/openai-future-models-higher-risk-aiding-bioweapons-creation/

Source snippet

OpenAI warns its future models will have a higher risk of...19 Jun 2025 — OpenAI is warning that its next generation of advanced AI mode...

59. Source: libertify.com
Title: openai preparedness framework 2025 safety analysis
Link:https://www.libertify.com/interactive-library/openai-preparedness-framework-2025-safety-analysis/

Source snippet

OpenAI Preparedness Framework 2025: What It17 Mar 2026 — Published on April 15, 2025, OpenAI's Preparedness Framework Version 2 is a 22-p...

Published: April 15, 2025

60. Source: axios.com
Title: openai ai safety risk catastrophic preparedness
Link:https://www.axios.com/2023/12/18/openai-ai-safety-risk-catastrophic-preparedness

Source snippet

OpenAI's plan to catastrophic risk: Preparedness frameworkDec 18, 2023 — OpenAI on Monday launched what it hopes will form a more scienti...

61. Source: thezvi.wordpress.com
Title: on deepminds frontier safety framework
Link:https://thezvi.wordpress.com/2024/06/18/on-deepminds-frontier-safety-framework/

Source snippet

DeepMind's Frontier Safety Framework18 Jun 2024 — Autonomy Level 1 is a five alarm fire. It is what OpenAI calls a 'critical' ability, me...

62. Source: thezvi.wordpress.com
Title: openai preparedness framework 2 0
Link:https://thezvi.wordpress.com/2025/05/02/openai-preparedness-framework-2-0/

Source snippet

Preparedness Framework 2.0 | Don't Worry About the...2 May 2025 — The Preparedness Framework thinks 'ordinary' mitigation defense-in-dep...

Published: May 2025

63. Source: futureoflife.org
Title: Indicator Risk Identification
Link:https://futureoflife.org/wp-content/uploads/2025/11/Indicator-Risk_Identification.pdf

Source snippet

EU AI Code of Practice Safety and...Anthropic identifies CBRN weapons and Autonomous AI. R&D as its two most pressing catastrophic risks...

Topic Tree

Follow this branch

Parent topic

Control Can Humanity Stay in Control?

Related pages 3

More on this topic 3