Within AI Governance
Can Humans Stay in Control of AI?
Keeping humans involved in important decisions may help ensure advanced AI expands human capability without replacing legitimate human judgement.
On this page
- Where AI advice meets human judgement
- Designing oversight for powerful systems
- Decisions that require human authority
Page outline Jump by section
Introduction
Human oversight is one of the central mechanisms proposed for keeping advanced AI systems connected to human judgement as their capabilities grow. In a future with superintelligence — AI systems that could exceed human abilities across many fields — oversight cannot simply mean a person checking a machine’s final answer. The challenge is designing systems where humans retain meaningful authority over goals, boundaries, interpretation and major decisions while still benefiting from AI’s extraordinary problem-solving ability.
The optimistic case for AI and human flourishing depends on this balance. If advanced AI can accelerate science, improve medicine, expand education and help solve global problems, humanity will need ways to guide those capabilities rather than merely react to them. Human oversight mechanisms aim to create that bridge: combining machine intelligence with human responsibility. Current AI governance frameworks already treat oversight as more than a symbolic “human in the loop” requirement, emphasising that people must be able to understand, monitor, challenge and override AI systems when necessary.[AI Act Service Desk]ai-act-service-desk.ec.europa.euAI Act Service Desk Article 14: Human oversight | AI Act Service DeskAI Act Service DeskArticle 14: Human oversight | AI Act Service DeskJune 13, 2024…
Where AI advice meets human judgement
The basic idea behind human oversight is simple: powerful AI should remain a source of advice, analysis and action under human authority rather than becoming an independent decision-maker whose choices cannot be meaningfully questioned.
For ordinary AI systems, this might mean a doctor reviewing an AI-generated diagnosis or a researcher checking an AI-generated hypothesis. For superintelligence-level systems, the problem becomes much harder. A system that can design experiments, negotiate complex strategies, write software, manage resources or propose policies may produce recommendations that humans struggle to evaluate directly.
This creates what researchers call a scalable oversight problem: how can humans supervise systems that may become too capable, too fast or too specialised for any individual person to fully understand? OpenAI’s public safety research describes scalable oversight as involving methods such as improved human-AI interfaces, verification tools and systems that can identify uncertainty and seek clarification from human supervisors.[OpenAI]OpenAIOpen AIHow we think about safety and alignment | Open AIHow we think about safety and alignment | OpenAI…
A useful distinction is that oversight is not the same as asking humans to approve every tiny action. A superintelligent system could potentially handle enormous amounts of routine reasoning while humans retain control over the questions that define direction and legitimacy:
- What goals should the system pursue?
- Which risks are unacceptable?
- Which decisions require democratic or expert approval?
- When should the system stop, explain itself or request human judgement?
- Who has the authority to override it?
The aim is not to reduce AI capability, but to ensure that greater intelligence remains connected to human purposes.
Designing oversight for powerful systems
Effective oversight is likely to require several layers rather than one single safeguard. No individual mechanism is likely to be sufficient if AI systems become far more capable than current models.
Human control built into the system
One approach is to design AI systems so that human intervention is possible by default. This includes clear controls, monitoring tools, access limits and the ability to pause or interrupt operations.
The European Union’s AI Act provides one current example of this governance direction. For high-risk AI systems, Article 14 requires systems to support human oversight by enabling people to monitor performance, understand limitations, avoid excessive reliance on AI outputs, override decisions and stop systems where necessary.[AI Act Service Desk]ai-act-service-desk.ec.europa.euAI Act Service Desk Article 14: Human oversight | AI Act Service DeskAI Act Service DeskArticle 14: Human oversight | AI Act Service DeskJune 13, 2024…
For future superintelligent systems, similar principles would likely need to become more advanced. A simple “approve” button may be meaningless if a human cannot understand what they are approving. Oversight tools may need to include:
- explanations of why a system reached a conclusion;
- uncertainty estimates showing where the AI may be unreliable;
- independent evaluation systems checking important decisions;
- records of actions and reasoning paths for later review;
- protected shutdown or containment procedures.
The difficult question is whether these tools remain effective when AI systems become highly strategic. A system that can outperform humans intellectually may also be better at presenting convincing explanations. Oversight therefore requires not only access to information, but confidence that the information itself is trustworthy.
AI systems supervising AI systems
Because humans may struggle to evaluate extremely advanced AI outputs, one proposed mechanism is to use AI assistants to help oversee other AI systems.
This idea appears in work on AI-assisted supervision and methods such as Constitutional AI. Anthropic’s Constitutional AI research explores training models using explicit principles and AI-generated critiques rather than relying only on large amounts of human feedback. The goal is to make AI behaviour easier to steer as systems become more capable.[anthropic.com]anthropic.comConstitutional AI: Harmlessness from AI Feedback \ AnthropicDecember 15, 2022…
AI-based oversight could help humans by:
- checking whether another AI follows agreed rules;
- identifying suspicious behaviour;
- comparing multiple possible solutions;
- translating complex technical reasoning into understandable summaries;
- highlighting disagreements or uncertainty.
However, this approach creates a further governance challenge: the overseer AI must itself be trustworthy. If one powerful system is used to judge another, humans still need ways to test, audit and challenge the supervisory layer.
For this reason, many researchers see future oversight as a hierarchy: multiple systems checking each other, combined with human institutions that retain ultimate authority.
Principles, constitutions and shared values
Another mechanism focuses on giving AI systems clearer principles to follow. Instead of relying only on examples of acceptable behaviour, researchers have explored whether AI systems can be guided by written principles describing important values and constraints.
Anthropic’s Constitutional AI work is one example. It uses a set of explicit principles to guide model behaviour and explores whether AI systems can critique and improve their own responses according to those principles.[anthropic.com]anthropic.comConstitutional AI: Harmlessness from AI Feedback \ AnthropicDecember 15, 2022…
A related question is who gets to define those principles. If superintelligent systems influence society, their guiding rules cannot simply reflect the preferences of a small group of developers or organisations. Anthropic has also explored approaches involving public input into AI constitutions, examining whether broader participation could shape AI behaviour.[anthropic.com]anthropic.comOctober 17, 2023…
This connects technical alignment with democratic legitimacy. Human oversight is not only about preventing accidents; it is also about ensuring that humanity has a meaningful role in deciding what powerful AI systems should value.
Decisions that require human authority
Not every decision should be treated equally. A key governance question is deciding where humans must remain directly responsible.
In an AI bloom scenario, many decisions could reasonably become AI-assisted. A superintelligent system might help scientists identify promising treatments, help engineers design cleaner energy systems or help governments analyse policy options. The value comes from expanding human ability.
But some decisions involve legitimacy, rights or irreversible consequences. These may require human authority even if AI systems provide better predictions.
Examples could include:
- decisions affecting fundamental rights or personal freedoms;
- choices about the distribution of AI-created wealth or resources;
- military decisions involving the use of force;
- changes to long-term social priorities;
- decisions that permanently affect future generations.
The reason is not that humans will always be better at technical judgement. A superintelligent system may outperform people in many forms of reasoning. The reason is that some decisions are not purely technical. They involve values, responsibility and collective consent.
A flourishing future would require both intelligence and wisdom: AI capability combined with institutions that allow humans to decide what kind of civilisation they want to build.
The limits of keeping humans “in the loop”
Human oversight can fail if it becomes a superficial procedure rather than genuine control.
One risk is automation bias: people may trust AI recommendations too much, especially when systems appear highly accurate. The EU AI Act explicitly recognises the danger that human operators may over-rely on AI outputs and requires oversight measures that account for this tendency.[AI Act Service Desk]ai-act-service-desk.ec.europa.euAI Act Service Desk Article 14: Human oversight | AI Act Service DeskAI Act Service DeskArticle 14: Human oversight | AI Act Service DeskJune 13, 2024…
Another problem is speed. A human reviewer may be able to evaluate one AI decision, but not thousands of complex actions happening rapidly. Oversight must therefore be designed around appropriate levels of abstraction. Humans may need to supervise goals, constraints and system behaviour rather than individual actions.
There is also a deeper philosophical challenge: if an AI system becomes vastly more capable than humans, meaningful oversight may depend on mechanisms that are not yet fully developed. Researchers continue to debate whether scalable oversight is primarily an engineering challenge or whether it involves deeper limits on human ability to evaluate intelligence greater than our own.[arXiv]arxiv.orgarXiv Modeling Human Beliefs about AI Behavior for Scalable OversightModeling Human Beliefs about AI Behavior for Scalable OversightFebruary 28, 2025…
A future where humans guide intelligence, rather than compete with it
Human oversight mechanisms are ultimately about preserving human agency in a world where intelligence may become abundant. The goal is not to keep humans performing tasks that machines can do better, but to ensure that the direction of civilisation remains connected to human values, rights and aspirations.
A successful superintelligence future would likely not involve humans manually controlling every action of advanced AI. Instead, it would involve carefully designed institutions, technical safeguards and decision systems where AI provides extraordinary capability while humans retain authority over purpose and meaning.
For the AI bloom vision to become reality, the question is not simply whether machines can become more intelligent than people. It is whether humanity can create a relationship with those systems in which greater intelligence leads to greater freedom, creativity, health and flourishing rather than a loss of human control.
Amazon book picks
Further Reading
Books and field guides related to Can Humans Stay in Control of AI?. Use these as the next step if you want deeper reading beyond the article.
The Alignment Problem: Machine Learning and Human Values
Finalist for the Los Angeles Times Book Prize A jaw-dropping exploration of everything that goes wrong when we build AI systems and the m...
Human Compatible: Artificial Intelligence and the Problem of...
A leading artificial intelligence researcher lays out a new approach to AI that will enable us to coexist successfully with increasingly...
Life 3.0: Being Human in the Age of Artificial Intelligence
'This is the most important conversation of our time, and Tegmark's thought-provoking book will help you join it' Stephen Hawking THE INT...
Rebooting AI: Building Artificial Intelligence We Can Trust
Two leaders in the field offer a compelling analysis of the current state of the art and reveal the steps we must take to achieve a robus...
eBay marketplace picks
Marketplace Samples
Live-tested eBay searches with available results related to this page.
Selected fromrobotics kit oneBay.co.uk.
Endnotes
1.
Source: OpenAI
Title: Open AIHow we think about safety and alignment | Open AI
Link:https://openai.com/safety/how-we-think-about-safety-alignment/
Source snippet
How we think about safety and alignment | OpenAI...
2.
Source: anthropic.com
Title: Constitutional AI: Harmlessness from AI Feedback \ Anthropic
Link:https://www.anthropic.com/news/constitutional-ai-harmlessness-from-ai-feedback?trk=public_post_comment-text
Source snippet
December 15, 2022...
Published: December 15, 2022
3.
Source: anthropic.com
Link:https://www.anthropic.com/research/collective-constitutional-ai-aligning-a-language-model-with-public-input
Source snippet
October 17, 2023...
Published: October 17, 2023
4.
Source: arxiv.org
Title: arXiv Modeling Human Beliefs about AI Behavior for Scalable Oversight
Link:https://arxiv.org/abs/2502.21262
Source snippet
Modeling Human Beliefs about AI Behavior for Scalable OversightFebruary 28, 2025...
Published: February 28, 2025
5.
Source: arxiv.org
Title: arXiv Limits of Safe AI Deployment: Differentiating Oversight and Control
Link:https://arxiv.org/abs/2507.03525
6.
Source: OpenAI
Title: how we monitor internal coding agents misalignment
Link:https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/
7.
Source: OpenAI
Title: reasoning models chain of thought controllability
Link:https://openai.com/index/reasoning-models-chain-of-thought-controllability/
8.
Source: OpenAI
Title: collective alignment aug 2025 updates
Link:https://openai.com/index/collective-alignment-aug-2025-updates/
9.
Source: anthropic.com
Title: Specific versus General Principles for Constitutional AI \ Anthropic
Link:https://www.anthropic.com/news/specific-versus-general-principles-for-constitutional-ai
10.
Source: OpenAI
Title: governance of superintelligence
Link:https://openai.com/index/governance-of-superintelligence/
11.
Source: anthropic.com
Title: Claude’s Constitution
Link:https://www.anthropic.com/news/claudes-constitution?stream=top
12.
Source: OpenAI
Title: how should ai systems behave
Link:https://openai.com/index/how-should-ai-systems-behave/
13.
Source: anthropic.com
Title: Constitutional AI: Harmlessness from AI Feedback \ Anthropic
Link:https://www.anthropic.com/news/constitutional-ai-harmlessness-from-ai-feedback
14.
Source: en.ai-act.io
Link:https://en.ai-act.io/article/high-risk-ai-systems/requirements-for-high-risk-ai-systems/human-oversight
15.
Source: ai-act-service-desk.ec.europa.eu
Title: AI Act Service Desk Article 14: Human oversight | AI Act Service Desk
Link:https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-14
Source snippet
AI Act Service DeskArticle 14: Human oversight | AI Act Service DeskJune 13, 2024...
Published: June 13, 2024
16.
Source: digital-strategy.ec.europa.eu
Title: eu Navigating the AI Act | Shaping Europe’s digital future
Link:https://digital-strategy.ec.europa.eu/en/faqs/navigating-ai-act
17.
Source: regulation-ai.eu
Title: article 14
Link:https://www.regulation-ai.eu/en/articles/article-14/
18.
Source: aisecurityandsafety.org
Title: scalable oversight
Link:https://aisecurityandsafety.org/en/guides/scalable-oversight/
19.
Source: ai-solutions.daviesmeyer.com
Title: scalable oversight
Link:https://ai-solutions.daviesmeyer.com/en/glossary/scalable-oversight
20.
Source: eur-lex.europa.eu
Title: eu Regulation
Link:https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=celex%3A32024R1689
21.
Source: ai-act-service-desk.ec.europa.eu
Title: eu Recital 73 | AI Act Service Desk
Link:https://ai-act-service-desk.ec.europa.eu/en/ai-act/recital-73
22.
Source: ai-act-service-desk.ec.europa.eu
Title: eu Recital 91 | AI Act Service Desk
Link:https://ai-act-service-desk.ec.europa.eu/en/ai-act/recital-91
23.
Source: alignmentsurvey.com
Title: Scalable Oversight | AI Alignment
Link:https://alignmentsurvey.com/[materials
24.
Source: europarl.europa.eu
Title: eu Texts adopted
Link:https://www.europarl.europa.eu/doceo/document/TA-9-2023-0236_EN.html
25.
Source: edps.europa.eu
Link:https://www.edps.europa.eu/data-protection/technology-monitoring/techsonar/scalable-oversight_en
26.
Source: alignmentlibrary.org
Title: Scalable Oversight
Link:https://alignmentlibrary.org/en/problems/other/scalable-oversight
Additional References
27.
Source: artificialintelligenceact.eu
Link:https://artificialintelligenceact.eu/article/14/
Source snippet
Article 14: Human Oversight | EU Artificial Intelligence ActToday — Part of Chapter III: High-Risk AI System ➔ Section 2: Requirements fo...
28.
Source: youtube.com
Title: Ex-Google CEO WARNS: “Humanity Is Running Out Of Time”
Link:https://www.youtube.com/watch?v=nn9DoI0BFJ4
Source snippet
This video selection explores scalable oversight frameworks, the alignment challenges of advanced autonomous systems, and the imperative...
29.
Source: nature.com
Link:https://www.nature.com/articles/s41746-026-02971-1
30.
Source: nist.gov
Link:https://www.nist.gov/speech-testimony/balancing-knowledge-and-governance-foundations-effective-risk-management-artificial
31.
Source: youtube.com
Title: Agentic AI and the Problem of Control: Who’s in Charge?
Link:https://www.youtube.com/watch?v=TZS98IGs_20
Source snippet
Superintelligent AI: Can Humans Control What They Build?...
32.
Source: youtube.com
Title: Superintelligent AI: Can Humans Control What They Build?
Link:https://www.youtube.com/watch?v=n8qy3G0OCMA
Source snippet
Ex-Google CEO WARNS: "Humanity Is Running Out Of Time"...
33.
Source: youtube.com
Title: How to Align AI: Put It in a Sandwich
Link:https://www.youtube.com/watch?v=5mco9zAamRk
Source snippet
Maja Trębacz - Scalable Oversight: A Practical Approach to Verifying Code at Scale...
34.
Source: service.betterregulation.com
Link:https://service.betterregulation.com/document/2026/08/02/742206
Source snippet
"14 Human oversight | Regulation 2024/1689/EU - Artificial Intelligence Act (EU AI Act) | Better RegulationToday — [https://service.betterr..."](https://service.betterr...")...
35.
Source: techcrunch.com
Link:https://techcrunch.com/2026/01/21/anthropic-revises-claudes-constitution-and-hints-at-chatbot-consciousness/
36.
Source: zhongzhuzhou.org
Link:https://www.zhongzhuzhou.org/blog/2026-03-31-2026-03-31-ConstitutionalAI-technical-review-en/



