Within Alignment
Who Gets To Set AI Values?
Rule-based AI principles aim to create consistent behaviour, but choosing those principles raises questions about whose values guide advanced systems.
On this page
- How principle based AI guidance works
- Conflicts between cultures and values
- Creating legitimate alignment standards
Page outline Jump by section
Introduction
Advanced AI systems may become powerful enough to influence decisions about health, education, science, security, public information and the wider direction of society. A central challenge is therefore not only making AI capable, but deciding what principles should guide that capability. “Constitutional AI” refers to approaches where an AI system is shaped by explicit written principles rather than only by examples of preferred behaviour or human ratings. The idea is attractive because rules can make values visible and reviewable. Yet it creates a deeper governance question: who has the authority to write the constitution?
The challenge is especially important for an AI-enabled human bloom. If advanced AI helps expand knowledge, improve health, reduce scarcity and increase human agency, its guiding principles will influence whether those benefits support diverse forms of flourishing or reflect the priorities of a narrow group. The central dispute is not simply which rules an AI should follow, but how humanity should decide those rules when cultures, political systems and moral traditions disagree.[anthropic.com]anthropic.comConstitutional AI: Harmlessness from AI Feedback \ AnthropicDecember 15, 2022…
How principle-based AI guidance works
Constitutional approaches attempt to move AI alignment from an invisible training process towards explicit standards. Instead of relying only on large numbers of human evaluations, developers provide a written set of principles that describe desired behaviour. The AI can then critique and revise its own responses according to those principles, creating a training process based partly on AI-generated feedback rather than only human labels.[anthropic.com]anthropic.comConstitutional AI: Harmlessness from AI Feedback \ AnthropicDecember 15, 2022…
The appeal is that principles provide a shared reference point. A rule such as “respect human dignity”, “avoid causing unnecessary harm”, or “be honest about uncertainty” is easier to discuss than a hidden pattern learned from millions of examples. In theory, this makes advanced systems more consistent and gives society something concrete to examine.
A constitutional approach can also help with a practical problem in AI development: human feedback is limited. People cannot label every possible situation a highly capable system might encounter. Researchers at Anthropic have explored whether models can learn broader behaviours from a smaller number of written principles, including principles focused on being helpful, harmless and aligned with broader human interests.[anthropic.com]anthropic.comSpecific versus General Principles for Constitutional AI \ AnthropicOctober 24, 2023…
However, a constitution does not remove the value question. It moves that question earlier in the process. Someone still has to decide:
- Which principles belong in the constitution?
- Which harms deserve priority?
- Whose understanding of fairness, dignity or freedom should guide decisions?
- How should conflicts between principles be resolved?
A constitution can make values clearer, but it cannot make value choices disappear.
Why different societies disagree about AI values
Human values overlap in important ways, but they are not identical. Many societies support broad goals such as avoiding unnecessary suffering, protecting human dignity and improving knowledge. Yet disagreement appears when principles collide.
For example, an AI assistant might need to balance:
- Freedom versus protection: Should an AI allow more individual choice even when decisions may create personal risks, or should it intervene more strongly to prevent harm?
- Privacy versus collective benefit: Should medical AI systems use more personal data to accelerate research, or should stronger privacy protections limit data access?
- Equality versus efficiency: Should AI systems prioritise equal outcomes across groups, or focus on maximising overall benefits even if gains are unevenly distributed?
- Transparency versus security: Should AI decision processes always be open, or can some information remain restricted to prevent misuse?
These conflicts are not technical bugs. They are political and moral disagreements that existed before AI. Advanced AI makes them more urgent because the systems may operate at a scale where embedded assumptions affect millions or billions of people.
International efforts show both agreement and disagreement. UNESCO’s global AI ethics framework places human rights, dignity, diversity, inclusion, transparency and human oversight at the centre of AI governance. At the same time, it recognises that implementation requires cooperation among different societies and institutions.[UNESCO]unesco.orgEthics of Artificial IntelligenceEthics of Artificial Intelligence - AI | UNESCO…
The difficulty is that even widely supported principles can be interpreted differently. “Fairness”, for example, might mean equal treatment, correcting historical disadvantages, or ensuring equal access to opportunities. An AI system cannot simply apply fairness without some interpretation of what fairness means in a particular context.
The problem of hidden value choices
One risk of constitutional AI is that a constitution may appear neutral while containing contested assumptions.
A set of principles written by a technology company, government or expert group may reflect the priorities of that institution. Even well-intentioned designers may unintentionally prioritise certain cultural ideas about individual rights, social responsibility, authority or acceptable risk.
This does not mean all principles are equally valid or that AI should have no constraints. A powerful AI system cannot operate without some boundaries. The challenge is creating boundaries that are legitimate, transparent and open to challenge.
The difference is similar to the difference between a private company creating internal workplace rules and a society creating laws. Both involve principles, but one has a much stronger expectation of public accountability.
This is why many AI governance discussions focus on participation rather than simply finding the “correct” values. A future where AI influences civilisation-scale decisions may require mechanisms that allow affected communities to contribute to how systems are governed. UNESCO’s approach emphasises multi-stakeholder governance, human oversight and inclusion as parts of responsible AI development.[UNESCO]unesco.orgEthics of Artificial IntelligenceEthics of Artificial Intelligence - AI | UNESCO…
When principles conflict inside an AI system
Even after agreeing on broad principles, advanced AI systems still need ways to handle conflicts.
Consider a medical AI assistant:
- A patient may want complete privacy.
- Researchers may want access to health data to discover treatments.
- Doctors may want transparency about how recommendations are produced.
- Hospitals may want efficiency and lower costs.
All of these goals can support human flourishing, but they cannot always be maximised simultaneously.
A constitutional AI system therefore needs more than a list of values. It needs a method for reasoning about trade-offs. Possible approaches include:
- Priority rules: Some principles always override others, such as preventing severe harm before improving convenience.
- Context-sensitive judgement: Different situations require different balances.
- Human review: Important conflicts are escalated rather than automatically resolved by AI.
- Public accountability: Decisions about major trade-offs are subject to oversight.
The European Union’s AI Act reflects this kind of layered approach. For high-risk AI systems, it requires mechanisms for human oversight so people can monitor, interpret and, where appropriate, override AI decisions.[AI Act Service Desk]ai-act-service-desk.ec.europa.euAI Act Service Desk Article 14: Human oversight | AI Act Service DeskAI Act Service DeskArticle 14: Human oversight | AI Act Service DeskJune 13, 2024…
Creating legitimate alignment standards
For advanced AI to contribute to human flourishing, the question is not only “what values should AI have?” but “how should humanity decide?”
A credible alignment framework may need several features.
Shared foundations without forced uniformity
Some principles may achieve broad agreement across societies: avoiding deliberate harm, maintaining human accountability, protecting basic rights and preserving the ability of people to challenge powerful institutions. These can form a foundation without requiring every culture to share identical moral views.
Democratic and institutional legitimacy
Powerful AI systems should not be governed solely by the preferences of the organisations that build them. Public institutions, civil society, researchers and affected communities may all need roles in shaping standards.
This does not mean every AI decision should become a political negotiation. It means the underlying rules governing powerful systems should have legitimacy beyond technical convenience.
Ability to revise principles
Human values change. Societies reconsider ideas about privacy, equality, disability, rights and responsibilities over time. A rigid AI constitution could preserve outdated assumptions.
A better model may treat AI principles as living frameworks: stable enough to provide guidance, but open to review as societies learn more about AI and its effects.
Distinguishing universal protections from contested preferences
One important governance challenge is separating basic protections from political or cultural preferences.
For example, preventing deception or protecting people from coercion may be widely accepted as a foundation. But questions about education, economic policy, speech norms or cultural expression may involve legitimate disagreement.
A flourishing-oriented AI future may therefore require systems that protect a shared moral floor while allowing diversity above that floor.
Why this matters for the AI bloom vision
The optimistic case for advanced AI depends on more than intelligence growth. Superintelligent systems could potentially accelerate scientific discovery, improve medicine, expand education and help humanity solve large-scale problems. But whether those capabilities produce broad flourishing depends partly on the principles that guide them.
A future of abundance could still fail to serve humanity if AI systems are shaped around narrow goals, concentrated power or one group’s vision of the good life. Conversely, carefully designed constitutional principles could help advanced AI remain compatible with human agency, diversity and long-term wellbeing.
The deepest challenge is therefore not writing a perfect list of rules. It is building legitimate processes for deciding, revising and applying principles in a world where humans themselves do not fully agree on what flourishing means. For advanced AI, alignment is ultimately a question of collective governance: how a diverse civilisation chooses to guide technologies powerful enough to influence its future.
Amazon book picks
Further Reading
Books and field guides related to Who Gets To Set AI Values?. Use these as the next step if you want deeper reading beyond the article.
Human Compatible
A leading artificial intelligence researcher lays out a new approach to AI that will enable us to coexist successfully with increasingly...
The Alignment Problem
Finalist for the Los Angeles Times Book Prize A jaw-dropping exploration of everything that goes wrong when we build AI systems and the m...
Superintelligence
This profoundly ambitious and original book picks its way carefully through a vast tract of forbiddingly difficult intellectual terrain.
Life 3.0
'This is the most important conversation of our time, and Tegmark's thought-provoking book will help you join it' Stephen Hawking THE INT...
eBay marketplace picks
Marketplace Samples
Live-tested eBay searches with available results related to this page.
Selected fromtechnology poster oneBay.co.uk.
Endnotes
1.
Source: anthropic.com
Title: Constitutional AI: Harmlessness from AI Feedback \ Anthropic
Link:https://www.anthropic.com/news/constitutional-ai-harmlessness-from-ai-feedback
Source snippet
December 15, 2022...
Published: December 15, 2022
2.
Source: unesco.org
Title: Ethics of Artificial Intelligence
Link:https://www.unesco.org/en/artificial-intelligence/recommendation-ethics?hub=70235
Source snippet
Ethics of Artificial Intelligence - AI | UNESCO...
3.
Source: anthropic.com
Title: Specific versus General Principles for Constitutional AI \ Anthropic
Link:https://www.anthropic.com/news/specific-versus-general-principles-for-constitutional-ai
Source snippet
October 24, 2023...
Published: October 24, 2023
4.
Source: anthropic.com
Link:https://www.anthropic.com/research/collective-constitutional-ai-aligning-a-language-model-with-public-input
5.
Source: anthropic.com
Title: Claude’s Constitution
Link:https://www.anthropic.com/news/claudes-constitution?stream=top
6.
Source: unesdoc.unesco.org
Title: document Viewer.xhtml
Link:https://unesdoc.unesco.org/in/documentViewer.xhtml?ark=%2Fark%3A%2F48223%2Fpf0000381137%2FPDF%2F381137eng.pdf.multi&file=%2Fin%2Frest%2FannotationSVC%2FDownloadWatermarkedAttachment%2Fattach_import_75c9fb6b-92a6-4982-b772-79f540c9fc39%3F_%3D381137eng.pdf&fullScreen=true&id=p%3A%3Ausmarcdef_0000381137&locale=en&updateUrl=updateUrl6452&v=2.1.196
7.
Source: unesdoc.unesco.org
Title: document Viewer.xhtml
Link:https://unesdoc.unesco.org/in/documentViewer.xhtml?ark=%2Fark%3A%2F48223%2Fpf0000381137%2FPDF%2F381137eng.pdf.multi&file=%2Fin%2Frest%2FannotationSVC%2FDownloadWatermarkedAttachment%2Fattach_import_e86c4b5d-5af9-4e15-be60-82f1a09956fd%3F_%3D381137eng.pdf&fullScreen=true&id=p%3A%3Ausmarcdef_0000381137&locale=fr&updateUrl=updateUrl2468&v=2.1.196
8.
Source: unesco.org
Title: Ethics of Artificial Intelligence
Link:https://www.unesco.org/en/artificial-intelligence/recommendation-ethics?hub=918
9.
Source: unesco.org
Title: Ethics of Artificial Intelligence
Link:https://www.unesco.org/en/artificial-intelligence/recommendation-ethics?hub=66778
10.
Source: unesco.org
Title: Ethics of Artificial Intelligence
Link:https://www.unesco.org/en/artificial-intelligence/recommendation-ethics?hub=158596
11.
Source: anthropic.com
Link:https://www.anthropic.com/constitution
12.
Source: unesco.de
Title: Artificial Intelligence
Link:https://www.unesco.de/en/topics/science/ethics-emerging-technologies/ai/
13.
Source: en.ai-act.io
Link:https://en.ai-act.io/article/high-risk-ai-systems/requirements-for-high-risk-ai-systems/human-oversight
14.
Source: ai-act-service-desk.ec.europa.eu
Title: AI Act Service Desk Article 14: Human oversight | AI Act Service Desk
Link:https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-14
Source snippet
AI Act Service DeskArticle 14: Human oversight | AI Act Service DeskJune 13, 2024...
Published: June 13, 2024
15.
Source: digital-strategy.ec.europa.eu
Title: eu Navigating the AI Act | Shaping Europe’s digital future
Link:https://digital-strategy.ec.europa.eu/en/faqs/navigating-ai-act
Source snippet
the AI Act | Shaping Europe’s digital futureJuly 27, 2026 — NAVIGATING THE AI ACT These questions and answers detail the AI Act’s goals...
Published: July 27, 2026
16.
Source: digital-strategy.ec.europa.eu
Link:https://digital-strategy.ec.europa.eu/lt/node/9745
Source snippet
Act | Shaping Europe’s digital futureJuly 27, 2026 — AI ACT The AI Act is the first-ever legal framework on AI, which addresses the risks...
Published: July 27, 2026
17.
Source: digital-strategy.ec.europa.eu
Title: eu A I Act | Shaping Europe’s digital future
Link:https://digital-strategy.ec.europa.eu/da/node/9745
Source snippet
Act | Shaping Europe’s digital futureJuly 24, 2026 — A RISK-BASED APPROACH The AI Act defines 4 levels of risk for AI systems: Image: pyr...
Published: July 24, 2026
18.
Source: consilium.europa.eu
Title: eu Artificial intelligence act
Link:https://www.consilium.europa.eu/en/policies/artificial-intelligence-act/
19.
Source: eur-lex.europa.eu
Title: eu Regulation
Link:https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=celex%3A32024R1689
20.
Source: ai-act-service-desk.ec.europa.eu
Link:https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-50
21.
Source: ai-act-service-desk.ec.europa.eu
Title: eu Recital 72 | AI Act Service Desk
Link:https://ai-act-service-desk.ec.europa.eu/en/ai-act/recital-72
22.
Source: ai-act-service-desk.ec.europa.eu
Title: eu Recital 73 | AI Act Service Desk
Link:https://ai-act-service-desk.ec.europa.eu/en/ai-act/recital-73
23.
Source: ai-act-service-desk.ec.europa.eu
Title: eu Recital 91 | AI Act Service Desk
Link:https://ai-act-service-desk.ec.europa.eu/en/ai-act/recital-91
24.
Source: plato.stanford.edu
Link:https://plato.stanford.edu/archives/spr2019/entries/legitimacy/
25.
Source: unesco.org.nz
Link:https://unesco.org.nz/knowledge-hub/ethics-of-artificial-intelligence-recommendation
Additional References
26.
Source: youtube.com
Link:https://www.youtube.com/watch?v=5Q46fJbjdSw
Source snippet
First AI Constitution arrives | Anthropic brings Claude's living doc | Mentor Sandy - Billion Hopes...
27.
Source: youtube.com
Title: Constitutional AI: Teaching Machines a Moral Compass for Safety
Link:https://www.youtube.com/watch?v=8Wzs3HxJQ0c
Source snippet
Constitutional AI Explained | Two-Phase Training (Supervised + RL-CAI) | Anthropic Paper...
28.
Source: papers.ssrn.com
Link:https://papers.ssrn.com/sol3/Delivery.cfm/6940538.pdf?abstractid=6940538&mirid=1
Source snippet
Constitutionalism: Why AI's 'Political Bias' Is a Social Contract Problem, Not Only a Technical One by Daniel Ziekenoppasser-Powell:: SS...
29.
Source: cambridge.org
Link:https://www.cambridge.org/core/journals/german-law-journal/article/from-principle-to-practice-value-alignment-in-ai-ethics-and-governance/3E75AC1B06FF6438A8D92C181C09D1E8
30.
Source: neurips.cc
Link:https://neurips.cc/virtual/2024/99059
31.
Source: gcedclearinghouse.org
Link:https://gcedclearinghouse.org/en/node/129375?level=7
32.
Source: cambridge.org
Link:https://www.cambridge.org/core/journals/data-and-policy/article/reversing-the-logic-of-generative-ai-alignment-a-pragmatic-approach-for-public-interest/8801BCA3832E848593E8D7F926C242CF
33.
Source: ppr.lse.ac.uk
Link:https://ppr.lse.ac.uk/articles/10.31389/lseppr.113
34.
Source: kclpure.kcl.ac.uk
Link:https://kclpure.kcl.ac.uk/portal/en/publications/decentralising-llm-alignment-a-case-for-context-pluralism-and-par/
35.
Source: researchportal.hkust.edu.hk
Link:https://researchportal.hkust.edu.hk/en/publications/democratizing-value-alignment-from-authoritarian-to-democratic-ai-2/



