Within Human Feedback

Why Humans Disagree About AI Answers

Human evaluators often disagree about quality, safety, and usefulness, revealing how difficult it is to capture shared values in feedback data.

74 sources 3 graphics
Preview for Why Humans Disagree About AI Answers

On this page

  • Where evaluator judgements diverge
  • What feedback data leaves unclear
  • Implications for future alignment

Introduction

Human feedback has become one of the main ways researchers try to align AI systems with human intentions. The evidence so far is mixed: feedback-based training can clearly make AI assistants more useful, less toxic, and better at following instructions, but the data used to guide these systems contains a deeper problem — humans often disagree about what the best answer actually is.[OpenAI]OpenAIinstruction followingAligning language models to follow instructions | OpenAIJanuary 27, 2022…Published: January 27, 2022

Preference Gaps illustration 1

These disagreements are not just annotation mistakes. They reveal a harder alignment question: when people provide feedback, are they expressing a shared human value, a personal preference, a temporary judgement, or a compromise between competing goals? For AI systems intended to support a future of broad human flourishing, this distinction matters. A system that learns only the majority preference may miss minority concerns, reward persuasive style over truth, or mistake agreement with users for genuine helpfulness.

The study of preference disagreement provides important evidence about both the strengths and limits of current alignment methods. It shows that human feedback is a valuable steering mechanism, but not a complete definition of what humans ultimately want from increasingly capable AI.

Where evaluator judgements diverge

Human feedback training usually works by asking people to compare possible AI outputs. An evaluator might choose which answer is more helpful, safer, clearer, or more truthful. These comparisons are powerful because many qualities that matter to humans are difficult to specify as simple rules. However, the resulting data reflects the judgements of particular people in particular situations, not a single universal standard of quality.[OpenAI]OpenAIinstruction followingAligning language models to follow instructions | OpenAIJanuary 27, 2022…Published: January 27, 2022

Disagreement appears for several reasons:

  • Different goals can conflict. One evaluator may prefer a concise answer, while another values detail. One may prioritise politeness, while another prefers direct criticism.
  • Some questions have no single correct style. Advice, explanations, and creative writing often involve legitimate differences in taste.
  • Safety judgements can depend on context. A refusal may appear cautious to one person and unhelpful to another.
  • People may disagree about deeper values. Different communities may have different views about fairness, acceptable risk, political neutrality, or the balance between individual freedom and collective welfare.

Research examining disagreements in human preference datasets has found that many disagreements are structured rather than random. A 2024 study on diverging preferences identified categories including unclear tasks, response style differences, refusal behaviour, and annotation errors. It argued that many disagreements are treated as “noise” by standard reward-modelling approaches even though they can contain meaningful information about different human priorities.[arXiv]arxiv.orgarXiv Diverging Preferences: When do Annotators Disagree and do Models Know?Diverging Preferences: When do Annotators Disagree and do Models Know?October 18, 2024…Published: October 18, 2024

This matters because AI alignment is not only a technical problem of predicting labels. It is also a social problem of deciding whose judgement counts, how conflicting views should be represented, and when disagreement signals uncertainty rather than failure.

14:46

What feedback data leaves unclear

Majority preference is not the same as human values

A common assumption in reinforcement learning from human feedback (RLHF) is that enough examples of human choices can approximate what people want. This works reasonably well for many everyday tasks. OpenAI’s InstructGPT research showed that models trained with human feedback were preferred over much larger unaligned models on many user requests and improved in areas such as instruction following and truthfulness.[OpenAI]OpenAIinstruction followingAligning language models to follow instructions | OpenAIJanuary 27, 2022…Published: January 27, 2022

However, preference data has limits. A person selecting one answer over another may not be expressing a deep value judgement. They may simply prefer a particular writing style, a familiar argument, or an answer that sounds confident.

This creates a potential gap between:

  • what humans immediately prefer, and
  • what would actually serve human interests over time.

For example, users may prefer an assistant that agrees with them, but a more valuable system might sometimes challenge mistaken assumptions. Similarly, a confident but incorrect answer may initially appear more satisfying than a careful response that admits uncertainty.

These cases show why alignment researchers distinguish between surface preferences and deeper goals such as truthfulness, autonomy, fairness, and human flourishing.

Preference Gaps illustration 2

Disagreement can contain useful information

Earlier alignment approaches often treated disagreement between annotators as a problem to reduce. If two people disagreed, the system might simply average their responses or choose the majority label. Newer research suggests disagreement itself can be informative.

A disagreement may indicate:

  • the question is ambiguous;
  • reasonable people are applying different values;
  • the model is operating in a morally sensitive area;
  • more context is needed before making a decision.

The challenge is that many current reward models compress these differences into a single score. A model may learn that one answer is “preferred” without understanding whether the preference was unanimous or deeply contested. Research on diverging preferences found that standard reward-modelling techniques can fail to distinguish between broad agreement and divided opinions.[arXiv]arxiv.orgarXiv Diverging Preferences: When do Annotators Disagree and do Models Know?Diverging Preferences: When do Annotators Disagree and do Models Know?October 18, 2024…Published: October 18, 2024

For future AI systems, preserving uncertainty may become as important as learning preferences. An AI assistant helping with ordinary writing can usually optimise for usefulness, but an AI system influencing scientific decisions, governance, or resource allocation may need to recognise when humanity itself lacks agreement.

Evidence from real-world preference evaluations

Human disagreement is visible not only in training datasets but also in attempts to measure AI quality.

One example is Chatbot Arena, a large-scale evaluation platform where users compare responses from different language models. Instead of relying only on fixed tests, it collects pairwise human preferences from real interactions. The project has gathered hundreds of thousands of votes and found that crowdsourced evaluations can produce useful rankings, with agreement between crowd judgements and expert ratings in their analysis.[sky.cs.berkeley.edu]sky.cs.berkeley.eduChatbot Arena – UC Berkeley Sky Computing LabApril 25, 2024…Published: April 25, 2024

However, the success of these evaluations also highlights their limits. A ranking tells researchers which system people prefer in a particular comparison, but it does not fully explain why. A model may win because it is clearer, more entertaining, more agreeable, or more aligned with the expectations of the evaluator population.

The development of Arena-Hard, a benchmark built from real user interactions, reflects this challenge. Researchers emphasised that good evaluation requires both separating model capabilities and matching real human preferences. Their analysis reported strong agreement between the benchmark and Chatbot Arena rankings, but the need for such methods shows that measuring preference itself is an ongoing research problem.[sky.cs.berkeley.edu]sky.cs.berkeley.eduArena Hard – UC Berkeley Sky Computing LabApril 26, 2024…Published: April 26, 2024

The challenge of pluralistic alignment

The deeper issue is that humanity does not have one perfectly consistent set of preferences. A globally deployed AI system will interact with people who differ in culture, language, values, experiences, and political beliefs.

This creates a question beyond ordinary RLHF: should AI systems aim to satisfy average preferences, follow democratic processes, represent specific users, or operate according to broader principles?

Research on social choice and AI alignment argues that methods from political philosophy and collective decision-making may help address how diverse human feedback should be combined. The problem resembles many existing human governance challenges: societies already need ways to make decisions when individuals disagree about what is fair or desirable.[arXiv]arxiv.orgSocial Choice Should Guide AI Alignment in Dealing with Diverse Human FeedbackApril 16, 2024…Published: April 16, 2024

Anthropic’s experiments with Collective Constitutional AI explored a related question by gathering public input on principles for AI behaviour. The project found areas of agreement as well as disagreement among participants, illustrating that even when people are asked to define general principles rather than rank individual answers, differences remain.[anthropic.com]anthropic.comOctober 17, 2023…Published: October 17, 2023

This suggests that future alignment may require systems that do more than imitate preferences. They may need mechanisms for handling disagreement transparently, explaining trade-offs, and allowing legitimate differences between users and communities.

Preference Gaps illustration 3

Implications for a future AI bloom

If advanced AI contributes to a period of abundance, scientific acceleration, better healthcare, and expanded human capability, alignment will require more than making systems pleasant or responsive. The central question is whether AI systems can reliably support human goals when those goals are complex, diverse, and sometimes conflicting.

Human preference disagreements provide evidence for both optimism and caution.

The optimistic lesson is that human feedback works. It has already improved AI assistants and demonstrated that people can provide useful guidance for shaping model behaviour.[OpenAI]OpenAIinstruction followingAligning language models to follow instructions | OpenAIJanuary 27, 2022…Published: January 27, 2022

The caution is that feedback is not a direct measurement of human flourishing. It is a limited window into human judgement. As AI systems become more powerful, the gap between a simple preference signal and the broader question of what humanity should want may become increasingly important.

A flourishing future with advanced AI may therefore depend on moving from preference imitation towards richer forms of alignment: systems that understand reasons behind feedback, represent disagreement honestly, protect minority interests, and remain accountable to the diverse humans they are meant to serve.

Amazon book picks

Further Reading

Books and field guides related to Why Humans Disagree About AI Answers. Use these as the next step if you want deeper reading beyond the article.

BookCover for The Alignment Problem

The Alignment Problem

By Brian Christian

Finalist for the Los Angeles Times Book Prize A jaw-dropping exploration of everything that goes wrong when we build AI systems and the m...

BookCover for Artificial Intelligence

Artificial Intelligence

By Stuart Jonathan Russell, Peter Norvig et al.

Rating: 4.5/5 from 10 Google Books ratings

Artificial intelligence: A Modern Approach, 3e,is ideal for one or two-semester, undergraduate or graduate-level courses in Artificial In...

BookCover for The Righteous Mind

The Righteous Mind

By Jonathan Haidt

'A landmark contribution to humanity's understanding of itself' The New York Times Why can it sometimes feel as though half the populatio...

eBay marketplace picks

Marketplace Samples

Live-tested eBay searches with available results related to this page.

UsingUSA

Selected fromrobotics poster oneBay.co.uk.

Endnotes

1. Source: OpenAI
Title: instruction following
Link:https://openai.com/index/instruction-following/

Source snippet

Aligning language models to follow instructions | OpenAIJanuary 27, 2022...

Published: January 27, 2022

2. Source: arxiv.org
Title: arXiv Training language models to follow instructions with human feedback
Link:https://arxiv.org/abs/2203.02155

3. Source: arxiv.org
Title: arXiv Diverging Preferences: When do Annotators Disagree and do Models Know?
Link:https://arxiv.org/abs/2410.14632

Source snippet

Diverging Preferences: When do Annotators Disagree and do Models Know?October 18, 2024...

Published: October 18, 2024

4. Source: sky.cs.berkeley.edu
Title: Chatbot Arena – UC Berkeley Sky Computing Lab
Link:https://sky.cs.berkeley.edu/project/chatbot-arena/

Source snippet

April 25, 2024...

Published: April 25, 2024

5. Source: arxiv.org
Title: arXiv Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Link:https://arxiv.org/abs/2403.04132

6. Source: sky.cs.berkeley.edu
Title: Arena Hard – UC Berkeley Sky Computing Lab
Link:https://sky.cs.berkeley.edu/project/arena-hard/

Source snippet

April 26, 2024...

Published: April 26, 2024

7. Source: arena.ai
Title: AIThe Arena-Hard Pipeline
Link:https://arena.ai/blog/arena-hard/

Source snippet

The Arena-Hard Pipeline...

8. Source: arxiv.org
Link:https://arxiv.org/abs/2404.10271

Source snippet

Social Choice Should Guide AI Alignment in Dealing with Diverse Human FeedbackApril 16, 2024...

Published: April 16, 2024

9. Source: anthropic.com
Link:https://www.anthropic.com/research/collective-constitutional-ai-aligning-a-language-model-with-public-input

Source snippet

October 17, 2023...

Published: October 17, 2023

10. Source: OpenAI
Title: row terms of use
Link:https://openai.com/es-US/policies/row-terms-of-use/

11. Source: OpenAI
Title: collective alignment aug 2025 updates
Link:https://openai.com/index/collective-alignment-aug-2025-updates/

12. Source: arena.ai
Title: Preference Proxy Evaluations
Link:https://arena.ai/blog/preference-proxy-evaluations

14. Source: arena.ai
Title: Chatbot Arena Conversation Dataset Release
Link:https://arena.ai/blog/chatbot-arena-conversation-dataset-release/

16. Source: anthropic.com
Title: Specific versus General Principles for Constitutional AI \ Anthropic
Link:https://www.anthropic.com/news/specific-versus-general-principles-for-constitutional-ai

17. Source: arena.ai
Title: Chatbot Arena Conversation Dataset Release
Link:https://arena.ai/blog/dataset

18. Source: anthropic.com
Title: Constitutional AI: Harmlessness from AI Feedback \ Anthropic
Link:https://www.anthropic.com/news/constitutional-ai-harmlessness-from-ai-feedback

19. Source: OpenAI
Title: our approach to alignment research
Link:https://openai.com/index/our-approach-to-alignment-research/

20. Source: OpenAI
Title: instruction following
Link:https://openai.com/index/instruction-following/?_hsmi=202743306

21. Source: OpenAI
Title: instruction following
Link:https://openai.com/index/instruction-following/?trk=article-ssr-frontend-pulse_x-social-details_comments-action_comment-text

22. Source: OpenAI
Title: instruction following
Link:https://openai.com/de-DE/index/instruction-following/

23. Source: help.openai.com
Title: 5722486 api data usage policies
Link:https://help.openai.com/en/articles/5722486-api-data-usage-policies

24. Source: help.openai.com
Title: 10306912 sharing feedback and api inputs and outputs with openai
Link:https://help.openai.com/en/articles/10306912-sharing-feedback-and-api-inputs-and-outputs-with-openai

25. Source: youtube.com
Title: Anthropic-AI Alignment: The Truth About Autonomous Behavior
Link:https://www.youtube.com/watch?v=K-oJytqi8AE

Source snippet

Reinforcement Learning From Human Feedback (RLHF) | Direct Preference Optimization (DPO) | Explained...

26. Source: youtube.com
Link:https://www.youtube.com/watch?v=RBWVB9ubXqM

Source snippet

Reinforcement Learning from Human Feedback explained with math derivations and the PyTorch code...

27. Source: deepwiki.com
Title: openai/following-instructions-human-feedback | Deep Wiki
Link:https://deepwiki.com/openai/following-instructions-human-feedback/1-overview

Additional References

28. Source: sciencedirect.com
Link:https://www.sciencedirect.com/science/article/pii/S0747563226000853?dgcid=rss_sd_all

Source snippet

Today — COMPUTERS IN HUMAN BEHAVIOR Volume 181, August 2026, 108988 Full length article Rubric-conditioned large language mo...

Published: August 2026

29. Source: nature.com
Link:https://www.nature.com/articles/s41586-026-10303-2

30. Source: youtube.com
Title: How to Align AI: Put It in a Sandwich
Link:https://www.youtube.com/watch?v=5mco9zAamRk

Source snippet

Aengus Lynch - AI evaluators give wrong labels when they disagree with the consequences...

31. Source: youtube.com
Link:https://www.youtube.com/watch?v=qGyFrqc34yc

Source snippet

How to Align AI: Put It in a Sandwich...

32. Source: researchgate.net
Link:https://www.researchgate.net/publication/382559377_Reinforcement_Learning_from_Human_Feedback_Whose_Culture_Whose_Values_Whose_Perspectives

33. Source: aclanthology.org
Link:https://aclanthology.org/2021.conll-1.18/

34. Source: aclanthology.org
Link:https://aclanthology.org/2026.findings-acl.1342/

35. Source: alphaxiv.org
Link:https://www.alphaxiv.org/abs/2506.19467

36. Source: zeroentropy.dev
Link:https://www.zeroentropy.dev/concepts/constitutional-ai/

37. Source: paperswithcode.com
Link:https://paperswithcode.com/paper/unlocking-transparent-alignment-through