Within AI Bias

Why Facial AI Failed Some People

The Gender Shades research exposed how facial analysis systems can fail unevenly across demographic groups.

60 sources 3 graphics
Preview for Why Facial AI Failed Some People

On this page

  • The study that exposed accuracy gaps
  • Why average accuracy hid unequal failures
  • What fairer testing requires

Introduction

The 2018 Gender Shades study became one of the clearest demonstrations that artificial intelligence systems can work differently for different groups of people. Researchers Joy Buolamwini and Timnit Gebru tested commercial facial analysis systems and found large accuracy gaps linked to both gender and skin tone. The poorest performance occurred for darker-skinned women, while lighter-skinned men were classified most accurately.[Proceedings of Machine Learning Research]proceedings.mlr.pressProceedings of Machine Learning ResearchGender Shades: Intersectional Accuracy Disparities in Commercial Gender ClassificationJanuary 21…

Face AI Gaps illustration 1

The importance of the study was not simply that some facial AI tools made mistakes. It showed that average accuracy figures could hide unequal failures affecting particular populations. For a future in which AI helps expand human capabilities — from healthcare and education to public services and scientific progress — this kind of evidence highlights a central challenge: advanced systems must be evaluated against the diversity of the people they serve.[Proceedings of Machine Learning Research]proceedings.mlr.pressProceedings of Machine Learning ResearchGender Shades: Intersectional Accuracy Disparities in Commercial Gender ClassificationJanuary 21…

The study that exposed accuracy gaps

Gender Shades was published in 2018 by Joy Buolamwini and Timnit Gebru in the Proceedings of Machine Learning Research. The researchers examined three commercial gender classification systems using a new evaluation dataset designed to test performance across combinations of gender and skin type. Their approach used the Fitzpatrick Skin Type scale, a dermatological classification system for skin tones, to examine how performance changed across groups.[Proceedings of Machine Learning Research]proceedings.mlr.pressProceedings of Machine Learning ResearchGender Shades: Intersectional Accuracy Disparities in Commercial Gender ClassificationJanuary 21…

The study introduced the idea of intersectional evaluation: testing people who belong to overlapping groups rather than analysing categories separately. A system might appear reasonably accurate when measured only by gender or only by skin tone, while still performing poorly for people affected by both factors.

The results revealed a striking pattern:

  • Darker-skinned women experienced the highest error rates, reaching up to 34.7% in the systems tested.
  • Lighter-skinned men experienced the lowest error rates, with a maximum of 0.8%.
  • The gap was not a small variation around an otherwise equal system; errors were concentrated among specific demographic groups.[Proceedings of Machine Learning Research]proceedings.mlr.pressProceedings of Machine Learning ResearchGender Shades: Intersectional Accuracy Disparities in Commercial Gender ClassificationJanuary 21…

The study also examined the datasets used to develop and evaluate facial AI. Earlier benchmark datasets were heavily weighted towards lighter-skinned subjects, creating a risk that systems would learn from a narrower picture of human appearance. The researchers found that two widely used datasets contained predominantly lighter-skinned subjects, with one dataset consisting of 79.6% lighter-skinned individuals and another 86.2%.[Proceedings of Machine Learning Research]proceedings.mlr.pressProceedings of Machine Learning ResearchGender Shades: Intersectional Accuracy Disparities in Commercial Gender ClassificationJanuary 21…

Why average accuracy hid unequal failures

A major lesson from Gender Shades was that a single performance number can give a misleading impression of fairness.

Imagine a facial analysis system that is 90% accurate overall. That figure may sound impressive, but it does not reveal whether the remaining errors are evenly distributed. If mistakes are concentrated among one group, the real-world experience of that group can be very different from the average user.

This matters because AI systems are increasingly used in settings where errors can have unequal consequences. Facial technologies have been considered for identity verification, security, policing, access control and other high-impact uses. The problem is not only whether a system makes mistakes, but who is most likely to experience those mistakes and what happens afterwards.

The Gender Shades findings helped shift discussion from “How accurate is the algorithm?” to a more demanding question: “Accurate for whom?” This change in evaluation standards became an important part of wider AI fairness research. Later work by the US National Institute of Standards and Technology (NIST) also found demographic differences across many face recognition algorithms, showing that unequal performance was not limited to one study or one type of system.[nist.gov]nist.govface recognition vendor test part 3demographic effectsFace recognition vendor test part 3::demographic effects | NISTDecember 1, 2019…Published: December 1, 2019

Face AI Gaps illustration 2

What caused the facial AI gaps?

Gender Shades did not claim that one single technical flaw explained all differences. Instead, it pointed to several connected issues involving data, testing and design choices.

Training data reflected limited representation

Machine learning systems learn patterns from examples. When the examples do not adequately represent the range of human variation, performance can become uneven.

The Gender Shades researchers found that existing benchmarks were not balanced across gender and skin type. A system trained or tuned using datasets that contain fewer examples of darker-skinned women may have fewer opportunities to learn the visual patterns needed for accurate classification.[Proceedings of Machine Learning Research]proceedings.mlr.pressProceedings of Machine Learning ResearchGender Shades: Intersectional Accuracy Disparities in Commercial Gender ClassificationJanuary 21…

However, later research has shown that the causes of performance gaps can be more complicated than simple dataset imbalance. IBM researchers investigating unequal gender classification accuracy found evidence that factors such as facial features, image characteristics and the way models interpret visual patterns can also contribute.[IBM Research]research.ibm.comIBM ResearchUnderstanding Unequal Gender Classification Accuracy from Face Images for arXiv - IBM ResearchNovember 30, 2018…Published: November 30, 2018

This matters because improving fairness requires more than adding more images. Developers need to understand where errors come from and test systems under realistic conditions.

What fairer testing requires

Gender Shades changed expectations around how AI systems should be evaluated. The study argued that testing should not rely only on overall accuracy but should report results across relevant demographic groups.[Proceedings of Machine Learning Research]proceedings.mlr.pressProceedings of Machine Learning ResearchGender Shades: Intersectional Accuracy Disparities in Commercial Gender ClassificationJanuary 21…

A stronger evaluation approach includes:

  • Balanced benchmarks: Test datasets should represent a broad range of people rather than relying on populations that are easier to collect.
  • Disaggregated results: Performance should be reported separately for different demographic groups instead of hiding differences inside one average score.
  • Independent audits: External researchers should be able to examine systems rather than relying only on company claims.
  • Clear limits on use: Some applications may require stronger evidence of reliability because mistakes can affect people’s opportunities, privacy or rights.

The study also demonstrated the value of external auditing. Research examining the impact of Gender Shades found that after the audit became public, several companies released updated systems and reduced some measured accuracy gaps on the benchmark used by the researchers. However, the researchers also emphasised that improved accuracy alone does not resolve broader questions about surveillance, privacy and appropriate use.[MIT Media Lab]media.mit.eduMIT Media LabOverview ‹ Actionable Auditing: Coordinated bias disclosure study — MIT Media Lab…

Face AI Gaps illustration 3

Why this matters for the future of AI

Gender Shades is a small but important case study in a much larger question about AI and human flourishing. More capable AI systems could eventually support major advances in medicine, science, education and productivity, but those benefits depend on systems being reliable for diverse populations.

A future of AI-enabled abundance would require more than powerful algorithms. It would require trustworthy systems that people can use without hidden groups carrying higher risks of failure. Facial AI showed that technical progress can move quickly while evaluation practices lag behind.

The lasting lesson of Gender Shades is not that AI cannot improve human life. It is that better futures require better measurements. If society wants advanced AI to expand opportunity rather than deepen existing inequalities, fairness testing must become part of building the technology itself rather than a correction applied afterwards.

Amazon book picks

Further Reading

Books and field guides related to Why Facial AI Failed Some People. Use these as the next step if you want deeper reading beyond the article.

BookCover for Invisible Women

Invisible Women

By Caroline Criado Perez

Winner of the 2019 Royal Society Science Book Prize Shortlisted for the 2019 Financial Times and McKinsey Business Book of the Year Award...

BookCover for Race After Technology

Race After Technology

By Ruha Benjamin

From everyday apps to complex algorithms, Ruha Benjamin cuts through tech-industry hype to understand how emerging technologies can reinf...

eBay marketplace picks

Marketplace Samples

Live-tested eBay searches with available results related to this page.

UsingUSA

Selected fromfacial recognition poster oneBay.co.uk.

Endnotes

1. Source: nist.gov
Title: face recognition vendor test part 3demographic effects
Link:https://www.nist.gov/publications/face-recognition-vendor-test-part-3demographic-effects

Source snippet

Face recognition vendor test part 3::demographic effects | NISTDecember 1, 2019...

Published: December 1, 2019

2. Source: nist.gov
Title: Face Projects | NIST
Link:https://www.nist.gov/programs-projects/face-projects

Source snippet

March 26, 2025 — FACE PROJECTS DESCRIPTION FACE TECHNOLOGY EVALUATIONS - FRTE/FATE To bring clarity to our testing scope and goals, what...

Published: March 26, 2025

3. Source: research.ibm.com
Link:https://research.ibm.com/publications/understanding-unequal-gender-classification-accuracy-from-face-images

Source snippet

IBM ResearchUnderstanding Unequal Gender Classification Accuracy from Face Images for arXiv - IBM ResearchNovember 30, 2018...

Published: November 30, 2018

4. Source: media.mit.edu
Link:https://www.media.mit.edu/projects/actionable-auditing-coordinated-bias-disclosure-study/overview/

Source snippet

MIT Media LabOverview ‹ Actionable Auditing: Coordinated bias disclosure study — MIT Media Lab...

5. Source: pages.nist.gov
Link:https://pages.nist.gov/frvt/html/frvt11.html

Source snippet

Recognition Technology Evaluation (FRTE) 1:1 VerificationJune 15, 2026 — FRTE 1:1 DEMOGRAPHIC DIFFERENTIALS SUMMARY This page summarizes...

Published: June 15, 2026

6. Source: nist.gov
Title: face recognition vendor test frvt
Link:https://www.nist.gov/programs-projects/face-recognition-vendor-test-frvt

7. Source: pages.nist.gov
Title: frvt demographics
Link:https://pages.nist.gov/frvt/html/frvt_demographics.html

8. Source: nist.gov
Title: facial recognition technology part iii ensuring commercial transparency accuracy
Link:https://www.nist.gov/speech-testimony/facial-recognition-technology-part-iii-ensuring-commercial-transparency-accuracy

9. Source: research.ibm.com
Link:https://research.ibm.com/publications/color-theoretic-experiments-to-understand-unequal-gender-classification-accuracy-from-face-images

10. Source: media.mit.edu
Title: full gender shades thesis 17
Link:https://www.media.mit.edu/publications/full-gender-shades-thesis-17/

11. Source: nist.gov
Title: demographic effects estimates automatic face recognition performance 0
Link:https://www.nist.gov/publications/demographic-effects-estimates-automatic-face-recognition-performance-0

12. Source: nist.gov
Title: demographic effects estimates automatic face recognition performance 1
Link:https://www.nist.gov/publications/demographic-effects-estimates-automatic-face-recognition-performance-1

13. Source: nist.gov
Title: demographic effects estimates automatic face recognition performance 2
Link:https://www.nist.gov/publications/demographic-effects-estimates-automatic-face-recognition-performance-2

14. Source: ibm.com
Link:https://www.ibm.com/docs/en/watsonx/w-and-w/2.2.0?topic=atlas-decision-bias

15. Source: betterworld.mit.edu
Title: algorithmic justice
Link:https://betterworld.mit.edu/algorithmic-justice/

16. Source: proceedings.mlr.press
Link:https://proceedings.mlr.press/v81/buolamwini18a.html

Source snippet

Proceedings of Machine Learning ResearchGender Shades: Intersectional Accuracy Disparities in Commercial Gender ClassificationJanuary 21...

17. Source: dataethicsrepository.iaa.ncsu.edu
Title: gender shades
Link:https://dataethicsrepository.iaa.ncsu.edu/2023/04/27/gender-shades/

18. Source: govinfo.gov
Title: Face Recognition Vendor Test Part 3: Demographic Effects
Link:https://www.govinfo.gov/app/details/GOVPUB-C13-5a5c1d28d99c5a718d50bdbbae85cbb9

19. Source: proceedings.mlr.press
Link:https://proceedings.mlr.press/v81/buolamwini18a

20. Source: proceedings.mlr.press
Link:https://proceedings.mlr.press/v81/

21. Source: genderedinnovations.stanford.edu
Link:https://genderedinnovations.stanford.edu/case-studies/facial.html

Additional References

22. Source: frontiersin.org
Title: Seminal work by Buolamwini
Link:https://www.frontiersin.org/journals/big-data/articles/10.3389/fdata.2025.1686452/full

Source snippet

Frontiers | Bias in AI systems: integrating formal and socio-technical approachesJanuary 8, 2026 — 3 EXAMPLES OF BIAS IN AI 3.1 FACIAL RE...

Published: January 8, 2026

23. Source: youtube.com
Title: Dissecting Racial Bias in an Algorithm that Guides [Health]({{ ‘health/’ | relative_url }}) Decisions for Millions
Link:https://www.youtube.com/watch?v=y6eo0FZIqjk

Source snippet

Ziad Obermeyer on tackling 'data bottleneck' to combat [AI bias]({{ 'ai-bias/' | relative_url }})...

24. Source: youtube.com
Title: Dissecting Algorithmic Bias | Ziad Obermeyer | AI FOR GOOD DISCOVERY
Link:https://www.youtube.com/watch?v=U5MlyFsMi-E

Source snippet

Dissecting Racial Bias in an Algorithm that Guides Health Decisions for Millions...

25. Source: pmc.ncbi.nlm.nih.gov
Link:https://pmc.ncbi.nlm.nih.gov/articles/PMC12823528/

Source snippet

EXAMPLES OF BIAS IN AI 3.1. FACIAL RECOGNITION SYSTEMS Facial recognition technologies remain among the most visible and critically exami...

26. Source: youtube.com
Title: Ziad Obermeyer on tackling ‘data bottleneck’ to combat AI bias
Link:https://www.youtube.com/watch?v=xWLGch3g3mA

Source snippet

How AI Bias Impacts Healthcare | Invisible Inputs Lesson (K–12)...

27. Source: congress.gov
Link:https://www.congress.gov/crs-product/R46586

28. Source: youtube.com
Title: Keynote Presentation: Dissecting Algorithmic Bias
Link:https://www.youtube.com/watch?v=JfKYO1W4uuA

Source snippet

Dissecting Algorithmic Bias | Ziad Obermeyer | AI FOR GOOD DISCOVERY...

29. Source: gendershades.org
Link:https://gendershades.org/overview.html

30. Source: dblp.org
Link:https://dblp.org/rec/conf/fat/BuolamwiniG18

31. Source: ellphacitizen.org
Link:https://www.ellphacitizen.org/academic-research/2018/2/12/gender-shades-intersectional-accuracy-disparities-in-commercial-gender-classification