Gender Shades
Joy Buolamwini and Timnit Gebru's audit of three commercial face-analysis products, which misclassified darker-skinned women up to 34.7 percent of the time and lighter-skinned men at most 0.8 percent.
What it was
Gender Shades was a paper by Joy Buolamwini, then a graduate student at the MIT Media Lab, and Timnit Gebru, who had done the work as a Stanford graduate student and was by then a postdoc at Microsoft Research. MIT News announced it on February 11, 2018, and the authors presented it at the first Conference on Fairness, Accountability and Transparency (FAT*) in New York on February 23-24, 2018.
The authors found that two widely used face benchmarks, IJB-A and Adience, were 79.6 and 86.2 percent lighter-skinned faces. They built a new benchmark of 1,270 parliamentarians from three African and three European countries, balanced by gender and by skin type on the dermatologists' Fitzpatrick scale, and ran it through the gender classifiers sold by Microsoft, IBM and Face++. All three did worst on darker-skinned women, with error rates of up to 34.7 percent, against a maximum of 0.8 percent for lighter-skinned men.
What it changed
The companies responded within months. IBM's Watson chief architect Ruchir Puri told MIT News in February 2018 that IBM had a new model "much more balanced in terms of accuracy across the benchmark that Joy was looking at." A January 2019 follow-up by Inioluwa Deborah Raji and Buolamwini, "Actionable Auditing," found that all three audited companies had released new versions within seven months and had cut error on darker-skinned women by 17.7 to 30.4 percent, and it tested Amazon and Kairos, which had not been named in the first audit.
In June 2020, during the protests after George Floyd's killing, IBM chief executive Arvind Krishna wrote to Congress on June 8 that IBM would no longer offer general-purpose facial recognition, and Amazon announced a one-year moratorium on police use of its Rekognition product two days later. The Verge's report on IBM's letter cited Buolamwini and Gebru's 2018 work as the research that "revealed for the first time the extent to which many commercial facial recognition systems (including IBM's) were biased."
The arguments it moved
What are the biggest risks from AI?
Gender Shades gave the argument that AI's main harms are present and measurable a concrete number and a method that others could repeat. Buolamwini told MIT News in February 2018 that "our benchmarks, the standards by which we measure success, themselves can give us a false sense of progress." The paper's method, publicly naming vendors and retesting them, became its own subject of study in Raji and Buolamwini's January 2019 "Actionable Auditing." Gebru has since set this kind of documented harm against arguments about future catastrophe, which she opposes in her work on longtermism.
Who should set the rules for AI?
The audit became evidence in the argument over whether companies or governments should restrict facial recognition. IBM's June 8, 2020 letter to Senators Cory Booker and Kamala Harris and three House members said IBM "firmly opposes and will not condone uses of any [facial recognition] technology" for mass surveillance or racial profiling and called for "a national dialogue" on police use. Amazon's June 10, 2020 moratorium statement said it hoped the pause "might give Congress enough time to implement appropriate rules."
Positions it bears on
-
Timnit Gebru, Facial recognition and surveillance
Facial analysis systems fail most on the people most exposed to surveillance, and she has argued since 2019 that police should not buy them.
Her stance on facial recognition rests on this audit, and her profile quotes its headline result.
-
Timnit Gebru, Bias is not only a data problem
When Yann LeCun said in June 2020 that machine learning systems are biased because their data is biased, Gebru objected that harms cannot be reduced to dataset composition; they also come from problem framing, deployment and who is in the room.
The paper measured dataset skew, and her June 2020 argument with Yann LeCun was over whether such skew is the whole explanation for biased systems.
Sources
- Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification (PMLR 81, FAT* 2018)
- Study finds gender and skin-type bias in commercial artificial-intelligence systems (MIT News, February 11, 2018)
- Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products (Raji and Buolamwini, AIES 2019)
- IBM will no longer offer, develop, or research facial recognition technology (The Verge, June 8, 2020)
- We are implementing a one-year moratorium on police use of Rekognition (Amazon, June 10, 2020)