# Timnit Gebru

> 1983– · Computer Scientist, AI Ethics Researcher
>
> **Recorded contribution:** AI ethics; dataset bias; co-authored "Stochastic Parrots"; founded DAIR Institute

## How to use this dossier

Read for a causal chain, not a hero story: inherited problem → contribution → mechanism → downstream capability → limit. Then close the page and complete the reconstruction exercise from memory.

## 1. Historical orientation

Eritrean-Ethiopian computer scientist Timnit Gebru researched computer vision, dataset documentation, and the social impacts of AI at Stanford, Microsoft Research, and Google. She co-founded Black in AI and later founded the Distributed AI Research Institute. Her co-authored work on Gender Shades, datasheets for datasets, and 'Stochastic Parrots' made data provenance, demographic performance gaps, environmental cost, and institutional power central technical questions. Gebru's work joins technical measurement to institutional accountability. Gender Shades made performance disparities empirically visible; Datasheets for Datasets then proposed documenting why a dataset exists, how it was collected, what populations it represents, and which uses its creators foresee or reject.

## 2. The problem inherited

AI systems were celebrated through average benchmark scores while hiding subgroup failures, undocumented datasets, extractive data practices, labor, scale costs, and institutional conflicts over publication.

## 3. The central contribution

Gebru advanced empirical audits and documentation practices that expose demographic disparities and the social and material conditions under which machine-learning systems are built.

## 4. Reconstruct the mechanism

1. Define the claimed task and identify populations for whom errors have different prevalence or consequences.
2. Construct an evaluation set with documented demographic labeling and limitations.
3. Measure error separately across intersectional groups instead of reporting only an average.
4. Trace disparities back through data, objectives, deployment context, governance, and available recourse.

## 5. What changed downstream

- Gender Shades demonstrated substantial intersectional performance differences in commercial gender classification.
- Datasheets and later critical work influenced model and dataset documentation, responsible-AI practice, and public scrutiny of research governance.
- Her research helped make dataset documentation, disaggregated evaluation, and the labor and environmental costs of large models normal subjects of technical review rather than external public-relations questions.

## 6. Attribution, limits, and uncertainty

- Gender Shades and Stochastic Parrots are multi-author works; labels for gender and skin type are limited operationalizations, not complete human identities.
- Audits reveal disparities but do not alone determine whether a system should exist, how categories should be defined, or which remedy affected communities prefer.
- Documentation has little force if organizations can ignore it, affected communities lack decision power, or audit access is restricted; demographic categories themselves can erase people or reproduce imposed classifications.

## 7. Reconstruction lab

Audit a classifier on an intersectional test matrix. Report sample counts and uncertainty, identify the worst-supported claim, and propose a non-model alternative alongside technical mitigation. Invite someone represented in the data to challenge the categories and permitted uses, then record which design decisions change rather than treating consultation as approval.

## 8. Evidence trail

- [Gender Shades](http://gendershades.org/) — MIT Media Lab
- [Datasheets for Datasets](https://arxiv.org/abs/1803.09010) — Communications of the ACM
- [On the Dangers of Stochastic Parrots](https://dl.acm.org/doi/10.1145/3442188.3445922) — ACM FAccT

---

*Research checked 2026-08-09. Dates, roles, and claims about living people are historical snapshots. Linked sources remain the authority; this dossier is original instructional synthesis.*
