# Yejin Choi

> ~1980– · Computer Scientist, NLP Researcher
>
> **Recorded contribution:** Common sense AI; MacArthur Fellow; UW NLP; moral reasoning in language models

## How to use this dossier

Read for a causal chain, not a hero story: inherited problem → contribution → mechanism → downstream capability → limit. Then close the page and complete the reconstruction exercise from memory.

## 1. Historical orientation

South Korean-born computer scientist Yejin Choi is a University of Washington professor and senior research director at the Allen Institute for AI. Her work investigates commonsense reasoning, social and moral dimensions of language, neuro-symbolic methods, and the limitations of large language models. Projects such as ATOMIC and Delphi made implicit social knowledge and its hazards explicit research objects.

## 2. The problem inherited

Language models can reproduce surface patterns while missing unstated causes, intentions, norms, and consequences that people rely on when interpreting everyday situations.

## 3. The central contribution

Choi and collaborators built datasets and models for commonsense and social reasoning while foregrounding the brittleness and ethical risk of treating learned judgments as moral authority.

## 4. Reconstruct the mechanism

1. Collect structured examples linking events to intents, reactions, causes, or likely effects.
2. Train models to retrieve or generate missing commonsense relations.
3. Evaluate on held-out situations and adversarially rephrased or culturally varied cases.
4. Audit whose norms and assumptions appear in labels before deploying any judgment.

## 5. What changed downstream

- Commonsense resources became useful training and evaluation infrastructure for NLP.
- Her work helped make social bias, moral pluralism, and model limitations part of technical evaluation rather than downstream disclaimers.

## 6. Attribution, limits, and uncertainty

- Datasets aggregate annotator judgments and cannot encode one universal commonsense or morality.
- A model predicting majority judgments is descriptive, not legitimate authority for consequential decisions, and may marginalize minority or contextual views.

## 7. Reconstruction lab

Collect ten everyday scenarios from three annotators and record intent and likely consequence. Measure disagreement, train a simple retrieval baseline, and refuse to collapse contested cases into one moral label. Change the scenario wording while preserving the underlying action and test whether predictions remain stable. Separate descriptive judgments about likely consequences from normative judgments about what ought to happen. Add a minority interpretation and retain it rather than averaging it away. The exercise exposes the core difficulty of computational commonsense: datasets encode whose expectations were solicited, models exploit linguistic shortcuts, and aggregate accuracy can conceal precisely the contested cases where automated moral authority would be most harmful. Publish annotator composition and uncertainty, allowing downstream users to see that disagreement is data rather than cleanup residue.

## 8. Evidence trail

- [Yejin Choi](https://homes.cs.washington.edu/~yejin/) — University of Washington
- [ATOMIC: An Atlas of Machine Commonsense](https://arxiv.org/abs/1811.00146) — AAAI
- [Delphi: Towards Machine Ethics and Norms](https://arxiv.org/abs/2110.07574) — arXiv

---

*Research checked 2026-08-09. Dates, roles, and claims about living people are historical snapshots. Linked sources remain the authority; this dossier is original instructional synthesis.*
