# Percy Liang

> ~1983– · Computer Scientist, AI Researcher
>
> **Recorded contribution:** HELM benchmarks; Stanford CRFM; foundation model evaluation; semantic parsing

## How to use this dossier

Read for a causal chain, not a hero story: inherited problem → contribution → mechanism → downstream capability → limit. Then close the page and complete the reconstruction exercise from memory.

## 1. Historical orientation

Stanford computer scientist Percy Liang works on machine learning, natural-language understanding, and foundation-model evaluation. He founded the Stanford Center for Research on Foundation Models and led the multi-institution HELM project, which evaluates language models across many scenarios and metrics rather than reducing performance to one leaderboard score. Earlier work includes semantic parsing and learning with latent structures. Liang's research repeatedly asks where an AI system's apparent capability actually comes from: model structure, data, prompts, human feedback, or the evaluation itself. That systems view culminated institutionally in Stanford's Center for Research on Foundation Models and technically in HELM's scenario-based evaluation framework.

## 2. The problem inherited

Foundation models were being deployed across tasks while evaluations used inconsistent prompts, narrow datasets, opaque model versions, and single aggregate metrics.

## 3. The central contribution

Liang and collaborators created HELM as a transparent, scenario-based framework for holistic model evaluation and reproducible comparison.

## 4. Reconstruct the mechanism

1. Define a scenario as a task, domain, language, population, and adaptation protocol.
2. Run versioned models with recorded prompts and decoding parameters.
3. Measure accuracy alongside calibration, robustness, fairness, bias, toxicity, efficiency, and other relevant dimensions.
4. Publish raw predictions and metadata so disagreements and trade-offs remain inspectable.

## 5. What changed downstream

- HELM helped establish multidimensional, transparent evaluation as infrastructure for foundation-model research.
- The work strengthened demands for model documentation, reproducibility, and disclosure of evaluation conditions.
- HELM helped normalize multidimensional model reporting—accuracy alongside calibration, robustness, fairness, efficiency, and other concerns—rather than compressing deployment readiness into one leaderboard number.

## 6. Attribution, limits, and uncertainty

- HELM is a large collaborative project and cannot cover every real deployment or newly emerging capability.
- Benchmark transparency does not remove contamination, metric gaming, unmeasured harms, or the need for domain experts and affected users.
- No benchmark can enumerate every population, language, task, or future use; publishing a matrix of scores can create false completeness when scenarios, prompts, and metrics embody contested choices.

## 7. Reconstruction lab

Evaluate two small models on one task using accuracy, calibration, latency, subgroup performance, and refusal behavior. Write the decision that changes depending on metric priority. Write a model card from the resulting matrix and identify one deployment decision the benchmark cannot justify without stakeholder evidence.

## 8. Evidence trail

- [Holistic Evaluation of Language Models](https://arxiv.org/abs/2211.09110) — Stanford CRFM
- [Percy Liang](https://cs.stanford.edu/~pliang/) — Stanford University
- [HELM](https://crfm.stanford.edu/helm/) — Stanford CRFM

---

*Research checked 2026-08-09. Dates, roles, and claims about living people are historical snapshots. Linked sources remain the authority; this dossier is original instructional synthesis.*
