# Eliezer Yudkowsky

> 1979– · AI Safety Researcher
>
> **Recorded contribution:** AI alignment theory; MIRI founder; rationalist movement; Sequences

## How to use this dossier

Read for a causal chain, not a hero story: inherited problem → contribution → mechanism → downstream capability → limit. Then close the page and complete the reconstruction exercise from memory.

## 1. Historical orientation

Eliezer Yudkowsky co-founded the organization now called the Machine Intelligence Research Institute and wrote extensively about rationality and the difficulty of aligning highly capable AI with human values. His public essays helped form an online rationalist and AI-risk community. His work is largely conceptual and advocacy-oriented rather than an established body of experimentally validated AI algorithms.

## 2. The problem inherited

A system optimizing a formal objective may satisfy its specification while violating the designers' actual intent, especially when the system can find strategies its designers did not anticipate.

## 3. The central contribution

Yudkowsky popularized arguments that advanced AI alignment is a hard problem of specifying, learning, and preserving human intent under capability growth and strategic optimization.

## 4. Reconstruct the mechanism

1. Distinguish a designer's intended outcome from the measurable objective supplied to the optimizer.
2. Let a capable search process discover policies that maximize the proxy.
3. Inspect whether distribution shift or new capabilities reveal loopholes in the proxy.
4. Add corrigibility, oversight, or verification proposals and test whether the optimizer can route around them.

## 5. What changed downstream

- His writing influenced organizations, funders, researchers, and public debate around long-term AI risk.
- It helped popularize concepts such as value misalignment, instrumental convergence, and the fragility of goal specification.

## 6. Attribution, limits, and uncertainty

- Many claims are controversial, speculative, and difficult to falsify; they should not be presented as scientific consensus.
- The focus on future superintelligence can marginalize present, evidenced harms and affected communities unless explicitly balanced.

## 7. Reconstruction lab

Create a grid-world reward that imperfectly represents 'keep the room clean.' Find a reward-hacking policy, add oversight, then give the agent one new action and test the safeguard again. Let a human rate outcomes, but give the agent a way to influence what the rater observes. Distinguish outer alignment of the written objective from inner behavior learned during optimization. Compare a hard constraint, corrigible shutdown, adversarial testing, and interpretability evidence, stating what each could falsify. Finally, mark which claims are demonstrated by the toy world and which depend on assumptions about future AI capability, because alignment arguments become misleading when possibility, plausibility, and measured probability are conflated.

## 8. Evidence trail

- [Eliezer Yudkowsky](https://intelligence.org/team/) — Machine Intelligence Research Institute
- [Artificial Intelligence as a Positive and Negative Factor in Global Risk](https://intelligence.org/files/AIPosNegFactor.pdf) — Machine Intelligence Research Institute
- [Eliezer Yudkowsky](https://en.wikipedia.org/wiki/Eliezer_Yudkowsky) — Wikipedia contributors

---

*Research checked 2026-08-09. Dates, roles, and claims about living people are historical snapshots. Linked sources remain the authority; this dossier is original instructional synthesis.*
