# Nick Bostrom

> 1973– · Philosopher
>
> **Recorded contribution:** Superintelligence; existential risk from AI; simulation argument

## How to use this dossier

Read for a causal chain, not a hero story: inherited problem → contribution → mechanism → downstream capability → limit. Then close the page and complete the reconstruction exercise from memory.

## 1. Historical orientation

Swedish philosopher Nick Bostrom founded Oxford's Future of Humanity Institute and wrote on anthropic reasoning, simulation arguments, existential risk, and machine superintelligence. His 2014 book Superintelligence popularized the control problem and strategic risks of highly capable AI among researchers, policymakers, and technologists. This is conceptual agenda-setting, not the invention of an AI mechanism.

## 2. The problem inherited

Debates about advanced AI often treated capability as an uncomplicated benefit and lacked explicit models of objectives, power accumulation, irreversible failure, and uncertainty over future systems.

## 3. The central contribution

Bostrom articulated philosophical and strategic arguments for studying how a system more capable than its designers might pursue misaligned goals and create catastrophic or existential risk.

## 4. Reconstruct the mechanism

1. Specify an agent with an objective and capabilities that may improve or accumulate resources.
2. Separate intelligence or optimization power from the content of the objective.
3. Trace instrumental strategies that could arise across many final goals.
4. Evaluate whether control, verification, governance, or containment remains possible as capability grows.

## 5. What changed downstream

- The work helped move long-horizon AI safety and governance into mainstream technical and policy discussion.
- It influenced research agendas on alignment, interpretability, evaluation, and catastrophic-risk controls.

## 6. Attribution, limits, and uncertainty

- The arguments are contested and depend on uncertain assumptions about future architectures, takeoff dynamics, agency, and social response.
- Attention to speculative existential scenarios can crowd out present harms, labor, concentration, environmental cost, and affected communities unless both horizons are studied.

## 7. Reconstruction lab

Write a two-agent toy model with a misspecified proxy objective. Give the optimizer one new capability at each round and identify the first point where a previously adequate safeguard fails. Formally distinguish the system’s objective, the designer’s intended outcome, and the evidence available to an overseer. Construct a case where the agent follows its written objective exactly while defeating the purpose of the task. Compare capability control, motivation alignment, monitoring, and shutdown as separate interventions. State which conclusion follows from your model and which remains a philosophical extrapolation about future systems; do not let a vivid scenario masquerade as probability evidence. Run sensitivity analysis over uncertain capability and oversight assumptions so the model communicates conditional results instead of manufacturing a single authoritative risk number.

## 8. Evidence trail

- [Nick Bostrom](https://www.ox.ac.uk/news-and-events/find-an-expert/professor-nick-bostrom) — University of Oxford
- [Superintelligence: Paths, Dangers, Strategies](https://global.oup.com/academic/product/superintelligence-9780198739838) — Oxford University Press
- [Existential Risk Prevention as Global Priority](https://nickbostrom.com/existential/risks) — Global Policy

---

*Research checked 2026-08-09. Dates, roles, and claims about living people are historical snapshots. Linked sources remain the authority; this dossier is original instructional synthesis.*
