# Jared Kaplan

> ~1985– · Physicist, AI Researcher, Co-founder of Anthropic
>
> **Recorded contribution:** Neural scaling laws co-author; Anthropic co-founder; physics-to-ML bridge

## How to use this dossier

Read for a causal chain, not a hero story: inherited problem → contribution → mechanism → downstream capability → limit. Then close the page and complete the reconstruction exercise from memory.

## 1. Historical orientation

Theoretical physicist Jared Kaplan became a Johns Hopkins professor, worked at OpenAI, and co-founded Anthropic. He was lead author of the 2020 paper Scaling Laws for Neural Language Models, which measured power-law relationships among language-model loss, parameter count, dataset size, and training compute. The work helped make model development more predictable and capital planning more quantitative. Kaplan's scaling-law work treated model performance as an empirical function of model size, dataset size, and compute. Smooth power-law relationships suggested that many training outcomes could be forecast across orders of magnitude, converting expensive frontier experiments into a resource-allocation problem.

## 2. The problem inherited

Researchers increased model, data, and compute budgets without a sufficiently clear empirical account of which resource was limiting or how far observed trends might continue.

## 3. The central contribution

Kaplan led an influential empirical scaling-law study and co-founded Anthropic, connecting quantitative model scaling with frontier-lab research and governance.

## 4. Reconstruct the mechanism

1. Train a family of otherwise comparable language models across parameter and data scales.
2. Measure held-out cross-entropy as a function of model size, tokens, and compute.
3. Fit power laws over the observed regime and estimate the resource-limiting frontier.
4. Allocate a fixed compute budget using the fitted relationships, then test extrapolations on new runs.

## 5. What changed downstream

- Scaling laws influenced model-size, dataset, and compute planning across frontier AI labs.
- They strengthened the expectation that capability and risk could change predictably with resource scale, motivating both investment and governance proposals.
- Scaling laws influenced how laboratories plan model, data, and compute budgets and supplied part of the empirical case for sustained investment in larger language models.

## 6. Attribution, limits, and uncertainty

- The paper has many authors; later compute-optimal work revised important allocation conclusions.
- A smooth average loss trend does not predict every capability, safety property, data regime, architecture change, environmental cost, or economic constraint.
- A fitted power law can break across architectures, data regimes, evaluation distributions, or post-training methods, and lower loss does not imply factuality, controllability, or social value.

## 7. Reconstruction lab

Train small language models at a three-by-three grid of sizes and token budgets. Fit log-log slopes, reserve one run for falsification, and report uncertainty rather than one extrapolated number. Fit competing exponents with confidence intervals, hold out the largest run, and state what result would falsify rather than merely update the claimed scaling relationship.

## 8. Evidence trail

- [Scaling Laws for Neural Language Models](https://arxiv.org/abs/2001.08361) — arXiv
- [Jared Kaplan](https://physics-astronomy.jhu.edu/directory/jared-kaplan/) — Johns Hopkins University
- [Anthropic company](https://www.anthropic.com/company) — Anthropic

---

*Research checked 2026-08-09. Dates, roles, and claims about living people are historical snapshots. Linked sources remain the authority; this dossier is original instructional synthesis.*
