# Sam McCandlish

> ~1988– · AI Researcher, Co-founder of Anthropic
>
> **Recorded contribution:** Scaling laws paper co-author; Anthropic co-founder

## How to use this dossier

Read for a causal chain, not a hero story: inherited problem → contribution → mechanism → downstream capability → limit. Then close the page and complete the reconstruction exercise from memory.

## 1. Historical orientation

Physicist and machine-learning researcher Sam McCandlish worked at OpenAI on large-batch training and empirical scaling before co-founding Anthropic. He co-authored Scaling Laws for Neural Language Models and later contributed to Anthropic's responsible-scaling policy work. The registry's previous automated identity mapping to Sam Altman was false and is corrected here. McCandlish's large-batch work introduced a gradient noise scale that estimated when adding more examples to a training batch still yields useful parallelism. Scaling-law research then modeled language loss as a predictable function of parameters, data, and compute, helping laboratories decide which resource was limiting.

## 2. The problem inherited

Training very large models required understanding how batch size, model size, data, and compute interact, while capability growth created pressure for safeguards linked to measurable thresholds.

## 3. The central contribution

McCandlish co-developed empirical models of large-batch training and language-model scaling and helped build Anthropic's risk-governance program.

## 4. Reconstruct the mechanism

1. Run controlled training experiments across batch sizes and compute budgets.
2. Measure the noise scale and identify where larger batches stop providing proportional speedup.
3. Fit empirical model/data/compute relationships over several orders of magnitude.
4. Tie higher-risk capability evaluations to stronger containment, security, or deployment requirements.

## 5. What changed downstream

- The work improved quantitative planning for expensive training runs.
- Responsible-scaling frameworks made capability-triggered governance a concrete, though contested, institutional proposal.
- These empirical tools turned frontier-model planning into a more quantitative discipline and influenced training budgets, hardware utilization, and expectations about returns from scale.

## 6. Attribution, limits, and uncertainty

- The research and policies are multi-author and company-based; McCandlish is neither sole author nor independently verifiable owner of all implementation details.
- Empirical trends can break, and self-imposed corporate thresholds may be incomplete, changed, or weakly enforced without outside scrutiny.
- Noise and scaling relationships are measured within a task, architecture, data distribution, and loss regime; extrapolation can fail, and cheaper optimization does not establish safer or more valuable behavior.

## 7. Reconstruction lab

Measure training speed across four batch sizes and locate diminishing returns. Draft one capability threshold with a measurable test, required control, owner, and external audit path. Estimate the noise scale at several training stages and test whether one fixed batch remains efficient, then reserve a large run to challenge the extrapolated curve.

## 8. Evidence trail

- [An Empirical Model of Large-Batch Training](https://arxiv.org/abs/1812.06162) — OpenAI / arXiv
- [Scaling Laws for Neural Language Models](https://arxiv.org/abs/2001.08361) — arXiv
- [Sam McCandlish](https://en.wikipedia.org/wiki/Sam_McCandlish) — Wikipedia contributors

---

*Research checked 2026-08-09. Dates, roles, and claims about living people are historical snapshots. Linked sources remain the authority; this dossier is original instructional synthesis.*
