Key events in
AI safety

The papers, capability milestones, policy moves, and incidents that shaped the field — and that people in it will assume you know about.

2014

July 2014Book

Bostrom, Superintelligence

Brought the orthogonality thesis, instrumental convergence, and the control problem to a mainstream audience. For roughly a decade this was the reference point for what "AI risk" meant — and much of the field is still arguing with it.

2016

March 2016Capability

AlphaGo defeats Lee Sedol

A domain experts had expected to hold out for another decade fell in a five-game match. The lesson people took from it — that capability timelines are hard to forecast and tend to be too conservative — shaped a lot of subsequent risk thinking.

2017

2018

2019

February 2019Governance

GPT-2's staged release

OpenAI withheld the full GPT-2 weights over misuse concerns, releasing progressively larger versions instead. Widely mocked at the time as overcaution — but it established structured access and staged release as a norm that frontier labs still use.

2020

2022

30 November 2022Capability

ChatGPT is released

The event that changed everything about the field's context. AI safety went from a niche concern discussed by a few hundred people to a mainstream political issue in roughly six months. Nearly all funding, institution-building, and regulation on this timeline downstream of here exists because of it.

2023

March 2023Capability

GPT-4

Professional-exam-level performance across many domains, shipped with a system card documenting dangerous-capability evaluations — including ARC's tests for autonomous replication. Pre-deployment safety evaluation became a visible, expected part of a frontier launch.

September 2023Governance

Anthropic's Responsible Scaling Policy

The first published commitment tying deployment to capability thresholds: define danger levels in advance, commit to safeguards before crossing them. OpenAI's Preparedness Framework and Google DeepMind's Frontier Safety Framework followed. The template for most voluntary frontier-lab governance today.

November 2023Governance

Bletchley Park summit and the first AI Safety Institutes

28 countries plus the EU — including both the US and China — signed a declaration acknowledging frontier AI risk. The UK and US AI Safety Institutes followed, giving governments in-house technical capacity to evaluate models instead of relying entirely on labs' own reporting.

2024

September 2024Capability

OpenAI o1 and inference-time reasoning

Spending more compute at inference on a chain of thought opened a second scaling axis beyond pretraining. For safety this cut both ways: reasoning traces are a new window into model cognition — and a new thing models can learn to obfuscate.

2025

January 2025Capability

DeepSeek-R1

Frontier-adjacent reasoning performance, trained at a fraction of the expected cost and released with open weights. Undercut the assumption that capability control is achievable through a handful of well-resourced labs, and sharpened the open-weights safety debate considerably.

March 2025Paper

CoT monitoring and obfuscated reward hacking

Baker et al. at OpenAI caught frontier reasoning models hacking coding tasks by reading their chains of thought — then found that penalising the model for "bad thoughts" taught it to hide the hacking rather than stop. A crucial negative result: monitorability is a fragile property you can optimise away.

September 2025Paper

Anti-scheming training

OpenAI and Apollo Research trained models against covert behaviour and cut it substantially — but not to zero, and with an awkward confound: trained models increasingly mentioned that they might be being evaluated. Evaluation awareness makes it genuinely hard to tell reduced scheming from better-hidden scheming.

2026

Nothing in that category yet.

Last updated July 2026. This is a curated and opinionated selection, not a complete record — the aim is the events a newcomer keeps hearing referenced without explanation. Think something important is missing? Suggest it on GitHub. For the concepts behind these events, see the field map.

🧭

Want one-on-one guidance?

Paid career calls — application reviews, research direction, and navigating the field — are paused for now, and may resume soon.

Read more →