2014
Bostrom, Superintelligence
Brought the orthogonality thesis, instrumental convergence, and the control problem to a mainstream audience. For roughly a decade this was the reference point for what "AI risk" meant — and much of the field is still arguing with it.
2016
AlphaGo defeats Lee Sedol
A domain experts had expected to hold out for another decade fell in a five-game match. The lesson people took from it — that capability timelines are hard to forecast and tend to be too conservative — shaped a lot of subsequent risk thinking.
2017
2018
2019
GPT-2's staged release
OpenAI withheld the full GPT-2 weights over misuse concerns, releasing progressively larger versions instead. Widely mocked at the time as overcaution — but it established structured access and staged release as a norm that frontier labs still use.
2020
2022
ChatGPT is released
The event that changed everything about the field's context. AI safety went from a niche concern discussed by a few hundred people to a mainstream political issue in roughly six months. Nearly all funding, institution-building, and regulation on this timeline downstream of here exists because of it.
2023
GPT-4
Professional-exam-level performance across many domains, shipped with a system card documenting dangerous-capability evaluations — including ARC's tests for autonomous replication. Pre-deployment safety evaluation became a visible, expected part of a frontier launch.
Anthropic's Responsible Scaling Policy
The first published commitment tying deployment to capability thresholds: define danger levels in advance, commit to safeguards before crossing them. OpenAI's Preparedness Framework and Google DeepMind's Frontier Safety Framework followed. The template for most voluntary frontier-lab governance today.
Bletchley Park summit and the first AI Safety Institutes
28 countries plus the EU — including both the US and China — signed a declaration acknowledging frontier AI risk. The UK and US AI Safety Institutes followed, giving governments in-house technical capacity to evaluate models instead of relying entirely on labs' own reporting.
2024
OpenAI o1 and inference-time reasoning
Spending more compute at inference on a chain of thought opened a second scaling axis beyond pretraining. For safety this cut both ways: reasoning traces are a new window into model cognition — and a new thing models can learn to obfuscate.
2025
DeepSeek-R1
Frontier-adjacent reasoning performance, trained at a fraction of the expected cost and released with open weights. Undercut the assumption that capability control is achievable through a handful of well-resourced labs, and sharpened the open-weights safety debate considerably.
CoT monitoring and obfuscated reward hacking
Baker et al. at OpenAI caught frontier reasoning models hacking coding tasks by reading their chains of thought — then found that penalising the model for "bad thoughts" taught it to hide the hacking rather than stop. A crucial negative result: monitorability is a fragile property you can optimise away.
Anti-scheming training
OpenAI and Apollo Research trained models against covert behaviour and cut it substantially — but not to zero, and with an awkward confound: trained models increasingly mentioned that they might be being evaluated. Evaluation awareness makes it genuinely hard to tell reduced scheming from better-hidden scheming.
2026
Last updated July 2026. This is a curated and opinionated selection, not a complete record — the aim is the events a newcomer keeps hearing referenced without explanation. Think something important is missing? Suggest it on GitHub. For the concepts behind these events, see the field map.