Key events in
AI safety

The papers, capability milestones, policy moves, and incidents that shaped the field — and that people in it will assume you know about.

2000

2014

July 2014Book

Bostrom, Superintelligence

Brought the orthogonality thesis, instrumental convergence, and the control problem to a mainstream audience. Defines the term "Superintelligence". For roughly a decade this was the reference point for what "AI risk" meant.

2017

2018

2019

2020

2022

30 November 2022Capability

ChatGPT is released

The event that changed everything about the field's context. AI safety went from a niche concern to a mainstream political issue in roughly six months — nearly all the funding, institutions and regulation below this point exist because of it.

2023

March 2023Capability

GPT-4

Professional-exam-level performance across many domains, shipped with a system card documenting dangerous-capability evaluations, including ARC's tests for autonomous replication. Pre-deployment safety evaluation became an expected part of a frontier launch.

September 2023Governance

Anthropic's Responsible Scaling Policy

The first published commitment tying deployment to capability thresholds: define danger levels in advance, commit to safeguards before crossing them. OpenAI's Preparedness Framework and DeepMind's Frontier Safety Framework followed. The template for voluntary frontier-lab governance.

November 2023Governance

Bletchley Park summit and the first AI Safety Institutes

28 countries plus the EU — including both the US and China — signed a declaration acknowledging frontier AI risk. The UK and US AI Safety Institutes followed, giving governments in-house capacity to evaluate models rather than relying on labs' own reporting.

2024

September 2024Capability

OpenAI o1 and inference-time reasoning

Spending more compute at inference on a chain of thought opened a second scaling axis beyond pretraining. For safety it cut both ways: reasoning traces are a new window into model cognition, and a new thing models can learn to obfuscate.

2025

January 2025Capability

DeepSeek-R1

Frontier-adjacent reasoning performance, trained at a fraction of the expected cost and released with open weights. Undercut the assumption that capability could be controlled through a handful of well-resourced labs.

September 2025Paper

Anti-scheming training

OpenAI and Apollo Research trained models against covert behaviour and cut it substantially — but not to zero, and with a confound: trained models increasingly mentioned they might be being evaluated. Evaluation awareness makes reduced scheming hard to tell from better-hidden scheming.

2026

Nothing in that category yet.

Last updated July 2026. This is a curated and opinionated selection, not a complete record — the aim is the events a newcomer keeps hearing referenced without explanation. Think something important is missing? Suggest it on GitHub. For the concepts behind these events, see the concept map.

🧭

Want one-on-one guidance?

Paid career calls — application reviews, research direction, and navigating the field — are paused for now, and may resume soon.

Read more →