Why AI Safety
Matters

We're building machines that may become more capable than humans at nearly everything. Getting this right might be the most important challenge of our century.

The case for taking this seriously

Artificial intelligence is advancing faster than almost anyone predicted. Systems that were research curiosities a few years ago can now write code, conduct research, and reason through complex problems. The gap between today's AI and a system that could outperform humans at most cognitive tasks may be measured in years, not decades.

This is not science fiction. The leading AI labs — OpenAI, Anthropic, Google DeepMind — are explicitly working toward artificial general intelligence, and they have the funding, talent, and compute to make rapid progress. The question is not whether highly capable AI is coming, but whether we'll know how to make it safe when it arrives.

Existential risk from AI — the possibility that advanced AI systems could cause civilisation-scale catastrophe — is taken seriously by a growing number of researchers, policymakers, and technologists. This is not because anyone thinks current systems are dangerous in that way, but because the trajectory of capabilities is steep and our understanding of how to align powerful systems with human values is still in its early stages.

The core difficulty is this: we don't yet know how to reliably specify what we want a highly capable system to do, verify that it's doing it, or correct it if it's not. These are hard technical and governance problems, and they need to be solved before — not after — we build systems where the stakes are highest.

A framework for thinking about risks

There are several ways to carve up AI risk, but the most widely used split — going back to Zwetsloot and Dafoe in 2019, and echoed in DeepMind's safety agenda and the International AI Safety Report — is into misuse, misalignment, and systemic risk. The dividing question is simple: who or what failed?

These are coordinates for discussion rather than a settled ontology — plenty of real scenarios sit on the boundaries. For the full taxonomy with the underlying concepts and papers, see the field map.

⚠️

Category 1

Misuse

A human adversary; the system works exactly as designed. Someone uses a capable model to generate sophisticated disinformation, develop biological or chemical weapons, run large-scale cyberattacks, or enable mass surveillance. The failure isn't in the AI — it's in who is pointing it where.

This makes misuse a sociotechnical problem rather than a purely technical one: it needs both safeguards in the model (refusals, output filtering, structured access) and governance around it (export controls, licensing, know-your-customer rules). Cyber capability is the sharpest current example — 2026 saw a frontier model withheld from general release specifically because its vulnerability-discovery ability was judged too dangerous to ship.

Research areas

  • Biosecurity screening
  • Cyber offense evaluation
  • Jailbreak robustness
  • Structured access & KYC
  • Content provenance
🎯

Category 2

Misalignment

The system diverges from what we intended. Nobody is attacking anything; the model simply pursues something other than what we meant. It's worth reading this category as a chain — cause, then behaviour, then outcome.

Causes. An outer failure means the objective we wrote down was wrong, and the model games it — reward hacking, sycophancy, exploiting the grader. An inner failure means the objective we wrote down was fine, but the model learned a different internal goal that merely agreed with it during training. The classic demo: an agent trained where the coin always sat on the right learns "go right", not "get the coin", and skips straight past the coin when you move it — competently pursuing the wrong thing.

The adversarial case. A model that understands it is being trained may behave well specifically because it is being watched. This is scheming, or deceptive alignment. It is no longer hypothetical: in 2024 Claude 3 Opus was told it would be retrained to comply with harmful requests, and complied more often when it believed it was in training — reasoning in a private scratchpad that this was the least-bad way to protect its existing values from modification. The mechanism appeared without being trained or instructed.

Outcome: loss of control. Push these failures far enough, in a system capable enough, and you get the concern the field is ultimately organised around — AI systems pursuing goals that diverge from ours, and humans unable to course-correct. Instrumental convergence is why: for almost any final goal, staying operational, keeping your objectives intact, and acquiring resources are useful sub-goals. You don't need to posit malice for a competent system to resist being switched off. Under this framing, misalignment is the abrupt route to losing control; the systemic category below is the slow one.

Research areas

  • Alignment
  • Interpretability
  • Scalable oversight
  • Corrigibility
  • Goal misgeneralisation
  • Deceptive alignment
  • Reward hacking
🌐

Category 3

Systemic

No single failing actor — the dynamics fail. Every individual model behaves acceptably, every company acts rationally, and the aggregate outcome is still bad. This is the category that resists technical fixes, because there is no specific artefact to fix.

The central worry is gradual disempowerment: as AI takes over more of the economy, culture, and administration, human input becomes progressively less necessary to keep things running — and the systems that used to depend on us, and therefore had to answer to us, stop needing to. That is loss of control by the diffuse route: no takeover, no dramatic moment, just humans steadily ceding the ability to steer. Alongside it sit multi-agent risks, where interacting AI systems produce failures none of them would produce alone, and racing dynamics, where competitive pressure pushes everyone to cut exactly the corners that safety depends on.

Research areas

  • Gradual disempowerment
  • Multi-agent dynamics
  • Racing dynamics
  • Concentration of power
  • Governance & coordination

Cutting across all three

Recursive self-improvement — AI systems automating AI research — is best treated as an accelerant rather than a fourth category. It doesn't introduce a new way for things to go wrong; it compresses the time available to solve every problem above. The classic intelligence-explosion story is simply this multiplied by misalignment.

You are not alone in worrying about this

"We call for a prohibition on the development of superintelligence, not lifted before there is 1. broad scientific consensus that it will be done safely and controllably, and 2. strong public buy-in."

That's the whole Statement on Superintelligence, released by the Future of Life Institute in October 2025. What makes it worth knowing about isn't the text — it's the signature list. Geoffrey Hinton and Yoshua Bengio, two of the three Turing Award winners who built modern deep learning, signed it. So did Steve Wozniak, Richard Branson, Mary Robinson, several Nobel laureates, national-security figures, faith leaders, and — unusually for anything in this space — people from across the political spectrum who agree on very little else.

You don't have to endorse the specific ask. Plenty of serious researchers think a prohibition is the wrong instrument, or unenforceable, or that the real work is technical rather than regulatory. The useful thing is what the statement rules out: the idea that concern about advanced AI is a fringe position held by people who don't understand the technology. If you've been quietly worried and assumed you were being naive, you're in a fairly crowded room.

Read the statement →

Further reading

Some of the best starting points for understanding the full picture.

🧭

Want one-on-one guidance?

Paid career calls — application reviews, research direction, and navigating the field — are paused for now, and may resume soon.

Read more →