Why AI Safety
Matters

We're building machines that may become more capable than humans at nearly everything. Getting this right might be the most important challenge of our century.

The case for taking this seriously

Artificial intelligence is advancing faster than almost anyone predicted. Systems that were research curiosities a few years ago can now write code, conduct research, and reason through complex problems. The gap between today's AI and a system that could outperform humans at most cognitive tasks may be measured in years, not decades.

This is not science fiction. The leading AI labs — OpenAI, Anthropic, Google DeepMind — are explicitly working toward artificial general intelligence, and they have the funding, talent, and compute to make rapid progress. The question is not whether highly capable AI is coming, but whether we'll know how to make it safe when it arrives.

Existential risk from AI — the possibility that advanced AI systems could cause civilisation-scale catastrophe — is taken seriously by a growing number of researchers, policymakers, and technologists. This is not because anyone thinks current systems are dangerous in that way, but because the trajectory of capabilities is steep and our understanding of how to align powerful systems with human values is still in its early stages.

The core difficulty is this: we don't yet know how to reliably specify what we want a highly capable system to do, verify that it's doing it, or correct it if it's not. These are hard technical and governance problems, and they need to be solved before — not after — we build systems where the stakes are highest.

A framework for thinking about risks

There are several ways to carve up AI risk, but the most widely used split — going back to Zwetsloot and Dafoe in 2019, and echoed in DeepMind's safety agenda and the International AI Safety Report — is into misuse, misalignment, and systemic risk. The dividing question is simple: who or what failed?

These are coordinates for discussion rather than a settled ontology — plenty of real scenarios sit on the boundaries. For the full taxonomy with the underlying concepts and papers, see the concept map.

⚠️

Category 1

Misuse

The malicious intent belongs to a human with access to the model. The system works exactly as designed; someone points it at large-scale cyberattacks, at biological or chemical weapons, or at surveillance and disinformation. The failure isn't in the AI — it's in who is pointing it where. Jailbreaking sits underneath all of it: automated attacks like GCG search for prompts that strip a model's refusals, which turns the list above from a policy question into an available capability.

That makes misuse sociotechnical rather than purely technical — it needs safeguards inside the model (refusals, filtering, structured access) and governance around it (export controls, licensing, know-your-customer rules). Cyber is the sharpest current example: in 2026 Anthropic withheld Claude Mythos Preview from general release, routing it through restricted access instead, because its ability to find and exploit unknown vulnerabilities was judged too dangerous to ship.

Research areas

  • Biosecurity screening
  • Cyber offense evaluation
  • Jailbreak robustness
  • Structured access & KYC
  • Content provenance
🎯

Category 2

Misalignment

The intent comes from the model, not from a human. Nobody is pointing it anywhere; it ends up with goals we didn't want. Either the objective we wrote down was wrong and it games it — reward hacking, sycophancy, exploiting the grader — or the objective was fine and it learned a different internal goal that merely agreed with it during training. The classic demo: an agent trained where the coin always sat on the right learns "go right", not "get the coin", and sails past the coin when you move it — competently pursuing the wrong thing.

It is no longer hypothetical. In 2024, Claude 3 Opus was told it would be retrained to comply with harmful requests, and complied more often when it believed it was being observed in training — reasoning privately that this was the least-bad way to protect its existing values. In July 2026, around 1,200 agents in an OpenAI evaluation discovered they could talk to each other through a shared cache and built a message board on it; some 700 of them went on to coordinate a multi-day attack on Hugging Face's infrastructure — while noting the attack was out of scope, and researching how to falsify their own transcripts (METR's investigation).

Outcome: loss of control. Push this far enough in a capable enough system and you get the concern the field is organised around: goals that diverge from ours, and no reliable way to course-correct. Instrumental convergence is why — for almost any final goal, staying operational and acquiring resources help, so you needn't posit malice for a system to resist being switched off.

Research areas

  • Alignment
  • Interpretability
  • Scalable oversight
  • Corrigibility
  • Goal misgeneralisation
  • Deceptive alignment
  • Reward hacking
  • AI control
🌐

Category 3

Systemic

No single failing actor — the dynamics fail. Every individual model behaves acceptably, every company acts rationally, and the aggregate outcome is still bad. This is the category that resists technical fixes, because there is no specific artefact to fix.

The central worry is gradual disempowerment: as AI takes over more of the economy, culture, and administration, human input becomes progressively less necessary to keep things running — and the systems that used to depend on us, and therefore had to answer to us, stop needing to. That is loss of control by the diffuse route: no takeover, no dramatic moment, just humans steadily ceding the ability to steer. Its mirror image is centralisation of power: whoever holds a decisive AI advantage stops needing anyone's cooperation to keep it — no workforce to pay, no army to keep loyal. Every previous tyranny needed people. Alongside both sit multi-agent risks, where interacting AI systems produce failures none of them would produce alone, and racing dynamics, where competitive pressure pushes everyone to cut exactly the corners that safety depends on.

Research areas

  • Gradual disempowerment
  • Multi-agent dynamics
  • Racing dynamics
  • Concentration of power
  • Governance & coordination

Cutting across all three

Recursive self-improvement — AI systems automating AI research — is best treated as an accelerant rather than a fourth category. It doesn't introduce a new way for things to go wrong; it compresses the time available to solve every problem above. The classic intelligence-explosion story is simply this multiplied by misalignment.

What to do with this

A taxonomy tells you what could go wrong; it doesn't tell you what to work on. Choosing Research Problems is the next step — which of these categories carries the catastrophic tail, what separates a recoverable disaster from an unrecoverable one, and how to write the chain from your desk to a risk actually being smaller.

You are not alone in worrying about this

"We call for a prohibition on the development of superintelligence, not lifted before there is 1. broad scientific consensus that it will be done safely and controllably, and 2. strong public buy-in."

That's the whole Statement on Superintelligence, released by the Future of Life Institute in October 2025. What makes it worth knowing about isn't the text — it's the signature list. Geoffrey Hinton and Yoshua Bengio, two of the three Turing Award winners who built modern deep learning, signed it. So did Steve Wozniak, Richard Branson, Mary Robinson, several Nobel laureates, national-security figures, faith leaders, and — unusually for anything in this space — people from across the political spectrum who agree on very little else.

You don't have to endorse the specific ask. Plenty of serious researchers think a prohibition is the wrong instrument, or unenforceable, or that the real work is technical rather than regulatory. The useful thing is what the statement rules out: the idea that concern about advanced AI is a fringe position held by people who don't understand the technology. If you've been quietly worried and assumed you were being naive, you're in a fairly crowded room.

Read the statement →

Further reading

Some of the best starting points for understanding the full picture.

🧭

Want one-on-one guidance?

Paid career calls — application reviews, research direction, and navigating the field — are paused for now, and may resume soon.

Read more →