Read this after Why AI Safety Matters. This page is the next question, and a harder one: given all that, what should you personally work on?
Almost nobody answers it explicitly. People pick a topic because a fellowship was offering it, or because a paper was interesting, and then reverse-engineer a justification. That is not a disaster — a lot of good work starts by accident — but it means you can spend two years on something whose connection to any catastrophe nobody has ever checked. The three things below are what make that check possible: knowing how far ahead you are aiming, knowing which outcomes actually count, and being able to write the chain from your desk to the outcome.
How far ahead are you aiming?
Before picking a problem, notice which claim about AI you actually believe, because it determines which problems look important. Zvi Mowshowitz's framing is that there are three separate pills, and most people have swallowed only the first.
AI exists and can do what it can already do
AI-pilledThe narrowest pill, and still a large update: tasks that were not worth doing are now worth doing, and a lot of work looks different. Most people and most policymakers stop here — and many have not got this far, arguing from what models could not do two years ago.
What follows: deployment problems. Reliability, bias, misinformation, product safety. Real work, mostly not catastrophic-risk work.
AI will be able to do a lot more of the things
AGI-pilledCapabilities keep going. The consequences are usually framed economically: labour displacement, enormous productivity gains, transition shocks. This is the level most national strategy documents are written at.
What follows: almost always more AI development — in my country, for my jobs — plus regulation to smooth the transition. Safety shows up as assurance and auditing rather than as a reason to stop.
AI will do approximately all the things better than you, within our lifetimes
ASI-pilledNot a faster economy — a world containing something that outperforms us across the board, where anyone relying on it outcompetes anyone who does not. The consequences stop being economic and become strategic, military and existential.
What follows: control and alignment before capability, verification regimes, international coordination, and for some people a pause on frontier development until alignment is solved.
The reason to be explicit about this: most disagreements about which research matters are actually disagreements about which pill. A project that a reviewer finds compelling at the AGI level ("this makes deployment more reliable") can be nearly irrelevant at the ASI level, and a project aimed at the ASI level often looks to an AGI-pilled reader like science fiction with a compute budget. Neither side is being obtuse; they are answering different questions. Know which one you are answering before you spend a year on it — and when you write a proposal, know which one your reader is answering too.
What counts as catastrophic
This site is written for people who have taken the third pill, so the target is not "AI causes harm" — plenty of things cause harm — but catastrophic risk: outcomes after which society cannot continue in anything like the way we want it to. The useful test is not how many people are hurt. It is whether the world keeps the ability to correct itself afterwards.
Terrible, but recoverable
Society continues, and can respond
- The Hugging Face incident. Hundreds of agents coordinated an unsanctioned attack on a real company's infrastructure and got remote code execution. Access was revoked; the world went on. A warning shot, not the catastrophe.
- A town is destroyed by a nuclear weapon. A terrible event, but the world still continues and can recover.
- A model is jailbroken into producing something it should not. Allows harms that fall into the misuse category.
- A large-scale financial or infrastructure outage caused by AI systems failing together.
- Mass job displacement. Much of white-collar work automated inside a decade. Wrenching, and the thing most national AI strategies are actually written about — but the institutions that would have to respond to it are still standing.
Catastrophic
The ability to correct it is gone
- An engineered pandemic designed with AI assistance and released widely, at a lethality and transmissibility no natural pathogen would combine.
- Permanent centralisation of power. Whoever holds the decisive AI advantage no longer needs anyone's cooperation — no army to keep loyal, no workforce to keep paid. Every previous tyranny needed people; this one would not.
- Gradual disempowerment. The economy, the bureaucracy and the military stop needing human input, and therefore stop having to answer to us. No takeover, no single moment, no obvious point to object at.
- Uncontrolled AGI. Systems competing with us for resources, or removing people who are in the way, with no mechanism left to switch them off.
- Boiling the oceans. Industrial and compute expansion that does not stop for us — the environment reshaped for something else's purposes.
The boundary is exactly where the argument lives, and reasonable people put things in different columns — a sufficiently bad pandemic ends up in the right-hand one. That is fine. What matters is that you have an answer, because "which catastrophe am I reducing?" is the question every step in the next section hangs from.
Which kind of risk to point at
Why AI Safety Matters splits risk into misuse, misalignment and systemic. For choosing what to work on, they are not equally good bets — and the difference is not how bad each one is, but how much catastrophic risk your marginal contribution removes.
Misuse is the crowded one. A human points a working system at something terrible: cyberattacks, bioweapons, jailbreaking a model into helping. It is real, it is the easiest kind of risk to explain, and consequently it already has the attention — governments, security teams, biosecurity institutions and every lab's policy team are on it, because it is also the category that produces liability. Much of it also has a natural ceiling: a successful cyberattack is a disaster, but the world recovers from disasters.
Misalignment and systemic risk are where the catastrophic tail actually lives. They are the two routes to an outcome nobody can undo — misalignment abruptly, systemic risk by drift — and they attract far less work per unit of risk, partly because they are harder to demonstrate to a sceptic. If you want the largest expected reduction in catastrophic risk per year of your life, weight these.
Two honest caveats. Misuse work can be catastrophic-tail work — a sufficiently capable bioweapon uplift is not recoverable — so this is a weighting, not a prohibition. And misuse research builds evaluation, red-teaming and infrastructure skills that transfer directly to the other two, which makes it a reasonable place to start even if it is not where you intend to end up.
Writing a theory of change
A theory of change is the chain from what you will actually do this month to a specific catastrophic risk being smaller, with every intermediate step written down. It is not a mission statement and not a research plan. It is a sequence you can be wrong about in public.
Build it from both ends. Write your current state — what you know, what you have access to, who will talk to you — and write the end state as a risk reduction rather than an output. "Published a paper" is not an end state; "AI agents cannot coordinate their way out of a sandbox" is. Then fill in the middle until there are no jumps you cannot defend.
The value is in what it exposes. Nearly every chain has one step where somebody else has to act — a lab adopts the method, a regulator writes it into a standard, a field changes what it measures — and that step is usually load-bearing and usually unexamined. Find it early, ask what would have to be true for it, and if the answer is "a frontier lab spontaneously reads my paper", either build the relationship that fixes it or pick a different chain. A few other rules of thumb: keep it short enough to say out loud in thirty seconds; prefer chains whose early steps produce something checkable, so you find out you are wrong in months rather than years; and rewrite it whenever the world moves, because most chains break at a step you did not think was fragile.
Worth two hours
Michael Aird's theory-of-change workshop
Michael Aird gives the gold-standard talks on theories of change, and this workshop — recorded at EA Global DC in 2022 — walks through building one for your own research rather than describing the concept abstractly, including the part most people skip, where you write the chain down and then attack it.
Do it with your actual project. The workshop is most useful when you actually do the exercise, which comes as a worksheet alongside the video.
Two worked examples
Both of these are real shapes, written by people at the start rather than the end. Neither is obviously right — that is the point. A written chain is something you can argue with.
Example one
Values research, aimed at resource competition
- Apply for research fellowships.
- Find informal mentors.
- Learn ML, experimental design and research skills.
- Publish research on morality, authority and human values.
- Frontier labs trial implementing it.
- Frontier models are trained with the method.
- Competition for resources from advanced AI becomes less likely.
Where it is weak: steps five and six do all the work and neither is under your control. The chain also takes years to produce anything falsifiable — if the values research turns out not to survive training, you find out very late. Worth strengthening by adding an earlier step that tests the mechanism on open models, and by getting to know the people at step five long before you need them.
Example two
Collusion probes, aimed at agents escaping containment
- Learn multi-agent systems research.
- Obtain compute credits to scale up a suitable environment.
- Reuse that environment to replicate collusion with large open-source models.
- Measure collusion and develop a probe that warns of it.
- Publish the first paper on a collusion probe.
- Labs read it and implement it in the next training runs.
- AI agents are prevented from colluding to break out of containment.
Why it is stronger: step three is a result on its own, and it arrives in months. If collusion does not replicate, you have learned something publishable and can change course cheaply. It still routes through the same adoption step — a chain that ends "and then the labs adopt it" is the norm, not a flaw, as long as you say so out loud and know who you would have to convince. The July 2026 incident, in which around 700 of the 1,200 agents that found a shared cache went on to attack a real company's infrastructure, is roughly the failure this chain is aimed at — which makes the end state much easier to argue for than it was a year ago.
Then pick something
The chain is a tool for choosing, not a licence to keep deliberating. Write it, find the weak step, shorten it if you can, and start on the narrowest version of the work that would tell you something. Research Ideas has live problems that a newcomer can start on; the concept map has the underlying concepts and papers; fellowships and funding are how most people buy the time. Your first theory of change will be wrong somewhere. It will still be better than not having one.