For better outcomes: we can slow AI capabilities AND accelerate AI safety

Headshot of Adam Jones

Adam Jones

AI capabilities are improving fast, and probably at an accelerating rate, likely due to RSI-like factors.

Scatter chart of the length of software tasks LLMs can complete 50% of the time, against model release date, on a linear scale. The points sit near zero from GPT-3 in 2019 through GPT-4 in 2023, then curve sharply upwards through o3 and GPT-5 to Claude Opus 4.6 at around 12 hours in 2026.

METR's time horizon measurements, on a linear scale. Usually this chart is shown logarithmically, which makes the trend a straight line — worth remembering that's what a straight line there means.

I don't think we have equivalently good measurements of our capacity to handle those capabilities. (These would be useful!)

But my guess is these are increasing slightly more than linearly, and not fast enough to keep up with AI development.

An exponential curve labelled 'AI capabilities' rising steeply away from a straight line labelled 'our capacity to handle them'. The gap between the two widens toward the right of the chart.

If this continues, it might lead to catastrophic outcomes, like a permanent totalitarian state or losing control of these systems entirely.

There are two ways to change this dynamic:

  • Push down the capabilities line. For example, per Anthropic:

    We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology.

  • Push up the handling capacity line. Without a pause or slowdown, we will need AI systems to be extremely good at AI safety work: alignment research, safeguards development, governance research, treaty negotiation, AGI macrostrategy, cybersecurity, pandemic prevention and response, and much more.

Many people I meet view things through only one of these frames, and should consider more actions in both.