
We Need Scalable AI Safety
We should build the capacity to scale our defenses before we urgently need to turn them up.
We’ve built a powerful engine for scaling AI capabilities: invest more capital and compute, and we can produce increasingly capable systems. We need to build a corresponding engine for safety.
Suppose we decided tomorrow to devote a much larger compute budget to making the world safer from AI. What could we spend it on effectively? We need programs that can absorb those resources and turn them into credible reductions in risk: continuously monitoring dangerous behavior, securing critical infrastructure, and researching better safeguards.
Proposals for pacing frontier AI development already explore how we might allocate compute differently. The AI Futures Project, for example, proposes minimum allocations of 70% for external inference and 25% for transparent safety research, leaving 5% for capabilities R&D. This would direct more resources toward using existing capabilities and studying their risks while slowing the development of new ones.
Where would you turn the dials?
Move a slider. The other allocations rebalance to keep the budget at 100%.
Train and develop more capable systems.
Put existing capabilities to work.
Observe activity and investigate warning signs.
Find, validate, and fix vulnerabilities.
Develop safeguards and evaluate what works.
Compute units are illustrative; these allocations are not estimates of safety. The pacing preset uses the essay’s 70/25/5 proposal; the other presets are examples. More compute only helps if we can turn it into useful work.
I think we should develop this idea further by identifying specific safety programs that could productively absorb much larger budgets. Along this line of thinking, let me share three possibilities that come to mind:
Scalable AI monitoring and sensing.
We could invest in systems that detect concerning AI activity early, from unexpected agent behavior to coordinated attempts to compromise infrastructure or spread beyond authorized environments. Think of this as an immune system: detecting threats and coordinating a response before they cascade. AI agents could help monitor, review, and report the behavior of other agents. For this to work, we would need appropriate visibility into their activity, safeguards for privacy, and ways to verify that the monitoring agents remain trustworthy. This includes verifying that agents that are part of the monitoring and sensing work don’t collude with the agents they oversee.
Scalable AI cyber defense.
Initiatives such as Project Glasswing and Daybreak suggest one direction: using AI to secure the software and infrastructure society depends on. Additional compute could support continuous testing, vulnerability discovery, and the development and verification of patches. In order to respond to AI swarm speed cyber attacks, we’ll need an always-on AI-speed defense. Teams of agents could work together to strengthen defenses and respond to threats identified by monitoring systems. The important question is whether additional defensive resources can help us close vulnerabilities and respond to attacks faster than adversaries can exploit them.
OpenAI describes its approach as a Defense Factory: an automated operation that continuously finds, validates, and fixes vulnerabilities. It also describes a “defender’s window,” during which defenders can use their access to frontier models and their own code to get ahead of attackers. The question is whether additional defensive resources can help us close vulnerabilities and respond to attacks faster than adversaries can exploit them.
Scalable safety research.
We could direct large amounts of compute toward conducting large-scale AI research on developing and testing better technical safeguards and governance mechanisms. Imagine tens of thousands of agents working on AI safety research, exploring promising directions and mechanisms, collectively performing the equivalent of thousands of human-years of research. This would be equivalent to many Navier-Stokes sized research focused on AI safety.
Scott Alexander described this related possibility, in a post he wrote a few years back:
“Consider the dumbest AI that can solve the alignment problem. It’s possible that this AI is no smarter than the top human researchers (because we can mass-produce it by the millions and run it for subjective centuries, and if we had a million top human researchers work on the problem for subjective centuries, probably they could solve it too). If the dumbest AI that can solve the alignment problem comes before the sorts of AIs that can precipitate the point of no return, then they can solve the alignment problem for us.”
Over time, one of the goals of this research could be to build a super-intelligent allocator of safety resources. As capabilities and threats evolve, it could identify where additional capital and compute would do the most good, evaluate the results, and adjust its recommendations. Building that feedback loop could itself be one of the most valuable outcomes of scalable safety research. One of the biggest challenges will be how do we keep human-trust in the loop when we hand off safety research to AI as well.
These are candidates for scalable safety. For each, we should ask what an additional dollar or unit of compute buys, where the bottlenecks emerge, and how we would verify that the added activity reduces risk. Discovering more vulnerabilities only helps if we can fix them; generating more research only helps if its findings are sound.
I’m skeptical that any single intervention will guarantee AI safety. Even substantial advances in mechanistic interpretability may leave uncertainty about how broadly deployed systems will behave - so we won’t have a Lean style verifiable proof of alignment. I expect we’ll need to model the risks and combine several approaches to reduce the likelihood of catastrophic failures. We’ll also need to understand where those approaches share weaknesses, so that one failure doesn’t undermine every layer of defense.
We can think of compute and capital allocation as a set of sliders. Some allocations accelerate new capabilities; others strengthen defenses, improve oversight, or support safety research. How we set those sliders could influence both the pace of AI development and our ability to manage its consequences.
I wanted to share this framing of safety as “scalable” to encourage folks in the field to think of other initiatives where we might be able to scale AI, and to add more slider candidates into the mix.
Building these programs would give companies, governments, and other funders concrete ways to invest in safer outcomes. We would still need to learn which approaches work and how their returns change with scale. But having effective programs ready to absorb resources would itself be valuable.
What if automating AI R&D triggers an intelligence explosion? identifies three priorities: gaining visibility into AI R&D automation, developing ways to steer and constrain an intelligence explosion, and preparing society to adapt to its effects. Scalable safety programs could help make parts of that agenda operational by providing systems we can fund, expand, evaluate, and improve as circumstances change.

If AI systems begin accelerating their own development through recursive self-improvement, our ability to respond will need to scale with them. We should build the capacity to scale our defenses before we urgently need to turn them up.