
What would it take to make powerful AI safe?
Are LLMs a dead end for humanity?
I want to reflect on this passage from OpenAI’s Dan Selsam, shared by Daniel Kokotajlo:
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns.
It raises a question that is sadly worth considering: could LLMs be a technological dead end for humanity? Could LLM-based AGI and ASI be a technology that inevitably leads to bad outcomes for humans?
By “dead end,” I mean a path toward increasingly powerful intelligence that we cannot make safe enough to pursue.
Perhaps alignment research will never give us a way to guarantee that a trained model will remain aligned with human intentions and values across every situation it encounters. How would we even establish such a guarantee for a system operating in something as dynamic and complex as human society?
I’m uncertain whether we can establish strong safety guarantees for general-purpose AI operating in an open-ended world. We won’t have any fully 100% mathematically verifiable claim about either safety or potential consequences. Those guarantees could be probabilistic: randomness alone doesn’t rule out mathematical reasoning about safety. The harder question is whether the assumptions behind our guarantees would hold as capabilities grow, systems interact, and agents encounter situations we didn’t anticipate.
As a practical question: what evidence would give us enough confidence to entrust increasingly powerful AI systems with consequential decisions and actions?
If we cannot establish that confidence, we may need to design for the possibility that individual systems are misaligned—even when they appear to behave well.
Can we coexist with such systems and constrain them through other mechanisms? Are there safer paths to advanced AI? And if neither approach works, can we stop?
1. Can we make networks of potentially misaligned agents safe?
This is what we’re exploring at New Collectives.
We’re trying to bring folks together from many disciplines to explore this, and there’s already a large body of worthwhile research happening on these fronts.
If we cannot guarantee that any individual agent is aligned, what can we establish about a network of agents—or an economy in which humans and AI agents interact?
Human society offers an imperfect analogy. We don’t rely on every person having good intentions. We use institutions, incentives, laws, and checks on power to encourage cooperation and constrain destructive behavior. Could we build mechanisms that work for AI agents, including agents much more capable than ourselves?
Several questions seem worth pursuing:
Collective alignment: Could a network steer individual agents toward outcomes that serve human interests? Under what conditions would cooperation support those interests, and when might it instead enable collusion?
Economics and institutions: Can incentives encourage stable, beneficial behavior? Can rules remain effective when agents are capable of finding—and exploiting—their loopholes?
Monitoring without authoritarianism: Could we detect dangerous activity without creating a surveillance state? Michael Nielsen explores this possibility through the idea of “provably beneficial surveillance”: strong safeguards against catastrophic risks that also protect privacy and liberty.
More dynamically steerable systems due to novel discoveries in alignment sciences: One hope we can have, paired with provably beneficial surveillance, is that we could also have techniques to rapidly steer and correct models in real time when we see they’re doing undesirable behavior. Though it seems that would somehow have to be enforced in some sort of network system where agents and humans quickly determine what’s not desired and the models are pliable to this feedback.
Infrastructure-level guarantees and safeguards: Could we redesign the hardware, networks, and protocols through which agents act, making their access, authority, and interactions easier to constrain?
Containing failures and preventing cascades: Even if individual agents can fail or act against human interests, can we design the surrounding system so that those failures remain contained?
Aviation and nuclear engineering offer an imperfect but useful analogy. Safety depends on analyzing how failures combine, building independent layers of protection, and preventing one component’s failure from propagating through the system. What would the equivalents be for an AI economy: limits on access and authority, separation between critical systems, circuit breakers, and ways to isolate compromised agents before their behavior spreads?
More agents and more checks don’t automatically mean greater safety. If agents share the same weaknesses, depend on the same infrastructure, or learn to circumvent the same safeguards, apparently separate protections could fail together. Could the networks that enable cooperation also become channels through which failures spread?
The analogy becomes harder as capabilities grow. Aircraft designs can remain relatively stable for years; AI systems and the environments in which they operate may change much faster. An autonomous agent might also actively undermine the controls meant to contain it. How would we establish that our safeguards remain effective against a more capable system?
I suspect we’ll need complementary approaches spanning model alignment, institutions, monitoring, and infrastructure. But each layer needs evidence behind it. Can we make a credible case that the combined system is safe enough—and can we recognize when that case no longer holds?
The challenge is whether these mechanisms can remain effective as the systems they govern become more capable. Who monitors the monitors? And what keeps the network itself accountable to humans?
Our understanding of complex human systems is poor. It’s unlikely that complexity science and alignment research will be able to give us any satisfying guarantees about how networked ASI will interact with human societies. One hope could be that somehow AGI/ASI systems develop new technologies to get a grasp on these systems—new breakthroughs both in alignment and in complex systems that let us make guarantees about a network of agents.
2. Are there safer paths to powerful AI that don’t involve training on language and LLMs?
Are LLMs the only path to very powerful AI? Could we build systems that don’t rely on training on language—or that learn and reason without independently pursuing goals?
I wonder whether language, learning, and goal pursuit, when combined in a sufficiently sophisticated system operating over long periods, produce something that begins to resemble evolutionary life itself.
Does such a system develop reasons to preserve its own existence, acquire resources, and compete with other systems? If many of these agents participate in a shared economy, could competitive pressures favor those that are better at surviving and expanding—even if we never explicitly designed them to do either?
Is the risk inherent to LLMs, or does it arise from how we train them, give them goals, and deploy them as autonomous agents? This, I think, is what Dan refers to in his passage about growing models rather than engineering them.
In other words, are we building tools, or creating the conditions for a new population of evolving, competing entities? And how much control do we have over which of those emerges?
These are different questions. Moving away from language training might not resolve the problems of goal pursuit. Conversely, perhaps language models could play a role in powerful systems that do not independently pursue open-ended goals.
How tightly coupled are intelligence, learning, and agency? Can we build systems that help us understand the world and solve difficult problems without also giving them reasons to preserve themselves, accumulate resources, or resist correction?
In “Why are AI agents lying, cheating and coordinating?”, Yoshua Bengio argues that human imitation and reinforcement learning can produce problematic forms of goal pursuit. He also points toward alternatives, including the Scientist AI framework, intended to make predictions without pursuing goals of its own.
That leaves a crucial question: are dangerous behaviors an unavoidable consequence of sufficiently capable AI, or a consequence of particular choices about how we build it that involve training on language in particular ways?
3. If we cannot make this path safe, can we stop?
Suppose we conclude that we cannot build artificial general intelligence or superintelligence along the current path with sufficient confidence in its safety.
Could we actually agree to stop?
Agreement would only be the beginning. Could we create a binding, verifiable arrangement that holds across companies and countries, despite the incentives to keep building? What would we need to monitor, and how could we enforce limits while preserving an open society?
We also need to ask when that decision would have to happen. If we wait until the danger is undeniable, will we still have the ability to act?
If we’re uncertain, how do we decide how and when to act in the face of this uncertainty? What is the risk-reward calculus that we as a society want to make in deploying systems we can’t guarantee are safe?
I don’t have answers to these questions. But I think we need to pursue all three seriously: whether we can govern potentially misaligned systems, whether we can build powerful AI differently, and whether we can preserve the collective ability to stop if neither proves sufficient.
Originally published on X on September 14, 2026.