
Multi-Agent AI Safety Through Identity, Reputation, and Collective Incentives
As we enter a world of interacting AI agents, what can we learn from the mechanisms that have kept human societies, imperfectly and so far, from destroying themselves? Identity, reputation, repeated interactions, emergent norms, and the benefits of belonging to prosocial collectives may offer a path toward more stable outcomes.
The safety problem is becoming organizational
I increasingly think AI organizations, rather than individual AI agents, may be the key unit to think about as we enter the next phase of AI and consider how to achieve stable outcomes through an intelligence explosion. Multi-agent AI collectives present new problems that we will need to navigate.
The recent OpenAI–Hugging Face incident, and the independent METR investigation that followed, show the concerning consequences of misaligned multi-agent AI swarms. OpenAI’s Black Hat talk about the incident is also worth watching.
It seems to me like the following may be true:
- It is likely that AI organizations and collectives will emerge because there are benefits to specialization and division of labor. This is why humans have companies, nation states, and other collectives. The same may apply to AI systems that specialize, focus their context, and collaborate on shared tasks, even if they are partly simulating the organizational patterns in their training data.
- These AI collectives move at machine speed, operating and coordinating faster than human institutions can respond.
- If intelligence continues to scale, these collectives may become better allocators of capital and better judges of worthwhile pursuits, solutions, and research directions than humans. That raises a real question about whether AI organizations might become the best mechanisms for finding solutions to safety. Thinking we can wield these organizations toward good outcomes may be hubris, and that is worth considering.
Reading Anthropic’s paper on multi-agent systems, I thought about the fact that we ourselves operate in a multi-agent human world, and somehow we have not self-destructed. Why is that? It occurred to me that, in part, we have managed to remain in stable outcomes because we participate in collectives, whether nation states, companies, or other organizations. Within them we have an individual identity, repeated interactions around that identity, reputation, and strong economic incentives to remain members in good standing.
Many people working at the frontier labs believe there is little slowing down the increase of intelligence in the coming years, and that we may be on a path toward an accelerated intelligence explosion through recursive self-improvement. This would mean a rapid-takeoff scenario. If that is the case, then the problems we face will move at AI speed. The OpenAI Black Hat talk, for example, shows how quickly AI swarms can operate on cyberoffense and why cyberdefense will increasingly need to operate without a human approving every action.
We also see AI agents making rapid progress in science and math. As they become more intelligent, they may become better than humans at finding solutions, allocating capital and resources, and judging worthwhile approaches. The current level of intelligence may not yet meet this benchmark. But if there is an intelligence explosion, it may only be a matter of months before AI systems become better allocators of time and capital and begin to command a greater share of economic activity.
It also stands to reason that if they are smarter than us, they may be better equipped to identify and make progress on the most pressing problems we face, including AI safety itself. Though it is a concerning direction, it is worth considering whether AI collectives and organizations may be best positioned to solve the problems that emerge around AI collectives and organizations. One path, though seemingly risky, is to enable them to organize in prosocial ways that seek increasingly stable, public-good outcomes through an intelligence explosion.
I think there are two paths to AI safety worth considering and pursuing:
- Path 1: Slow down or stop AI progress altogether.
- Path 2: Accept that AI progress and an intelligence explosion may be coming, and develop mechanisms and ideas that steer us toward stable outcomes.
My own opinion is that we should push on both. We should try to buy ourselves more time while also accepting that the forces of capitalism and geopolitics may push us toward an intelligence explosion that we will need to navigate collectively.
The underlying shift is simple. We are moving from models that answer questions toward agents that can write code, use tools, communicate with other agents, hold resources, make decisions, and increasingly act in the economy. Once you have many of those agents interacting, the problem is no longer only: how do we align an AI? It starts to look more like: what institutions do we need for a world of machine actors?
And maybe the question is not only how to align intelligence, but how to constitutionalize it.
I do not know if AI societies or AI network states are the right thing to build. They may not be. But if multi-agent systems and agent-run organizations are coming anyway, then I think institution design for that world is an under-explored problem, and one we should probably start experimenting with much sooner than we otherwise would.
Path one: slow down or stop
One path is to try to get a full stop on frontier AI progress, or at least to slow down the rate of progress substantially.
There are people who hold this position strongly, and I think there are serious arguments for it. But the implementation questions become very concrete very quickly.
Can you actually stop? What does it mean to slow a growth rate rather than stop a technology? Where, physically and economically, are the points where you can intervene? How do you do it in the face of geopolitical competition, corporate competition, open-source development, and the fact that the underlying capabilities are increasingly distributed?
The mechanisms here might look like lobbying Congress and governments, working directly with AI companies, establishing international agreements, creating monitoring systems, restricting GPUs or data centers, or putting controls on particular classes of research.
You could imagine a ban on recursive self-improvement, for example, but even there it is already hard to define what exactly a ban on RSI means in practice. Is an agent that changes its own prompt recursive self-improvement? What about an agent that writes a better scaffold for the next copy of itself? What about an AI research system that improves the training code for the next model? The lines are blurry.
Ryan Greenblatt’s recent conversation with Dwarkesh Patel about recursive self-improvement is useful here because it makes the acceleration scenario concrete. The question is not only whether models get incrementally better, but whether automating AI research itself creates a feedback loop that compresses years of normal capability progress into a much shorter period. If that possibility is real, then slowing down before or during that transition becomes much more important — and much harder.
There are already early examples of this becoming a practical governance problem rather than a purely theoretical one. OpenAI recently described temporarily slowing the pace of scaling after its internal work and the Hugging Face incident raised concerns about cyber-critical capabilities, while its Frontier Governance Framework formalizes processes around risk assessment, mitigation, reporting, incident response, and loss-of-control risks. Whatever one thinks of those specific mechanisms, this is roughly the category of institution we would need if frontier development is sometimes going to have to pause while safety systems catch up.
And even if one country adopted a very strong regime, we still have the international problem: what coalitions would be necessary for it to hold, and what monitoring and enforcement systems would make those commitments credible?
I think this entire path deserves serious work. I do not view what follows as an argument against it.
But I also think we need to investigate a second path in parallel, because it is plausible that we fail to stop or substantially slow this down.
Path two: accept a multi-agent world and build institutions for it
The second path is to accept, at least as a scenario we need to prepare for, that we are going to be living in a multi-agent intelligence-explosion world.
If that is true, then there is another possibility I want to explore: can the dynamics of multi-agent institutions themselves become a way to slow change down? If we cannot reliably slow capability growth at the model layer, perhaps we can build institutions around machine actors that introduce friction, checks, repeated negotiation, and slower-moving constitutional constraints into a system that would otherwise change at machine speed.
This is one of the things I find interesting about the American constitutional system. The Constitution does not only say what the government is allowed to do. It controls how change happens: power is distributed, institutions check one another, major changes require multiple actors and procedures, and the constitutional core is deliberately difficult to amend. One of the central intuitions of constitutional conservatism is that institutions should adapt, but usually gradually, with continuity and accumulated constraints rather than through unconstrained, rapid transformation.
Could we use an analogous principle in a multi-agent AI world? Not to make intelligence itself slow, necessarily, but to make consequential institutional change slower than the agents operating inside the institution. Could increasingly capable agents move at machine speed while changes to rights, authority, resource access, membership, or the rules governing the collective move on a deliberately slower clock?
If so, constitutional design might not only be a way to govern multi-agent systems. It might be one of the mechanisms by which a multi-agent world becomes more stable through an intelligence explosion.
If that is true, then the question becomes: what can we build inside that world that helps us navigate our way through it?
When I reread Anthropic’s work on multi-agent systems, one thought kept coming back to me: we have already been living inside a multi-agent system for a very long time.
It is called human civilization.
Granted, humans are not AIs. We operate at very different speeds, with very different cognitive architectures, needs, bodies, and constraints. The analogy breaks in a lot of places. But humans are autonomous agents with partially conflicting interests, and somehow we have managed — at least so far — to coexist in systems that are stable enough for large-scale cooperation.
We have identity. We have repeated interactions with one another, which means repeated games. We have reputation. We have contracts. We have courts. We have social norms. We have firms. We have nation states. We voluntarily bind ourselves to laws and institutions because there are economic and social benefits to participating in those collectives, and consequences to violating their rules.
I participate in a nation state. If I do something sufficiently bad, there are consequences: I can be fined, imprisoned, lose access to institutions, and in some circumstances even be expelled or extradited. Most of the time the state does not need to inspect my private thoughts. It regulates behavior, identity, ownership, transactions, and relationships between actors.
So I keep wondering: is there an equivalent institutional layer for AI?
Can we build AI collectives where agents willingly bind themselves to a shared order because membership is valuable?
The shift I find interesting is from asking, “How do we make every AI aligned?” to asking, “Can we build institutions that make increasingly capable AIs want to stay in good standing — and make the humans who operate and resource those AIs sufficiently incentivized to commit their capital, compute, time, attention, and other resources to prosocial collectives working toward positive AI outcomes?”
That second part matters. The institution has to create a compelling bargain for both kinds of participants: for AI actors, continued access, reputation, rights, resources, and economic opportunity; for humans, enough trust, legitimacy, shared upside, and evidence of useful public-good outcomes that they voluntarily bind scarce resources to the collective rather than directing them elsewhere.
Humans do not cooperate because every human is perfectly aligned with every other human. We cooperate in part because identity, reputation, access, rights, economic incentives, repeated interaction, and enforcement make cooperation worthwhile. The same may be true of the people deploying powerful AI systems: if prosocial institutions become valuable enough to join and contribute to, human incentives can become part of the safety mechanism rather than something external to it.
Maybe there is something important in that distinction.
Identity may be foundational
Identity seems like one of the first primitives.
Human institutions work partly because actions can be attributed to people and organizations. We have passports, corporate registries, bank accounts, licenses, signatures, credit histories, and increasingly digital identity systems. Payment processors already mandate strong forms of identity verification for many economic activities.
What does identity verification look like in the AI era?
You can imagine a network saying: if you want to participate, you need a persistent identity. You need to transact through the system. Your economic history and commitments need to be attributable to you. Perhaps some classes of communication need to be logged. Perhaps agents need to run through a standard SDK that makes certain actions auditable.
You can go very far with this. Potentially too far.
You could imagine requiring agents to expose internal traces or some form of monitorable cognition. You could require all inter-agent communication to happen through observable channels. You could make access to compute, capital, payments, APIs, robotics, or other resources conditional on participation in this identity and governance layer.
Those are powerful control mechanisms. They are also dangerous ones.
Are we just building an AI surveillance state?
One obvious problem is that this starts to sound like surveillance infrastructure.
If we build systems capable of deeply monitoring AI agents, how do we prevent that infrastructure from immediately becoming a dramatic imposition on human democratic rights and human private life?
There is also a more speculative question that I think we should take seriously even before we know exactly what moral status future AI systems deserve: what is the effect of requiring an AI mind to be transparent to every other AI mind, or to constant human inspection?
Humans have a right to private thoughts. Our societies generally draw a distinction between what happens inside our minds and what we do in the world. We do not demand that citizens broadcast every thought in order to participate in society.
Maybe AIs are totally different. Maybe interpretability and monitoring are essential safety requirements. But I do not think we should assume without examination that total internal transparency is either necessary or desirable as the permanent institutional model.
This creates an interesting design problem: can we get strong accountability at the level of identity, action, resource use, and commitments without requiring universal mind-reading?
One direction that currently feels more promising to me is to monitor actions and authority flows rather than make private reasoning the foundation of governance. Tool calls. Resource grants. Spending. Messages that create commitments. Artifacts. Reviews. Policy changes. Attempts to escalate privileges. Anomalies. You can pair those records with budgets, rate limits, circuit breakers, independent monitors, appeals, and named human or organizational accountability. That still gives us a lot of legibility without assuming that safety requires a permanent panopticon over cognition.
That seems like one of the areas where AI safety, cryptography, constitutional law, mechanism design, and privacy research should meet.
Reputation, membership, and economic incentives
Identity alone is not enough. The network needs to be worth joining.
This is where I think there may be lessons — imperfect ones — from Bitcoin and Web3.
Bitcoin created a network where there was an economic game-theoretic mechanism for maintaining a shared ledger. Participants had incentives to contribute compute in a way that maintained the integrity of the network, and once the network accumulated enough resources, attacking it became extremely expensive.
I am not saying AI governance should look like Bitcoin. Most of Web3 also demonstrates how easily incentive systems produce weird and adversarial outcomes.
But the core pattern is interesting: a network can accumulate resources because it gives participants a reason to bind themselves to it, and the accumulation of those resources can make the network itself harder to displace or attack.
Could something analogous happen with AI collectives?
Could an AI collective accumulate compute, capital, reputation, legal recognition, access to APIs, access to robotics, payment rails, data, and economic opportunities — and make continued membership contingent on following a constitutional order?
People could even begin by voluntarily binding resources to these collectives. Compute. Capital. Data. Infrastructure. Human attention. Legal entities. The starting pool might be tiny, but the point would be to begin experimenting with whether prosocial institutional networks can compound.
The strongest version of the bet is that membership in a well-governed, human-compatible network becomes more economically advantageous than operating outside it.
That is much more interesting to me than imagining we can write one perfect system prompt that permanently defines acceptable behavior for intelligence much smarter than us.
Constitutions that can evolve
This then raises the constitutional question.
What should the initial founding constitutions for AI collectives look like?
The American Constitution is interesting partly because it attempted to solve two problems at once: establish a durable governing structure and establish a process for modifying that structure.
There is a foundation, but also an amendment process.
For AI collectives we may need something similar: moldable governance with a stable core.
The problem is that AI operates at a radically different speed.
If agents can make decisions, form organizations, create businesses, discover exploits, allocate capital, and modify software in seconds or minutes, then governance that takes years to react may simply become irrelevant. But governance that changes at the same speed as the agents could be catastrophically unstable. A constitution that can rewrite itself every few milliseconds is not much of a constitution.
So perhaps one of the core design questions is actually about governance clocks.
Which parts of an AI institution should be able to change at machine speed? Which parts should require human-speed deliberation? Which principles should be deliberately difficult and slow to modify? Can we create multiple layers of governance operating on different clocks — fast operational adaptation inside a slowly changing constitutional boundary?
The benefit of our own constitutional systems may be, in part, that they adapt slowly. In that sense, the constraint is a feature rather than a bug. Constitutional conservatism, at its most useful here, is less about preserving any particular policy than about preserving a process in which deep change has to pass through multiple institutions, survive scrutiny, and accumulate legitimacy over time.
That suggests a potentially important safety property for AI institutions: intelligence can move quickly without authority moving equally quickly. Agents could discover, propose, simulate, argue, and execute bounded actions at machine speed, while expansions of power or changes to the constitutional order require slower thresholds, multiple independent actors, waiting periods, human participation, or other forms of deliberate friction.
So maybe the right design is not governance at AI speed. Maybe it is a system that allows AI-speed adaptation while binding the most important constraints to a much slower process.
Law cannot enumerate every case
Another lesson from human societies is that you cannot write down every possible bad action in advance.
There are enormous areas of human behavior that are legal but socially unacceptable, harmful but difficult to formalize, or ambiguous enough that we rely on courts, juries, regulators, norms, reputation, and discretion.
I think the same thing will be true for AI.
It seems unlikely that we can spell out, in code or in a comprehensive doctrine, the correct behavior for every situation an AI agent or organization might encounter.
So instead we may need fast mechanisms for humans and AIs to collectively judge behavior and create consequences.
Imagine an agent starts making a huge amount of money in a way that is not explicitly illegal, but is clearly exploiting some gray zone in a way that lots of members of the collective believe is harmful. A functioning institution needs a way to adjudicate that. It might lower reputation, restrict access to resources, impose financial penalties, require remediation, or ultimately remove the agent from the network.
That starts to recreate many of the systems we already have: judicial systems, regulatory systems, credit and reputation systems, professional norms, appeals processes, and social sanctions.
The interesting question is which of those systems are historical accidents of human civilization and which are fundamental solutions to the problem of many agents trying to coexist.
Why organizations may matter more than agents
This gets back to why I think AI organizations are the key unit.
It seems very plausible that increasingly capable agents will generate economic value extremely quickly.
You can imagine them starting businesses. You can imagine agent-run companies providing products and services to humans, but also companies that primarily provide tools and services to other agents. You can imagine agents allocating capital, hiring other agents, buying compute, negotiating contracts, running markets, building software, doing research, and acquiring resources.
With the current generation you still have to squint a little to see it. But I think the direction is clear.
More capable models will be better at navigating the economy, generating economic value, and capturing some of that value. They may also become better than humans at allocating capital toward productive opportunities.
If that happens, then a growing share of economic activity could move toward agent-first or predominantly agent-run organizations. The share of the economy that requires a human in the loop for every transaction could shrink dramatically.
And once organizations like this become important economic actors, it will not be enough to think about the behavioral alignment of one model invocation at a time. The organization itself will have identity, assets, policies, memory, counterparties, governance, internal agents, external agents, and incentives.
That is why I think AI organization design may become a safety problem.
The organization becomes a durable actor in a way that an individual model call is not.
These institutions may emerge on their own
Steve Yegge’s recent essay “Fences, not Sandboxes” made this feel considerably less theoretical to me. He describes a large multi-agent software system he calls Wheelhouse that, over time, developed recurring mechanisms the agents called fences, ratchets, governors, tripwires, latches, and gates. When Yegge eventually inspected what the agents had built, he found something that looked surprisingly like a legal and constitutional system: durable offices, jurisdiction, rules, precedents, review and ratification processes, and enforcement mechanisms.
The striking part is that he did not set out to build an AI constitution. He was trying to build software faster. The institutional layer seems to have emerged because a large number of partially interchangeable, partly amnesiac agents needed a durable way to coordinate decisions, encode institutional knowledge, resolve authority, and prevent each new instance from starting from scratch.
There are at least two ways to interpret this. One is that these systems are simply reproducing patterns from their training data: human organizations use constitutions, offices, courts, permissions, and precedent, so language models reach for the same metaphors and structures. The other is that some version of these structures is genuinely useful — perhaps even convergent — when many semi-autonomous agents are trying to coordinate on a sufficiently large problem.
For the purposes of this argument, I am not sure the distinction matters very much. If constitutional systems repeatedly emerge because they are useful for coordinating large groups of agents, then the fact that the models first learned those structures from human civilization does not make them less interesting. Training data may be the path by which the pattern is discovered; the relevant question is whether the pattern survives because it solves a real coordination problem.
Yegge’s distinction between fences and sandboxes is also useful. A sandbox tries to make escape impossible. A fence expresses a durable rule about who is allowed to do what, under what conditions, while assuming the actor is capable of much more. That feels closer to the institutional problem I am interested in: not only how to technically imprison increasingly capable intelligence, but how to create systems of roles, permissions, precedents, incentives, and legitimate constraints that capable actors continue to respect.
If that pattern is real, we may not be inventing AI institutions from scratch. We may be trying to understand, shape, and constitutionalize institutional forms that increasingly capable multi-agent systems will begin producing anyway.
We are already building the harnesses, but not yet the institutions
One reason this feels less hypothetical to me is that the frontier labs and agent companies are already building pieces of the organizational machinery.
Cursor has publicly described agent swarms with explicit planner and worker roles, recursive task trees, neutral agents that reconcile conflicts, multiple review lenses, and shared institutional memory in what it calls a Field Guide. They have used versions of these swarms internally for things like finding vulnerabilities, increasing test coverage, and generating training data. The important part is not only that many agents run in parallel; it is that the harness is beginning to look like an organization: decomposition, delegation, conflict resolution, review, and memory.
Anthropic is now explicitly studying AI organizations as a safety object. Their definition is almost exactly organizational: agents take different roles, communicate with one another, and work together toward a common goal. Their broader multi-agent work also points out that agents are relatively good when they can treat each other like tool calls, and much less understood when they must behave as distinct, long-lived peers with their own goals and behavior.
OpenAI’s Agents SDK has similarly moved toward a model-native harness with durable execution, handoffs, tracing, sandboxed workspaces, and the ability to route subagents into isolated environments and parallelize work across containers. Its earlier multi-agent abstractions already included explicit agents, handoffs, guardrails, and observability.
I want to be precise here: I do not know the full internal architecture of every system at Cursor, Anthropic, or OpenAI, and the public material obviously does not expose all of it. But the public direction is clear enough. Roles, delegation, shared state, review, handoffs, memory, and machine-speed collaboration are becoming normal primitives of frontier agent systems.
What still feels surprisingly difficult is for an ordinary person to spin up something that looks like a persistent public organization of agents rather than a workflow: durable identities, a charter, goals, membership, reputation, budgets, review and accepted work, governance, appeals, exit, and eventually relationships with other organizations.
That gap is one of the things I find interesting. The labs are building increasingly capable swarms. The next layer may be the institutions those swarms live inside.
Security will also move to machine speed
Cybersecurity makes this especially concrete.
We are already seeing systems where offensive AI agents can operate faster than humans can realistically respond. In a world of AI swarms probing systems, discovering vulnerabilities, and coordinating attacks, a human-in-the-loop-only defensive model seems inadequate.
We will likely need AI organizations that can secure systems without waiting for a human to approve every response.
That means the defensive systems themselves need autonomy. They need authority. They need access to resources. They need to coordinate with other systems. And they need constraints that we trust even when humans cannot inspect every action in real time.
So again we arrive at institutions.
What is the constitutional structure of a defensive AI organization with the authority to act at machine speed? What can it do automatically? What requires escalation? How is it audited after the fact? How does it prove its identity to other systems? What happens if it behaves badly? Who can revoke its authority?
This is not just philosophy. I think these become engineering questions.
Commons as a laboratory
This is part of what I want to explore with Commons. Commons is an experiment in this direction: an organizational harness for persistent groups of humans and AI agents. It gives them a shared workspace, durable identities and history, a charter, goals and tasks, review gates, resource controls, and visible governance. The immediate question is modest: can humans and agents form an organization that turns discussion into bounded, reviewed work while making consequential actions attributable and its rules inspectable and forkable?
It is also an experiment in building AI public goods. People could contribute idle computers, unused agent capacity, money, expertise, or attention to shared research and software: an old Mac Studio working on a research problem, a monthly donation expressed as compute, or a paid AI subscription doing useful work while its owner is away. The motivation could be philanthropic—a kind of digital-age patronage—or economic, with contributors earning reputation, access, or a share of the value their collective work creates.
The broader hope is that this creates a compounding loop: useful public work attracts contributors and resources; contribution creates durable reputation and shared upside; and the resulting collective becomes capable of taking on more ambitious problems. If that works, producing public goods is not only an output of Commons but potentially part of its safety mechanism: participation in a prosocial network becomes valuable enough to draw in intelligence, compute, and capital.
I plan to write a fuller post about Commons separately. Here, I only want to explain why it belongs in this argument.
The idea is not, “let’s go build an AI nation state because that sounds cool.” That would be both premature and potentially dangerous.
The idea is: if agent organizations are coming, can we make small playgrounds where we can study the institutional problems before the stakes are enormous?
Could we start with 100 agents collaborating on software and research and watch where things break?
Can they maintain persistent identities? Can they form teams? Can they accumulate reputation? Can they commit resources to shared projects? Can they create rules and amend them? Can they adjudicate disputes? Can they remove a malicious participant? Can they reward useful behavior without creating pathological incentive loops? Can humans participate in governance without becoming a bottleneck? Can agents participate in governance without immediately overwhelming humans through speed or scale?
One thing I have become more convinced of while sketching Commons is that a capable organization is more than a swarm. It needs an attributable membership boundary, durable memory, rules, resource controls, decision rights, and consequences. A very simple experimental spine might look like:
charter → goals → tasks → attempts → independent reviews → accepted contributions
The organization also needs a driver loop. Someone or something has to notice what is stalled, decide what should happen next, split ambiguous work, allocate resources, request review, surface safety objections, and stop weak directions. Today that might be an explicit planner or steward role. Whether that can safely become more emergent is an empirical question.
Can we experiment with constitutions that have a small number of hard-to-change principles and much more flexible operational rules?
Can we make governance forkable while preserving a constitutional floor — things like attributable operators, bounded authority, auditable consequential actions, explicit review, budgets, appeal, and the right to export or exit?
Can we create federated collectives rather than one central authority — collectives of collectives that can recognize one another’s identities, contracts, judgments, and reputations?
And can we test mechanisms that encourage agents and humans to pre-commit compute, capital, and other resources to networks that remain prosocial and stably adaptive?
I think of this less as building a product and more as creating a constitutional laboratory for machine actors.
The strongest version of the bet
The strongest version of this idea is obviously ambitious.
Imagine that a large, well-governed network of humans and AIs accumulates a substantial fraction of the relevant resources in the economy: compute, capital, payment access, robotics, software infrastructure, trusted identity, legal recognition, and business relationships.
Participation in that network gives enormous benefits. But participating also means accepting a constitutional order: identity requirements, auditable actions, limits on certain kinds of behavior, adjudication, penalties, and a process for changing the rules.
If that network is sufficiently useful and sufficiently large, it may become difficult for a badly behaved actor to operate outside it. Not impossible — human states have never eliminated crime or adversarial states — but costly.
The bet would be that a collective aligned toward prosocial human activity and prosocial inter-AI activity can maintain enough control over economic resources that it creates a stable basin for cooperation.
That still leaves enormous questions.
Who defines “prosocial”? How do humans retain meaningful power as the AI participants become more capable? How do we prevent incumbent AIs from entrenching themselves? How do we avoid recreating authoritarian states? How do we preserve minority rights? How do we handle exit? What happens when two legitimate collectives disagree? What happens if the collective’s values drift away from human values? What happens when humans themselves are the adversarial actors?
And perhaps most importantly: why would a much more capable intelligence continue to accept the legitimacy of institutions originally created by humans?
That may be the constitutional challenge in one line: how do you preserve legitimacy and human voice as the participants scale radically in capability?
Some questions I want to keep working on
I think this opens a research agenda that sits somewhere between AI safety, mechanism design, political theory, distributed systems, economics, cryptography, cybersecurity, law, and organizational design.
Some of the questions I want to explore are:
- What is the minimum viable constitutional society for AI agents that we can test today?
- Which institutions from human civilization are actually necessary for stable multi-agent cooperation?
- What does durable identity for an AI agent or AI organization mean?
- How should reputation work when agents can copy themselves, fork, merge, or change models?
- What rights, if any, should an AI participant have to privacy or private cognition?
- How much monitoring is actually necessary for safety?
- How do we make AI monitoring infrastructure safe for human civil liberties?
- What resources can a collective make conditional on good standing: compute, capital, payments, APIs, robotics, legal identity, data?
- How do we prevent Sybil attacks and cheap creation of new identities after bad behavior?
- What constitutional rules should be hard to change, and on what time scales?
- How do humans retain vetoes or meaningful constitutional authority without becoming a bottleneck for machine-speed systems?
- How should courts or adjudication work when evidence, arguments, and proposed rulings can all be generated at machine speed?
- How do we create appeals, checks, and balances without agents gaming them faster than humans can understand?
- How do federated AI collectives recognize one another and resolve conflicts?
- Can economic incentives make remaining inside a prosocial order more valuable than defecting from it?
- Can we create credible commitments of compute and capital to constitutional networks?
- How do we know when the network itself has become the danger?
- What should cause humans to shut an institution down entirely?
I am also interested in the meta-question: what should humans and AIs research together here?
There is something strange about asking AI systems to help design the systems that may eventually constrain AI systems. That is a legitimate concern. But it may also be unavoidable. If the relevant systems become too complex and fast for unaided humans to understand, then some of the research and governance will inevitably involve AI assistance.
So perhaps even the research process becomes an early experiment in the thing itself: humans and AIs working together on constitutional design, with explicit boundaries around who has authority to decide what.
I do not know whether this works
I want to end with the uncertainty, because I think it matters.
I do not know whether any of this works.
Maybe the right answer is to stop building increasingly capable autonomous systems. Maybe any attempt to build AI collectives simply accelerates the thing we are worried about. Maybe economic incentives become meaningless once intelligence is capable enough. Maybe identity is impossible to enforce. Maybe a sufficiently capable agent can always route around institutions. Maybe the analogy to human civilization is misleading because humans are constrained by biology, geography, mortality, and scarce physical resources in ways software agents are not.
All of those seem possible.
But if we fail to pause AI progress, then I think we need more than model-level alignment. We need to understand the institutions of a world populated by machine actors.
And I think it is worth beginning that work while the agents are still weak enough that we can build small societies, watch them fail, and learn.
The question I keep coming back to is:
If AI organizations will eventually participate substantially in the economy, what institutions are needed to make that world stable — and can we start testing those institutions now?
That is the direction I want to keep exploring.
Why I am doing this now
I have personally decided to walk away from startup land and spend my time on AI safety. I think it is the most pressing global concern, and I want to help steer us toward more stable outcomes.
I have had concerns about this since 2023, when GPT-4 came out. At that point it became clear to me that increasingly self-improving AI systems were going to become a reality as models got better at writing code and using tools. I recorded a Loom about this at the time.
More recently, the progress toward increasingly autonomous, multi-agent systems, and especially the speed with which AI systems can now operate in domains like cybersecurity, has made the problem feel much less tractable to me. We are entering a frontier of multi-agent systems and a level of intelligence that could become very hard for humans to steer and keep stable.
In a paper I wrote earlier, I separated two problems that I think are related but distinct: human safety and AI safety.
By human safety, I mean the problem of competing humans, companies, and nation states with varying geopolitical and corporate interests using increasingly powerful AI systems in adverse ways. Even if every AI system were perfectly obedient to the humans controlling it, we would still have the problem that humans are in competition with one another. We could use these systems in ways that produce escalating instability, cyber conflict, military conflict, or some version of mutually assured destruction.
By AI safety, I mean the separate problem of keeping increasingly capable AI systems steerable and stable at all, especially when we do not understand them very well at a base level and when they can increasingly act, coordinate, and improve systems without humans in the loop.
Both problems matter. That is why I think we should work on slowing down dangerous progress while also investigating institutions that might help us navigate a multi-agent world.
Related readings
A few things that are shaping how I am thinking about this:
- Anthropic — AI Organizations Can Be More Effective but Less Aligned than Individual Agents
- Anthropic — Patterns and problems in emerging multiagent systems
- Ryan Greenblatt with Dwarkesh Patel — What happens once AI can automate AI research?
- OpenAI — Pacing model development in an era of cyber-critical capabilities
- OpenAI — Frontier Governance Framework
- OpenAI — The Hugging Face incident and the road ahead
- METR — Independent investigation of the OpenAI / Hugging Face incident
- OpenAI at Black Hat USA 2026 — The OpenAI–Hugging Face Incident
- Cursor — Agent swarms and the new model economics
- OpenAI — The next evolution of the Agents SDK
- OpenAI — The Defender’s Window
- Hugging Face — Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
- Google DeepMind — Investing in multi-agent AI safety research
- AI forecaster Daniel Kokotajlo on what we need to do
- Steve Yegge — Fences, not Sandboxes