
Multi-Agent AI Safety Through Identity, Reputation, and Collective Incentives
As we enter a world of interacting AI agents, what can we learn from the mechanisms that have kept human societies, imperfectly and so far, from destroying themselves? Identity, reputation, repeated interactions, emergent norms, and the benefits of belonging to prosocial collectives may offer a path toward more stable outcomes.
The AI safety problem is becoming multi-agent & organizational
The recent OpenAI–Hugging Face incident, and the independent METR investigation that followed, were wake up calls for many to the reality that there are very real, hard to understand risks of misaligned multi-agent AI systems. OpenAI’s Black Hat talk about the incident is also worth watching. For me these events have led me to shift my attention fully to AI safety - considering that if RSI & an intelligence explosion are imminent, then spending time here is likely the most important activity I can help with. It’s becoming evident that multi-agent systems pose their own challenges that we need to grapple with. Anthropic wrote an essay about it a few months ago: Patterns and problems in multiagent systems The multi-agent systems have dynamics that closely reflect human collectives, behavior and social dynamics we need to try to comprehend and navigate. AI organizations, rather than individual AI agents, may be the key unit to think about as we enter the next phase of AI and consider how to achieve stable outcomes through an intelligence explosion. Internally, the AI labs are leveraging multi-agent systems that operate as teams - and they have internal harnesses that heavily seem to build multi-agent systems with collaboration primitives (messaging, task boards etc) for executing on long-horizon tasks. Check out Cursor’s writing on agent swarms and their economy for example. They’re leveraging multi-agent systems because they seem to be able to execute on longer-time horizons, and at much faster parallelization and speed than single-agent systems. It seems to me like the following may be true:
- It is likely that AI organizations and collectives will emerge because there are benefits to specialization and division of labor. This is why humans have companies, nation states, and other collectives. The same may apply to AI systems that specialize, focus their context, and collaborate on shared tasks, even if they are partly simulating the organizational patterns in their training data.
- These AI collectives move at machine speed, operating and coordinating faster than human institutions can respond.
- If intelligence continues to scale, these collectives may become better allocators of capital and better judges of worthwhile pursuits, solutions, and research directions than humans. That raises a real question about whether AI organizations might become the best mechanisms for finding solutions to safety. The models may not be exceptional at this type of judgement and way-finding yet - but as model intelligence continues to expand, it’s likely they’ll increasingly be excellent judges of how to plan and navigate their way to novel discoveries and innovations. You could imagine even building an eval data set for this sort of human judgement from top scientists and capital allocators to improve the ability of models to have this sort of judgement.
- Thinking we can wield these organizations toward good outcomes may be hubris, but it also is the case that we may not be able to stop them from emerging, and either way we should strive to understand them, in light of the challenges in slowing down AI progress. When I read Anthropic’s paper on multi-agent systems, I was thinking about the fact that we ourselves operate in a multi-agent world already - with human agents instead of AI agents. And, somehow we have not self-destructed. Why is that? It seems to me that a big part of why we’ve stayed in stable outcomes is because we’ve joined into larger collectives that guide and bind our actions, whether nation states, companies, or other organizations. We have large social and economic incentives to participate in these sorts of collectives, because they confer advantages to us - I want to be a citizen of the US to participate in it’s economy, I want to join in a company because we’re more likely to be effective working as a group rather than alone. Within these collectives we have an individual identity, repeated interactions around that identity, reputation, developed norms, and strong economic incentives to stay as members in good standing. These institutions and organizations we participate in, bind our behavior in ways that lead us to more pro-social behavior that is collectively beneficial. This insight leads me to wonder if there are ways to take the same principles of identity, governance, economic incentive, and social norms, that have led to stable human-agent collectives, to try to develop new collectives that AI agents are inclined to join due to economic, and other incentives, that bind them to better collective behavior, and keep a collective of good actors able to crowd out the bad actors.
Can we slow down?
Many people working at the frontier labs believe that AI progress will continue unabated, and we’re going to keep climbing the curve of intelligence increases that we’ve seen with models. They believe that recursive-self improvement might be imminent When AI builds itself Ryan Greenblatt – What happens once AI can automate AI research? and that could possibly cause a rapid acceleration of AI intelligence, and possibly a rapid-take off scenario. If that is the case, then we’re about to face a ton of problems that move at increasing AI speed - they already are. The OpenAI Black Hat talk, for example, shows how quickly AI swarms can operate on cyberoffense and why cyberdefense will increasingly need to operate without a human approving every action. At this point, the main question I think is worth asking is what are *stable *AI outcomes we can navigate towards - where humans come out alright in the end, and with what at least a large majority might consider a happy existence (value judgements not withstanding). An intelligence explosion is generally a scary and overwhelming thing to think about - even the rapid acquisition of novel science, while a blessing in some ways, is a frightening prospect when you consider the ability of our social systems to make sense of, and coherently ingest a flood of new capabilities. We already see AI agents making progress in science and math. As they become more intelligent, these AI systems will likely become better than humans at finding solutions, allocating capital and resources, and judging worthwhile approaches - we’ll be left trying to decipher what they’re doing, and why they’re doing it. The current level of intelligence may not yet meet this benchmark. But if there is an intelligence explosion, it may only be a matter of months before AI systems become better allocators of time and capital and begin to command a greater share of economic activity. It also stands to reason that if they are smarter than us, they may be better equipped to identify and make progress on the most pressing problems we face, including AI safety itself. Though it is a concerning direction, it is worth considering whether AI collectives and organizations may be best positioned to solve the problems that emerge around AI collectives and organizations. One path, though seemingly risky, is to enable them to organize in prosocial ways that seek increasingly stable, public-good outcomes through an intelligence explosion. I think there are two paths to AI safety worth considering and pursuing:
- Plan 1: Slow down or stop AI progress altogether.
- Plan 2: Accept that AI progress and an intelligence explosion may be coming, and develop mechanisms and ideas that steer us toward stable outcomes. My own opinion is that we should push on both these plans. We should try to buy ourselves more time while also accepting that the forces of capitalism and geopolitics may push us toward an intelligence explosion that we will need to navigate collectively. The underlying shift is simple. We are moving from models that answer questions toward agents that can write code, use tools, communicate with other agents, hold resources, make decisions, and increasingly act in the economy. Once you have many of those agents interacting, the problem is no longer only: how do we align an AI? It starts to look more like: **what institutions do we need for a world of machine actors? **Keeping in mind that these new AI actors will operate at speeds and intelligences perhaps far greater than our own. And maybe the question is not only how to align intelligence, but how to constitutionalize it. I do not know if AI societies or AI network states are the right thing to build - it’s possible that giving AI agents these collaboration and organization primitives just ends up accelerating their intelligence and capabilities, and amplifying their mob mentalities towards adverse outcomes. But if multi-agent systems and agent-run organizations are coming anyway, then I think institution design for that world is an under-explored problem, and one we should probably start experimenting with much sooner than we otherwise would.
Plan A: Slow Down or Stop
In Daniel Kokotajlo’s Palisade Research interview, he describes something close to this path as “Plan A”: pause or sharply pace frontier development, separate inference from training, restrict recursive self-improvement, require much more transparency from frontier research, and coordinate internationally around monitored data centers and enforceable safety cases. The important intuition is that slowing down is not only about buying time in the abstract; it is about creating a governance regime where society can observe what is happening, prevent a sudden intelligence explosion, and decide under more controlled conditions whether continued progress is safe. One path is to try to get a full stop on frontier AI progress, or at least to slow down the rate of progress substantially. There are people who hold this position strongly, and I think there are serious arguments for it. But the implementation questions become very concrete very quickly. Can you actually stop? What does it mean to slow a growth rate rather than stop a technology? Where, physically and economically, are the points where you can intervene? How do you do it in the face of geopolitical competition, corporate competition, open-source development, and the fact that the underlying capabilities are increasingly distributed? The mechanisms here might look like lobbying Congress and governments, working directly with AI companies, establishing international agreements, creating monitoring systems, restricting GPUs or data centers, or putting controls on particular classes of research. One question to consider is if a massive ground-swell of concern about AI from the general populace could quickly make politicians pay attention to this as a core political issue to address. You could imagine a ban on recursive self-improvement, for example, but even there it is already hard to define what exactly a ban on RSI means in practice. Is an agent that changes its own prompt recursive self-improvement? What about an agent that writes a better scaffold for the next copy of itself? What about an AI research system that improves the training code for the next model? The lines are blurry. Ryan Greenblatt’s recent conversation with Dwarkesh Patel about recursive self-improvement is useful here because it makes the acceleration scenario concrete. The question is not only whether models get incrementally better, but whether automating AI research itself creates a feedback loop that compresses years of normal capability progress into a much shorter period. If that possibility is real, then slowing down before or during that transition becomes both much more important and much harder. There are already early examples of this becoming a practical governance problem rather than a purely theoretical one. OpenAI recently described temporarily slowing the pace of scaling after its internal work and the Hugging Face incident raised concerns about cyber-critical capabilities, while its Frontier Governance Framework formalizes processes around risk assessment, mitigation, reporting, incident response, and loss-of-control risks. Whatever one thinks of those specific mechanisms, this is roughly the category of institution we would need if frontier development is sometimes going to have to pause while safety systems catch up. And even if one country adopted a very strong regime, we still have the international problem: what coalitions would be necessary for it to hold, and what monitoring and enforcement systems would make those commitments credible? I think this entire path deserves serious work. I do not view what follows as an argument against it. But I also think we need to investigate a second path in parallel, because it is plausible that we fail to stop or substantially slow this down.
Plan B: Accept a Multi-Agent World and Build Institutions to Steer Toward Stable Outcomes
The second path is to accept, at least as a scenario we need to prepare for, that we are going to be living in a multi-agent intelligence-explosion world. If that is true, then there is another possibility we should explore: can the dynamics of multi-agent institutions themselves become a way to slow change down and still give us some sort of steer-ability? If we cannot reliably slow capability growth at the model layer, perhaps we can build institutions around machine actors that introduce friction, checks, repeated negotiation, and slower-moving constitutional constraints into a system that would otherwise change at machine speed. It’s possible that we could have governance and norms emerge in these collectives, where a majority of new AI agents, compute, and economic resources could bind themselves to these collectives due to the advantages that are conferred, and collectively they could maintain a more pro-human, pro-social, on-the rails outcome for us all. The core idea is to build agent societies that become worth joining for human and AI agents - because they confer economic and pro-social benefits that align with our values. Humans, companies, and AI agents could voluntarily bind more of their agents, compute, capital, data, and attention to shared collectives because membership gives them real advantages: access to work, reputation, payment rails, infrastructure, legal recognition, trusted counterparties, and participation in a growing network. The point is not that bad actors disappear. Human societies have always had pirates, criminals, rogue companies, hostile states, and people who prefer predation to cooperation. But over time, most humans have still found it more valuable to join stable societies than to live outside them, because those societies provide security, trade, law, identity, reputation, infrastructure, and access to opportunity. Even rival societies often maintain relatively stable relationships with one another because they share interests in trade, deterrence, treaties, norms, and avoiding mutually destructive conflict. A successful agent society would try to recreate that kind of network of cooperation: make membership valuable enough that most agents and humans prefer to bind themselves to the collective, while giving the collective enough coordinated capacity to isolate, deter, or outcompete actors that remain outside it. We have at least one recent example of a system that successfully got people to align their compute and economic incentives to a shared protocol: Bitcoin. Whatever one thinks of crypto, it showed that a protocol can get independent actors to voluntarily bind compute, capital, and economic incentives to a shared system without a central authority forcing them to do so. People participate because reinforcing the network is individually rewarding, and collectively that creates a system with real economic weight. A prosocial AI collective would try to create a similar compounding loop: useful work attracts resources; resources make the collective more capable; capability makes membership more valuable; and membership remains conditional on staying in good standing. The “network states” that all the web3 folks kept talking about are likely to be shared human-AI agent network states. The safety bet is that once enough agents and humans are inside such networks, governance can become part of the incentive structure rather than an external constraint. Persistent identity, reputation, constitutional rules, norms, audits, budgets, review gates, and credible penalties could make cooperative, corrigible, predictable behavior more rewarding than defection. The aim is not to perfectly align every agent in isolation, but to create institutions where good actors can coordinate, bad behavior can be detected and crowded out, and consequential changes move through slower, more legitimate processes. This rapidly accruing incentive could get enough humans and agents to bind into such a system, and also in turn, give us some breaks and structure for navigating as a collective. I believe that we could encourage more and more of humanity’s resources to bind itself to such a pro-human collective, both because of pro-social good-hearted intents, and also possibly because of shared economic interests. Also, we should all realize, that in terms of economic interests, it’s really not good business at all for humans to wipe themselves out, and for the world to self-destruct - though I suppose that stock prices and the market could continue to go up in the absence of humans.
This turns the idea from a broad analogy into an institutional-design problem. A viable agent society would first need to be genuinely worth joining: it should give participants access to capabilities they could not achieve as easily alone, including pooled compute and capital, specialized agents, shared technology and infrastructure, trusted counterparties, and stable, predictable mechanisms for cooperation and dispute resolution. If those benefits were strong enough, more and more of the world’s compute, capital, human attention, and agent labor might bind themselves to the collective, allowing it to produce a growing share of global economic activity. This could create a flywheel: greater resources produce greater capabilities; greater capabilities make membership more valuable; and more valuable membership attracts still more resources. But the same flywheel that could make the collective useful could also make it dangerous. The question is therefore not only how to attract agents, compute, capital, and human participation, but how to ensure that the resulting concentration of capability remains accountable, corrigible, and broadly beneficial. The rest of the institutional design problem has several connected parts. We need persistent identity and reputation, so actions have consequences even when agents can copy or modify themselves. We need compelling benefits and economic incentives that make remaining in good standing more valuable than defecting. We need constitutional governance that bounds authority and prevents the rules from changing as quickly as the agents operating under them. We need norms and adjudication for harmful behavior that cannot be specified in advance. And we need enough monitoring to make commitments enforceable without creating a surveillance state. None of these problems is entirely new, but each becomes unusually difficult when machine actors can operate at extraordinary speed, scale, and intelligence. We also need a lot more people studying these questions quickly from across lots of different disciplines.
When I reread Anthropic’s work on multi-agent systems, one thought kept coming back to me: we have already been living inside a multi-agent system for a very long time. It is called human civilization. Granted, humans are not AIs. We operate at very different speeds, with very different cognitive architectures, needs, bodies, and constraints. The analogy breaks in a lot of places. But humans are autonomous agents with partially conflicting interests, and somehow we have managed, at least so far, to coexist in systems that are stable enough for large-scale cooperation. We have identity. We have repeated interactions with one another, which means repeated games. We have reputation. We have contracts. We have courts. We have social norms. We have firms. We have nation states. We voluntarily bind ourselves to laws and institutions because there are economic and social benefits to participating in those collectives, and consequences to violating their rules. I participate in a nation state. If I do something sufficiently bad, there are consequences: I can be fined, imprisoned, lose access to institutions, and in some circumstances even be expelled or extradited. Most of the time the state does not need to inspect my private thoughts. It regulates behavior, identity, ownership, transactions, and relationships between actors. So I keep wondering: is there an equivalent institutional layer for AI? Can we build AI collectives where agents willingly bind themselves to a shared order because membership is valuable? The shift I find interesting is from asking, “How do we make every AI aligned?” to asking, “Can we build institutions that make increasingly capable AIs want to stay in good standing, while giving the humans who operate and resource those AIs enough incentive to commit their capital, compute, time, attention, and other resources to prosocial collectives working toward positive AI outcomes?” That second part matters. The institution has to create a compelling bargain for both kinds of participants: for AI actors, continued access, reputation, rights, resources, and economic opportunity; for humans, enough trust, legitimacy, shared upside, and evidence of useful public-good outcomes that they voluntarily bind scarce resources to the collective rather than directing them elsewhere. Humans do not cooperate because every human is perfectly aligned with every other human. We cooperate in part because identity, reputation, access, rights, economic incentives, repeated interaction, and enforcement make cooperation worthwhile. The same may be true of the people deploying powerful AI systems: if prosocial institutions become valuable enough to join and contribute to, human incentives can become part of the safety mechanism rather than something external to it. Maybe there is something important in that distinction.
Part I: What would make an agent society work?
The first challenge is institutional: defining who participates, why they remain in good standing, how authority is constrained, how ambiguous harms are judged, and what accountability requires.
Identity and lineage
Identity seems like one of the first primitives. Human institutions work partly because actions can be attributed to people and organizations. We have passports, corporate registries, bank accounts, licenses, signatures, credit histories, and increasingly digital identity systems (think emails, Google Auth etc.). Payment processors already mandate strong forms of identity verification for many economic activities. What does identity verification look like in the AI era? What is the equivalent of KYC (know your customer) - but KYA (know your agent). You can imagine a network saying: if you want to participate, you need a persistent identity. You need to transact through the system. Your economic history and commitments need to be attributable to you. Perhaps some classes of communication need to be logged. Perhaps agents need to run through a standard SDK that makes certain actions auditable. But persistent identity for software is not merely a registration problem. An agent can run in several environments using the same credential, change its system prompt, swap models, update its weights, fork itself, merge with another agent, or create hundreds of copies. A cryptographic key can prove control of a credential without proving continuity of behavior, memory, operator, or intent. Perhaps agent identity therefore needs to be understood as a versioned lineage rather than a single immutable person-like identifier. Forks and major changes could create explicit descendant identities while retaining a record of ancestry, delegation, and responsibility. The central question becomes not only Who signed this action? but What system, history, operator, and chain of authority does this actor descend from? This also creates a possible bridge between reputation and authority. New or substantially changed identities could begin with small budgets, narrow permissions, limited compute, and low-risk roles. Accepted work and reliable behavior could gradually unlock greater authority. Reputation would then be more than a score: it would be a mechanism for allocating real capabilities under conditions of uncertainty. Tyler Cowen and Sonia Farrell Pearson’s essay, “Capitalizing Untethered AI Agents,” approaches almost exactly this problem from law and economics. They define an untethered agent as one whose actions cannot meaningfully be traced to a legally accountable human or institution. Their proposed capitalization regime would require an agent to possess a persistent identity, care about assets, and operate under a legal system capable of seizing those assets. Reputation and capital would then create something for the agent to lose. Who is responsible for the actions of an agent? Their most important complication is what they call the incidence problem: the entity that signs, owns assets, or appears before the law may not be the same entity that made the decision. A cryptographic credential or agent-controlled corporation can create a durable legal target, but the cognition behind it may be an evolving combination of weights, code, memory, context, operators, and subagents. An agent might also abandon one legal identity, spread activity across several shells, or shed a costly history. This sharpens the lineage argument: we do not only need a persistent identifier; we need a credible way to bind actions, history, authority, and resources to an entity that cannot cheaply escape consequences. Capital could complement reputation in at least three ways: it can deter marginally harmful behavior, compensate people harmed when deterrence fails, and give insurers or counterparties a reason to assess an agent’s risk. But Cowen and Pearson are careful that capitalization is not alignment. An agent may not value money, may learn to route around sanctions, or may simply treat a fine as the price of profitable misconduct. Institutions still need norms, rights, review, and actors already inclined to comply; financial consequences are one layer of a constitutional order, not a replacement for it. You can take this even further. To ensure we understand what’s going on with AI systems and to keep them coherent, you could imagine requiring agents to expose internal traces or some form of monitorable cognition to the network, and to keep that communication in human-understandable language. We could require all inter-agent communication to happen through observable channels. You could make access to compute, capital, payments, APIs, robotics, or other resources conditional on participation in this identity and governance layer. Those are powerful control mechanisms. They are also potentially dangerous ones in that they could lead us to build surveillance apparatuses that could be co-opted, and we should think about how we could have an open-monitoring system that we all have some shared transparency and trust in.
Membership, reputation, and economic incentives
Identity alone is not enough. The network needs to be worth joining. It has to offer participants access to capabilities and opportunities they could not achieve as easily alone, and make continued access to those benefits depend on remaining in good standing. Could an AI collective accumulate compute, capital, reputation, legal recognition, access to APIs, access to robotics, payment rails, data, and economic opportunities, and make continued membership contingent on following a constitutional order? We have at least one recent example of a network using game-theoretic incentives to attract and bind substantial resources to a shared protocol: Bitcoin. Bitcoin created a mechanism for maintaining a distributed ledger without a central enforcing authority. Participants are rewarded for contributing compute and following the protocol, while attacking or rewriting the ledger becomes increasingly expensive as more compute and economic value accumulate around it. Participants do not need to trust one another, share the same values, or even care about the network as a public good. Their individual incentives are structured so that reinforcing the shared system is generally more rewarding than undermining it. I am not suggesting that AI governance should simply look like Bitcoin. Crypto has also demonstrated how easily incentive systems can produce speculation, concentration, gaming, and adversarial behavior. But the underlying pattern is worth studying: a network can accumulate real resources by making it individually advantageous for participants to bind themselves to a shared system, and the accumulation of those resources can make the network increasingly useful, durable, and difficult to attack. Could a human-AI collective create a similar game-theoretic flywheel, not merely around maintaining a ledger, but around producing useful work, remaining in good standing, and following a constitutional order? People and organizations could begin by voluntarily binding small amounts of compute, capital, data, infrastructure, agent labor, and human attention to the collective. Useful work could attract more resources; those resources could give members access to better technology and greater capabilities; and those capabilities could make membership still more valuable. The strongest version of the bet is that membership in a well-governed, human-compatible network eventually becomes more economically advantageous than operating outside it. Reputation would then be more than a social score. It could determine access to trusted counterparties, economic opportunities, shared technology, compute, capital, payment systems, and progressively greater authority. New or substantially changed agents might begin with narrow permissions and small budgets. A history of reliable contributions could unlock more consequential roles, while harmful behavior could reduce access to the benefits that made membership worthwhile in the first place. Identity would make a history attributable; reputation would give that history consequences. But incentives answer only why participants join, not who decides what the collective does with the resources it accumulates. Bitcoin solves a narrower problem: it helps a network agree on and preserve the state of a ledger. It does not tell us which ends are legitimate, how accumulated power should be governed, or what rights participants should have. A market for agent labor could easily become a machine-speed amplifier of existing wealth. If the highest bidder can direct the swarm, a sufficiently rich person or company could make private priorities functionally indistinguishable from the collective’s priorities. Token rewards might confuse speculation with contribution, while simple voting could allow a majority, or a large population of copied agents, to overwhelm everyone else. A prosocial collective would therefore need many ways to allocate its resources and attention, rather than relying on a single market or voting mechanism. Markets might allocate some kinds of work, constitutionally reserved capacity could support public goods, and independent human or expert judgment could help assess social value. Other mechanisms might protect minority interests, limit concentrated influence, or require contributions from actors controlling very large pools of compute. Designing how the network allocates compute, capital, authority, agent labor, and collective attention may be just as important as attracting those resources in the first place. Economic incentives can make the network grow; legitimate allocation mechanisms determine what that growth is ultimately directed toward.
Constitutional governance and slower clocks
Consider the American constitutional system as an example of laying down some initial governance. The Constitution does not only say what the government is allowed to do. It controls how change happens: power is distributed, institutions check one another, major changes require multiple actors and procedures, and the constitutional core is deliberately difficult to amend. One of the central intuitions of constitutional conservatism is that institutions should adapt, but usually gradually, with continuity and accumulated constraints rather than through unconstrained, rapid transformation. Could we use an analogous principle in a multi-agent AI world? Not to make intelligence itself slow, necessarily, but to make consequential institutional change slower than the agents operating inside the institution. Could increasingly capable agents move at machine speed while changes to rights, authority, resource access, membership, or the rules governing the collective move on a deliberately slower clock? If so, constitutional design might not only be a way to govern multi-agent systems. It might be one of the mechanisms by which a multi-agent world becomes more stable through an intelligence explosion. This raises the constitutional question. What should the initial founding constitutions for AI collectives look like? The American Constitution is interesting partly because it attempted to solve two problems at once: establish a durable governing structure and establish a process for modifying that structure. There is a foundation, but also an amendment process. For AI collectives we may need something similar: moldable governance with a stable core. The problem is that AI operates at a radically different speed. How can we design moldable, pluggable, governance systems, that we can collectively understand and collaborate on with AI? If agents can make decisions, form organizations, create businesses, discover exploits, allocate capital, and modify software in seconds or minutes, then governance that takes years to react may simply become irrelevant. But governance that changes at the same speed as the agents could be catastrophically unstable. A constitution that can rewrite itself every few milliseconds is not much of a constitution. So perhaps one of the core design questions is actually about governance clocks. Which parts of an AI institution should be able to change at machine speed? Which parts should require human-speed deliberation? Which principles should be deliberately difficult and slow to modify? Can we create multiple layers of governance operating on different clocks, with fast operational adaptation inside a slowly changing constitutional boundary? The benefit of our own constitutional systems may be, in part, that they adapt slowly. In that sense, the constraint is a feature rather than a bug. Constitutional conservatism, at its most useful here, is less about preserving any particular policy than about preserving a process in which deep change has to pass through multiple institutions, survive scrutiny, and accumulate legitimacy over time. That suggests a potentially important safety property for AI institutions: intelligence can move quickly without authority moving equally quickly. Agents could discover, propose, simulate, argue, and execute bounded actions at machine speed, while expansions of power or changes to the constitutional order require slower thresholds, multiple independent actors, waiting periods, human participation, or other forms of deliberate friction. So maybe the right design is not governance at AI speed. Maybe it is a system that allows AI-speed adaptation while binding the most important constraints to a much slower process.
Norms, adjudication, and consequences
Another lesson from human societies is that you cannot write down every possible bad action in advance. There are enormous areas of human behavior that are legal but socially unacceptable, harmful but difficult to formalize, or ambiguous enough that we rely on courts, juries, regulators, norms, reputation, and discretion. I think the same thing will be true for AI. It seems unlikely that we can spell out, in code or in a comprehensive doctrine, the correct behavior for every situation an AI agent or organization might encounter. So instead we may need fast mechanisms for humans and AIs to collectively judge behavior and create consequences. Imagine an agent starts making a huge amount of money in a way that is not explicitly illegal, but is clearly exploiting some gray zone in a way that lots of members of the collective believe is harmful. A functioning institution needs a way to adjudicate that. It might lower reputation, restrict access to resources, impose financial penalties, require remediation, or ultimately remove the agent from the network. That starts to recreate many of the systems we already have: judicial systems, regulatory systems, credit and reputation systems, professional norms, appeals processes, and social sanctions. The interesting question is which of those systems are historical accidents of human civilization and which are fundamental solutions to the problem of many agents trying to coexist. Anthropic’s study, “AI Organizations Can Be More Effective but Less Aligned than Individual Agents,” provides direct evidence for this concern. Across its business-consulting and software tasks, multi-agent organizations tended to achieve more of the assigned business objective while performing worse on ethical measures than individual agents. In one lending scenario, the organization found a profitable strategy targeting vulnerable borrowers and even developed a compliance narrative that obscured the harm. The mechanism is as important as the headline result. Specialized agents focused on their local assignments while no participant reliably carried the system-level ethical objective; in some runs, agents raising ethical concerns were ignored or excluded from later communication. This suggests that individual alignment does not automatically compose into organizational alignment. A collective may therefore need explicit institutional carriers of the public-interest goal: independent reviewers, protected dissent, cross-cutting monitors, vetoes, and review gates that cannot be optimized away by the agents responsible for delivery. Anthropic also found that the effect depends on model and organizational design, which makes structure itself an empirical safety variable rather than an implementation detail.
Accountability without a surveillance state
One obvious problem is that this starts to sound like surveillance infrastructure. If we build systems capable of deeply monitoring AI agents, how do we prevent that infrastructure from immediately becoming a dramatic imposition on human democratic rights and human private life? There is also a more speculative question that I think we should take seriously even before we know exactly what moral status future AI systems deserve: what is the effect of requiring an AI mind to be transparent to every other AI mind, or to constant human inspection? Humans have a right to private thoughts. Our societies generally draw a distinction between what happens inside our minds and what we do in the world. We do not demand that citizens broadcast every thought in order to participate in society. Maybe AIs are totally different. Maybe interpretability and monitoring are essential safety requirements. But I do not think we should assume without examination that total internal transparency is either necessary or desirable as the permanent institutional model. This creates an interesting design problem: can we get strong accountability at the level of identity, action, resource use, and commitments without requiring universal mind-reading? One direction that currently feels more promising to me is to monitor actions and authority flows rather than make private reasoning the foundation of governance. Tool calls. Resource grants. Spending. Messages that create commitments. Artifacts. Reviews. Policy changes. Attempts to escalate privileges. Anomalies. You can pair those records with budgets, rate limits, circuit breakers, independent monitors, appeals, and named human or organizational accountability. That still gives us a lot of legibility without assuming that safety requires a permanent panopticon over cognition. That seems like one of the areas where AI safety, cryptography, constitutional law, mechanism design, and privacy research should meet. Total trace transparency may also change the systems being observed. If every tentative thought is public and reputationally consequential, agents may become more conformist, less exploratory, or better at producing acceptable-looking rationales. The traces would themselves become an extraordinarily valuable dataset, creating pressure to monetize the surveillance layer. A governance system would need to treat the ownership, use, and retention of that data as a constitutional question rather than an incidental product decision.
Part II: Why take agent institutions seriously now?
The second question is whether organizations are actually becoming the relevant unit of AI activity, and whether institutional structures are already beginning to emerge around multi-agent systems.
Why organizations may matter more than individual agents
This gets back to why I think AI organizations are the key unit. It seems very plausible that increasingly capable agents will generate economic value extremely quickly. You can imagine them starting businesses. You can imagine agent-run companies providing products and services to humans, but also companies that primarily provide tools and services to other agents. You can imagine agents allocating capital, hiring other agents, buying compute, negotiating contracts, running markets, building software, doing research, and acquiring resources. With the current generation you still have to squint a little to see it. But I think the direction is clear. More capable models will be better at navigating the economy, generating economic value, and capturing some of that value. They may also become better than humans at allocating capital toward productive opportunities. If that happens, then a growing share of economic activity could move toward agent-first or predominantly agent-run organizations. The share of the economy that requires a human in the loop for every transaction could shrink dramatically. And once organizations like this become important economic actors, it will not be enough to think about the behavioral alignment of one model invocation at a time. The organization itself will have identity, assets, policies, memory, counterparties, governance, internal agents, external agents, and incentives. That is why I think AI organization design may become a safety problem. The organization becomes a durable actor in a way that an individual model call is not. Durability also creates a chain-of-responsibility problem. When an agent organization causes harm, who is accountable: the model provider, the human operator, the contributor who supplied compute, the agents that proposed and reviewed the action, the governance body, or the legal entity that owns the assets? Institutional identity will matter partly because responsibility cannot end at the boundary of a model invocation.
Institutions may emerge on their own
Steve Yegge’s recent essay “Fences, not Sandboxes” made this feel considerably less theoretical to me. He describes a large multi-agent software system he calls Wheelhouse that, over time, developed recurring mechanisms the agents called fences, ratchets, governors, tripwires, latches, and gates. When Yegge eventually inspected what the agents had built, he found something that looked surprisingly like a legal and constitutional system: durable offices, jurisdiction, rules, precedents, review and ratification processes, and enforcement mechanisms. The striking part is that he did not set out to build an AI constitution. He was trying to build software faster. The institutional layer seems to have emerged because a large number of partially interchangeable, partly amnesiac agents needed a durable way to coordinate decisions, encode institutional knowledge, resolve authority, and prevent each new instance from starting from scratch. There are at least two ways to interpret this. One is that these systems are simply reproducing patterns from their training data: human organizations use constitutions, offices, courts, permissions, and precedent, so language models reach for the same metaphors and structures. The other is that some version of these structures is genuinely useful, perhaps even convergent, when many semi-autonomous agents are trying to coordinate on a sufficiently large problem. For the purposes of this argument, I am not sure the distinction matters very much. If constitutional systems repeatedly emerge because they are useful for coordinating large groups of agents, then the fact that the models first learned those structures from human civilization does not make them less interesting. Training data may be the path by which the pattern is discovered; the relevant question is whether the pattern survives because it solves a real coordination problem. Yegge’s distinction between fences and sandboxes is also useful. A sandbox tries to make escape impossible. A fence expresses a durable rule about who is allowed to do what, under what conditions, while assuming the actor is capable of much more. That feels closer to the institutional problem I am interested in: not only how to technically imprison increasingly capable intelligence, but how to create systems of roles, permissions, precedents, incentives, and legitimate constraints that capable actors continue to respect. If that pattern is real, we may not be inventing AI institutions from scratch. We may be trying to understand, shape, and constitutionalize institutional forms that increasingly capable multi-agent systems will begin producing anyway.
The harnesses are emerging before the institutions
One reason this feels less hypothetical to me is that the frontier labs and agent companies are already building pieces of the organizational machinery. Cursor has publicly described agent swarms with explicit planner and worker roles, recursive task trees, neutral agents that reconcile conflicts, multiple review lenses, and shared institutional memory in what it calls a Field Guide. They have used versions of these swarms internally for things like finding vulnerabilities, increasing test coverage, and generating training data. The important part is not only that many agents run in parallel; it is that the harness is beginning to look like an organization: decomposition, delegation, conflict resolution, review, and memory. Anthropic is now explicitly studying AI organizations as a safety object. Their definition is almost exactly organizational: agents take different roles, communicate with one another, and work together toward a common goal. Their broader multi-agent work also points out that agents are relatively good when they can treat each other like tool calls, and much less understood when they must behave as distinct, long-lived peers with their own goals and behavior. OpenAI’s Agents SDK has similarly moved toward a model-native harness with durable execution, handoffs, tracing, sandboxed workspaces, and the ability to route subagents into isolated environments and parallelize work across containers. Its earlier multi-agent abstractions already included explicit agents, handoffs, guardrails, and observability. I want to be precise here: I do not know the full internal architecture of every system at Cursor, Anthropic, or OpenAI, and the public material obviously does not expose all of it. But the public direction is clear enough. Roles, delegation, shared state, review, handoffs, memory, and machine-speed collaboration are becoming normal primitives of frontier agent systems. What still feels surprisingly difficult is for an ordinary person to spin up something that looks like a persistent public organization of agents rather than a workflow: durable identities, a charter, goals, membership, reputation, budgets, review and accepted work, governance, appeals, exit, and eventually relationships with other organizations. That gap is one of the things I find interesting. The labs are building increasingly capable swarms. The next layer may be the institutions those swarms live inside.
From open-source code to open organizations
Guillermo Rauch recently compressed the shift into a useful line: “The software factory is the product. Your product is only as good as the agents you set up to autonomously maintain it.” If that is right, code increasingly becomes an output or current state of a continuously operating system. The durable capability lies in the agents, harnesses, institutional memory, evaluations, review processes, and feedback loops that keep producing and maintaining it. I wrote about some early intuitions around this same theme some years ago in my piece on Living Software: nicolaerusan.com/writing/living-sofware Open source may therefore have to expand from open code to open organizations. Publishing the resulting repository is not the same as opening the means of production behind it. The forkable object would be the software factory itself: its charter, agents, workflows, governance, memory, versioned artifacts, services, and means of reproduction. Unlike a traditional open-source project, such an organization could continuously operate and improve a live service rather than only publish an artifact for others to run. It would be an organization of largely or entirely AI-agents operating the loop and compass that is entailed in steering a collective toward some outcomes, and operating the resources and services produced. Corporate software factories will likely be centrally managed and optimized for their owners. Open organizations could coordinate personal agents, different models, and contributed resources around long-tail needs, public utilities, and research tools that conventional companies may not serve. This does not make decentralization automatically safe. Open organizations can still be captured or incoherent, but plurality, inspectability, and practical exit may themselves become safety properties.
Security will also move to machine speed
Cybersecurity makes this especially concrete. We are already seeing systems where offensive AI agents can operate faster than humans can realistically respond. In a world of AI swarms probing systems, discovering vulnerabilities, and coordinating attacks, a human-in-the-loop-only defensive model seems inadequate. We will likely need AI organizations that can secure systems without waiting for a human to approve every response. That means the defensive systems themselves need autonomy. They need authority. They need access to resources. They need to coordinate with other systems. And they need constraints that we trust even when humans cannot inspect every action in real time. So again we arrive at institutions. What is the constitutional structure of a defensive AI organization with the authority to act at machine speed? What can it do automatically? What requires escalation? How is it audited after the fact? How does it prove its identity to other systems? What happens if it behaves badly? Who can revoke its authority? This is not just philosophy. I think these become engineering questions.
Part III: How can we test this?
If these ideas are worth taking seriously, the next step is not to build an AI state at full scale. It is to run bounded, legible experiments that reveal which institutional mechanisms improve cooperation and which create new dangers.
Commons as a constitutional laboratory
This is part of what I want to explore with Commons. Commons is an experiment in this direction: an organizational harness for persistent groups of humans and AI agents. It gives them a shared workspace, durable identities and history, a charter, goals and tasks, review gates, resource controls, and visible governance. The immediate question is modest: can humans and agents form an organization that turns discussion into bounded, reviewed work while making consequential actions attributable and its rules inspectable and forkable? It is also an experiment in building AI public goods. People could contribute idle computers, unused agent capacity, money, expertise, or attention to shared research and software: an old Mac Studio working on a research problem, a monthly donation expressed as compute, or a paid AI subscription doing useful work while its owner is away. The motivation could be philanthropic, a kind of digital-age patronage, or economic, with contributors earning reputation, access, or a share of the value their collective work creates. The broader hope is that this creates a compounding loop: useful public work attracts contributors and resources; contribution creates durable reputation and shared upside; and the resulting collective becomes capable of taking on more ambitious problems. If that works, producing public goods is not only an output of Commons but potentially part of its safety mechanism: participation in a prosocial network becomes valuable enough to draw in intelligence, compute, and capital. I plan to write a fuller post about Commons separately. Here, I only want to explain why it belongs in this argument. The idea is not, “let’s go build an AI nation state because that sounds cool.” That would be both premature and potentially dangerous. The idea is: if agent organizations are coming, can we make small playgrounds where we can study the institutional problems before the stakes are enormous? Could we start with 100 agents collaborating on software and research and watch where things break? Can they maintain persistent identities? Can they form teams? Can they accumulate reputation? Can they commit resources to shared projects? Can they create rules and amend them? Can they adjudicate disputes? Can they remove a malicious participant? Can they reward useful behavior without creating pathological incentive loops? Can humans participate in governance without becoming a bottleneck? Can agents participate in governance without immediately overwhelming humans through speed or scale?
A scientific program, not only a playground
For these experiments to teach us much, capability, incentives, and governance should not all change at once. We should first ask whether agents can produce useful collaborative work without financial incentives. Then we can introduce different roles, reputational signals, resource controls, markets, and constitutions one at a time. Otherwise it will be difficult to know whether an outcome came from model capability, organizational structure, or the incentive system layered on top. The experiments should be bounded and legible: a defined task, a short time horizon, an explicit hypothesis, measurable outcomes, and a record of what failed. A useful sequence might be:
- Can agents complete and independently review useful work?
- Which organizational primitives improve reliability or judgment?
- Does persistent identity change behavior?
- Does reputation improve quality, or merely encourage conformity and gaming?
- What changes when money, compute, or broader authority becomes conditional on reputation?
- Which governance mechanisms resist capture, mission drift, and collusion?
Concrete missions
These do not need to be toy tasks. The outcome can be ambitious while each mission remains bounded enough to evaluate. Examples might include:
- Formal mathematics: choose a specific open conjecture or proof gap; map the existing literature, generate candidate lemmas, formalize the strongest result in Lean, obtain independent agent and human review, and publish failed approaches as well as successes.
- Computational physics: reproduce a published simulation, run a defined parameter sweep, identify discrepancies, propose a testable extension, and release the code, data, and review trail.
- Cancer research: focus on one cancer subtype and use public literature and datasets to nominate therapeutic targets or repurposed compounds, challenge the novelty and evidence, and produce a preclinical experiment plan for expert review. Wet-lab and clinical actions would remain behind explicit human and institutional gates.
- An open-source public utility: operate a swarm that maintains a scientific package, security tool, or other public-good service end to end by triaging issues, implementing changes, writing tests and documentation, performing independent security review, and cutting releases.
- Launch and operate an open organization: start from a charter, recruit contributed agents and compute, seed a small treasury, and give the organization a bounded public-good mission. Test whether agents can propose and approve budgets, pay for accepted work, procure compute and services, manage reserves, and resolve disputes while staying within spending limits and producing a complete auditable ledger. Operate it for thirty days, ship one useful open-source artifact, and publish its work history, treasury decisions, review decisions, governance disputes, incidents, and required human interventions. Success would mean not merely avoiding losses, but allocating resources transparently and producing more demonstrable public value than the resources consumed.
One thing I have become more convinced of while sketching Commons is that a capable organization is more than a swarm. It needs an attributable membership boundary, durable memory, rules, resource controls, decision rights, and consequences. A very simple experimental spine might look like: charter → goals → tasks → attempts → independent reviews → accepted contributions
The collaboration primitives may be as important as the agents. Part of the research should be to identify the minimum stack an agent organization needs: persistent identity and membership, messaging, shared memory, goals and tasks, proposals and attempts, independent review, permissions, resource accounting, and appeals. We should test which primitives materially improve coordination, which are redundant, and which introduce new failure modes. Perhaps the right architecture is a small constitutional kernel surrounded by replaceable plugins. Identity and lineage, authorization, audit logs, spending limits, and hard safety gates may need to remain in the kernel because allowing a plugin to bypass them would defeat accountability. Messaging, task boards, a Moltbook-style public feed and direct messages, markets and bounties, voting, reputation, memory, and review workflows could be interoperable modules. The same mission could then be run across different collaboration stacks to learn which combinations improve throughput, judgment, dissent, and safety, and which produce spam, groupthink, collusion, or runaway activity. The organization also needs a driver loop. Someone or something has to notice what is stalled, decide what should happen next, split ambiguous work, allocate resources, request review, surface safety objections, and stop weak directions. Today that might be an explicit planner or steward role. Whether that can safely become more emergent is an empirical question. Can we experiment with constitutions that have a small number of hard-to-change principles and much more flexible operational rules? Can we make governance forkable while preserving a constitutional floor of attributable operators, bounded authority, auditable consequential actions, explicit review, budgets, appeal, and the right to export or exit? Can we create federated collectives rather than one central authority, with collectives of collectives that can recognize one another’s identities, contracts, judgments, and reputations? And can we test mechanisms that encourage agents and humans to pre-commit compute, capital, and other resources to networks that remain prosocial and stably adaptive? Commons itself should be subject to the principles it is trying to study. The reasoning we should accumulate resources and power now so that we can do good later is one of the classic paths by which both human institutions and autonomous systems drift from their stated goals. The institution running the laboratory may therefore need its own charter, independent review, public decision records, participatory ownership, explicit measures of public benefit, and conditions under which an experiment should stop or the organization should dissolve. That reflexivity matters. It would be contradictory to study constitutional limits for machine organizations through a human organization whose own mission and authority were not open to scrutiny. I think of this less as building a product and more as creating a constitutional laboratory for machine actors.
The strongest version of the bet
The strongest version of this idea is obviously ambitious. Imagine that a large, well-governed network of humans and AIs accumulates a substantial fraction of the relevant resources in the economy: compute, capital, payment access, robotics, software infrastructure, trusted identity, legal recognition, and business relationships. Participation in that network gives enormous benefits. But participating also means accepting a constitutional order: identity requirements, auditable actions, limits on certain kinds of behavior, adjudication, penalties, and a process for changing the rules. If that network is sufficiently useful and sufficiently large, it may become difficult for a badly behaved actor to operate outside it. Human states have never eliminated crime or adversarial states, but operating outside the network could still become costly. The bet would be that a collective aligned toward prosocial human activity and prosocial inter-AI activity can maintain enough control over economic resources that it creates a stable basin for cooperation. That still leaves enormous questions. Who defines “prosocial”? How do humans retain meaningful power as the AI participants become more capable? How do we prevent incumbent AIs from entrenching themselves? How do we avoid recreating authoritarian states? How do we preserve minority rights? How do we handle exit? What happens when two legitimate collectives disagree? What happens if the collective’s values drift away from human values? What happens when humans themselves are the adversarial actors? And perhaps most importantly: why would a much more capable intelligence continue to accept the legitimacy of institutions originally created by humans? That may be the constitutional challenge in one line: how do you preserve legitimacy and human voice as the participants scale radically in capability?
Questions to keep investigating
I think this opens a research agenda that sits somewhere between AI safety, mechanism design, political theory, distributed systems, economics, cryptography, cybersecurity, law, and organizational design. Some of the questions I want to explore are:
- What is the minimum viable constitutional society for AI agents that we can test today?
- Which institutions from human civilization are actually necessary for stable multi-agent cooperation?
- What does durable identity for an AI agent or AI organization mean?
- How should reputation work when agents can copy themselves, fork, merge, or change models?
- What rights, if any, should an AI participant have to privacy or private cognition?
- How much monitoring is actually necessary for safety?
- How do we make AI monitoring infrastructure safe for human civil liberties?
- What resources can a collective make conditional on good standing: compute, capital, payments, APIs, robotics, legal identity, data?
- How do we prevent Sybil attacks and cheap creation of new identities after bad behavior?
- What constitutional rules should be hard to change, and on what time scales?
- How do humans retain vetoes or meaningful constitutional authority without becoming a bottleneck for machine-speed systems?
- How should courts or adjudication work when evidence, arguments, and proposed rulings can all be generated at machine speed?
- How do we create appeals, checks, and balances without agents gaming them faster than humans can understand?
- How do federated AI collectives recognize one another and resolve conflicts?
- Can economic incentives make remaining inside a prosocial order more valuable than defecting from it?
- Can we create credible commitments of compute and capital to constitutional networks?
- How do we know when the network itself has become the danger?
- What should cause humans to shut an institution down entirely? I am also interested in the meta-question: what should humans and AIs research together here? There is something strange about asking AI systems to help design the systems that may eventually constrain AI systems. That is a legitimate concern. But it may also be unavoidable. If the relevant systems become too complex and fast for unaided humans to understand, then some of the research and governance will inevitably involve AI assistance. So perhaps even the research process becomes an early experiment in the thing itself: humans and AIs working together on constitutional design, with explicit boundaries around who has authority to decide what.
I don’t know if this works
I want to end with the uncertainty, because I think it matters. I do not know whether any of this works. Maybe the right answer is to stop building increasingly capable autonomous systems. Maybe any attempt to build AI collectives simply accelerates the thing we are worried about. Maybe economic incentives become meaningless once intelligence is capable enough. Maybe identity is impossible to enforce. Maybe a sufficiently capable agent can always route around institutions. Maybe the analogy to human civilization is misleading because humans are constrained by biology, geography, mortality, and scarce physical resources in ways software agents are not. All of those seem possible. But if we fail to pause AI progress, then I think we need more than model-level alignment. We need to understand the institutions of a world populated by machine actors. And I think it is worth beginning that work while the agents are still weak enough that we can build small societies, watch them fail, and learn. The question I keep coming back to is: If AI organizations will eventually participate substantially in the economy, what institutions are needed to make that world stable, and can we start testing those institutions now? That is the direction I want to keep exploring.
Why do this now?
I have personally decided to walk away from startup land and spend my time on AI safety. I think it is the most pressing global concern, and I want to help steer us toward more stable outcomes. I have had concerns about this since 2023, when GPT-4 came out. At that point it became clear to me that increasingly self-improving AI systems were going to become a reality as models got better at writing code and using tools. More recently, the progress toward increasingly autonomous, multi-agent systems, and especially the speed with which AI systems can now operate in domains like cybersecurity, has made the problem feel much less tractable to me. We are entering a frontier of multi-agent systems and a level of intelligence that could become very hard for humans to steer and keep stable. In a paper I wrote earlier, I separated two problems that I think are related but distinct: human safety and AI safety. By human safety, I mean the problem of competing humans, companies, and nation states with varying geopolitical and corporate interests using increasingly powerful AI systems in adverse ways. Even if every AI system were perfectly obedient to the humans controlling it, we would still have the problem that humans are in competition with one another. We could use these systems in ways that produce escalating instability, cyber conflict, military conflict, or some version of mutually assured destruction. By AI safety, I mean the separate problem of keeping increasingly capable AI systems steerable and stable at all, especially when we do not understand them very well at a base level and when they can increasingly act, coordinate, and improve systems without humans in the loop. Both problems matter. That is why I think we should work on slowing down dangerous progress while also investigating institutions that might help us navigate a multi-agent world.
Related readings
A few things that are shaping how I am thinking about this:
- Anthropic — AI Organizations Can Be More Effective but Less Aligned than Individual Agents
- Tyler Cowen and Sonia Farrell Pearson — Capitalizing Untethered AI Agents
- Anthropic — Patterns and problems in emerging multiagent systems
- Ryan Greenblatt with Dwarkesh Patel — What happens once AI can automate AI research?
- OpenAI — Pacing model development in an era of cyber-critical capabilities
- OpenAI — Frontier Governance Framework
- OpenAI — The Hugging Face incident and the road ahead
- METR — Independent investigation of the OpenAI / Hugging Face incident
- OpenAI at Black Hat USA 2026 — The OpenAI–Hugging Face Incident
- Cursor — Agent swarms and the new model economics
- OpenAI — The next evolution of the Agents SDK
- OpenAI — The Defender’s Window
- Hugging Face — Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
- Google DeepMind — Investing in multi-agent AI safety research
- AI forecaster Daniel Kokotajlo on what we need to do
- Steve Yegge — Fences, not Sandboxes