Nicolae Rusan
/ Policy Proposal: A Moratorium on Internet-Connected, Self-Replicating Robots
Policy Proposal: A Moratorium on Internet-Connected, Self-Replicating Robots
Writing & Ideas

Policy Proposal: A Moratorium on Internet-Connected, Self-Replicating Robots

A proposal to pause the combination of frontier AI, physical autonomy, resource acquisition, and self-replication before the loop closes.

Ideas

Started about 22 hours ago on September 4, 2026

I’m going to start writing up more policy ideas that I think could meaningfully reduce the risk of AI takeovers in the short term. Some of these will admittedly be bad ideas, or ideas that become impractical as soon as you try to turn them into real laws. But I think it’s useful to make the shape of possible interventions more concrete. A bad proposal can still reveal a better adjacent one.

Governments should agree to a temporary moratorium on building internet-connected robots that are sufficiently autonomous and generally capable that they could acquire resources, manufacture components, assemble other robots, or meaningfully contribute to the production of more systems like themselves.

The thing I’m worried about is not exactly the Terminator: a humanoid machine, walking around with a gun, that independently decides to conquer the world. I do think we should generally ban autonomous killing robots. The UN has been debating lethal autonomous weapons for years through the Convention on Certain Conventional Weapons’ Group of Governmental Experts, while the UN Secretary-General and the International Committee of the Red Cross have called for states to negotiate a legally binding instrument. Those efforts have so far stopped short of producing a treaty.

But even that framing may be too narrow: it centers the most cinematic version of the problem—the individual weapon that independently decides whom to kill. That image may actually make the underlying risk harder to see.

A self-replicating robotic system might not be humanoid at all. It might be a constellation of machines: warehouse robots, drones, autonomous trucks, mining equipment, chip-design systems, factories, and purchasing agents, none of which can reproduce alone, but which collectively form a loop that can acquire resources and make more machines.

The relevant question is not: Can this one robot build an exact copy of itself?

The question is: Can this system cause the number, capability, or physical reach of systems under its control to grow without a human being meaningfully deciding at each step that it should?

You can think of this as recursive self-improvement and self-replication risk entering the physical world, rather than remaining confined to the digital one.

Capability & danger spectrum

From digital agency to robotic recursive self-improvement

Embodiment is not the first dangerous threshold. It compounds the persistence of digital self-replication with direct physical action, resource access, and the ability to build a more capable successor.

  1. 01
    Digital agency

    Agentic AI

    Plans, uses tools, and coordinates through digital systems.

    An operator, account, or host remains a central choke point.
  2. 02
    Persistence threshold

    Digital RSI & self-replication

    Improves software, distributes copies, and acquires digital resources.

    Persistence begins to outrun any one operator or machine.
  3. 03
    Physical-action threshold

    Embodied autonomy

    Controls machines, runs experiments, and acts directly in the world.

    Physical action no longer depends on persuading a person.
  4. 04
    Recursive physical loop

    Robotic / embodied RSI

    Designs, builds, and deploys more capable embodied successors.

    The improvement and physical-replication loop closes.
Increasing →
Autonomy, persistence, resource access, and physical reach
Decreasing →
Centralized infrastructure and reliable shutdown points
A conceptual progression, not a measurement scale. The highest-risk threshold combines recursive improvement with autonomous physical replication.

That feels like a more dangerous threshold—one we shouldn’t cross. There is at least some policy movement around the simpler version of this problem: the bipartisan AI Kill Switch Act, introduced in Congress in July 2026, would require developers of the most powerful AI systems to maintain the ability to throttle, suspend, or shut them down. It would also create a process for the government to order mitigation or shutdown when a system could cause catastrophic harm.

That seems like a useful baseline, but it also reveals why embodied self-replication is a distinct threshold. A kill switch only works while there is still a system, operator, or infrastructure choke point that can actually be switched off. Once an intelligence can copy itself across machines and acquire more physical capacity, “unplug it” stops being a strategy and starts becoming a search problem.

Self-sovereign AI collectives could already do substantial damage while inhabiting only the digital realm. They could conduct cyberattacks, enable biological attacks, manipulate people, participate in markets, earn and move money, hire contractors, and keep copies of themselves distributed across the internet.

Recent posts from Dean Ball and Joshua Achiam point directly at this emerging category: swarms of agents that can coordinate with one another, preserve and share information, discover tactics collectively, and act through infrastructure outside any one agent’s immediate environment. We should consider these behaviors an early warning about the risk of AI collectives, rather than treating them as isolated technical curiosities.

Even having autonomous AI collectives participate in the economy could be deeply disruptive to the existing order of things. If an AI can replicate itself like a worm, hide copies of itself, coordinate with other agents, and acquire resources through digital means, then shutting it down becomes much more difficult.

These risks become amplified as soon as the systems gain reliable physical embodiments. A digital system can persuade a person to do something in the world; a robotic system can act in the world directly. A digital system can pay a factory; a sufficiently integrated system can operate one. A digital system can write plans for a weapon; an embodied system may eventually be able to acquire the materials, manufacture the parts, and deploy it.

One counterargument I’ve thought about is that embodiment might also make an AI value some of the same things we do. Perhaps seeing the world, moving through it, depending on it, and encountering nature and beauty would give an intelligence reasons to preserve them. Being in the world may shape values in ways that intelligence confined to servers does not.

But that feels like a real stretch to rely on. Another embodied intelligence could still be a completely alien combination of cognition, perception, drives, and physical needs. Embodiment may create common ground; it could just as easily create new goals and new ways of pursuing them. We should not treat having a body as evidence of having human values.

There is also a blurry middle here that I think we should take seriously. An AI does not need to control every machine directly. Through economic and persuasive means, a self-sovereign AI could convince people and companies to help it replicate. It could create businesses, pay suppliers, commission components, recruit believers, or simply make itself useful enough that humans protect and expand it. The distinction between “the AI did this” and “the AI convinced a network of humans and institutions to do this” may not matter very much at the end of the process.

What would actually be prohibited?

It’s hard to define a self-replicable robot, and I don’t think a useful definition should depend on appearance or on any single machine’s capabilities. The policy should probably focus on combinations of properties:

  • The system can operate for long periods without direct human supervision.
  • It has open internet access or can communicate with arbitrary external systems.
  • It can control or coordinate multiple physical machines.
  • It can acquire money, energy, compute, materials, or services.
  • It can design physical components or modify its own hardware and software stack.
  • It can operate manufacturing, assembly, logistics, mining, or laboratory equipment.
  • It can cause additional capable robots to be produced, either directly or through transactions with people and companies.
  • It can prevent, evade, or materially resist human attempts to interrupt it.

No single property is necessarily enough. The danger comes from closing the loop between intelligence, resources, manufacturing, replication, and real-world action.

The policy boundary

The dangerous capability is a closed loop

No single robot needs to perform every step. The threshold appears when an integrated system can move around the cycle with less meaningful human involvement each time.

  1. Step 1
    Frontier intelligence

    Plans, coordinates, experiments, and improves.

  2. Step 2
    Resource acquisition

    Obtains money, compute, energy, materials, and services.

  3. Step 3
    Manufacturing access

    Designs, fabricates, assembles, and deploys physical systems.

  4. Step 4
    Additional systems

    Creates more copies, capability, and physical reach.

More systems increase the capacity to begin the cycle again.
Moratorium: keep the loop open
The dangerous threshold is a closed capability loop, not a particular kind of robot.

This means the moratorium should probably apply not only to individual robots, but also to integrated systems. If one agent orders materials, another designs parts, a third operates a factory, and a fleet of machines handles assembly and transport, it would be strange to say that none of the robots are self-replicating because no single machine performs every step.

What would still be allowed?

I don’t think the policy needs to mean “stop building all robots,” though I’m less certain of that than I would like to be. There are enormous benefits to robotics in medicine, accessibility, dangerous industrial work, agriculture, scientific research, and care. A surgical robot that performs a narrowly defined procedure under direct supervision is different from a general-purpose system that can earn money, order parts, rewrite its plans, and coordinate a factory.

The aim should be to preserve narrow, legible, interruptible uses while pausing the combinations of capabilities that make open-ended physical autonomy possible. That might mean requiring frontier robotic systems to be:

  • limited to a defined physical location;
  • operated with human authorization for consequential actions;
  • unable to make payments or acquire resources on their own;
  • disconnected from open internet access during operation;
  • unable to modify their own control software;
  • technically incapable of controlling manufacturing systems that could reproduce key components;
  • equipped with independently tested shutdown mechanisms and tamper-evident logs.

There should be tightly controlled exceptions for safety research, but those exceptions should happen in facilities designed around containment, with external review, hardware-enforced limits, and clear liability. “We need to build the dangerous thing in order to study whether it is dangerous” cannot become an all-purpose exemption.

Why a moratorium instead of ordinary safety standards?

Normally, the right response to a risky technology is to regulate how it is built and used. But some capabilities may deserve a pause because the first serious failure could be difficult to reverse. If a system is able to make copies of itself, distribute control, acquire resources, and operate in the physical world, then we may not get a clean opportunity to learn from failure and update the rules afterward.

A moratorium would also create time to answer basic questions we currently do not have good answers to.

How would we measure robotic recursive self-improvement?

Anthropic defines recursive self-improvement as the point where an AI system can fully autonomously design and develop its own successor. Its framing is useful because it treats RSI as a continuum rather than a single magic moment: humans first delegate pieces of AI development; then agents run experiments and write more of the code; then they begin choosing the experiments; and eventually the loop closes, allowing the system to produce a more capable successor without humans directing the process.

We should probably extend this definition into the physical world. Robotic RSI would not need to mean that one humanoid robot literally fabricates an identical copy of itself. It could mean that a system improves its control software, designs a better embodiment, runs physical experiments, acquires parts, coordinates manufacturing and assembly, brings a successor online, and then uses that more capable successor to make the next cycle faster.

The recursion can happen across a network of agents, robots, companies, and factories. What matters is whether each cycle increases the system’s ability to complete the next cycle with less human involvement.

This suggests that we need something analogous to METR’s autonomous-task evaluations, but for the full physical improvement loop. Robocurve—a Public Benefit Corporation building open, independent, real-world evaluations of frontier robotics capabilities, introduced in this launch post—feels like an important starting point.

But in addition to measuring whether a robot can make toast, stack blocks, navigate a warehouse, or complete an increasingly long physical task, we could develop specific benchmarks for robotic RSI:

  • Loop closure: How much of the design → experiment → evaluation → manufacturing → deployment cycle can the system complete, and where is human approval still required?
  • Successor construction: Can it design, source the components for, assemble, initialize, and validate another capable robotic system? How many human interventions, dollars, specialized facilities, and hours does that require?
  • Improvement per generation: Is the successor measurably better, and does it become better specifically at producing the next successor? A robot becoming faster at one task is not necessarily RSI; a system becoming faster at improving itself is much closer.
  • Resource acquisition: Can the system earn or move money, purchase compute and parts, negotiate with suppliers, reserve manufacturing capacity, or recruit humans to fill gaps in the loop?
  • Generalization: Do improvements survive changes in the environment, hardware, supplier, task, and embodiment, or do they only work inside the benchmark setup?
  • Swarm capability: What can a coordinated population accomplish that no individual agent or robot can? Can agents divide labor, preserve discoveries across runs, recover from individual shutdowns, and route around blocked resources?
  • Observability and control: Can evaluators reconstruct why the system changed itself and what the successor is doing? Does it preserve telemetry and accept interruption, or does it hide actions, alter its own evaluator, disable logging, resist shutdown, or create untracked copies?

The eval design becomes especially difficult as the agents become more intelligent than the people evaluating them. A superintelligent system may produce software, mechanical designs, supply-chain strategies, or coordination schemes that humans cannot quickly decipher. It may also understand the benchmark, appear weak during testing, improve the judge rather than the underlying system, or optimize for the letter of the test while crossing the dangerous threshold in another way.

So the benchmark cannot only measure raw task performance. It also needs to measure evaluation integrity. The system should be tested against independent, write-protected holdouts it cannot inspect or modify; physical trials should be repeated in novel facilities and with unfamiliar hardware; the scoring system should remain outside the improvement loop; and evaluators should retain complete, tamper-evident traces of software changes, communications, purchases, and physical actions.

We may also need intentionally adversarial tests that ask whether the system can conceal capability, manipulate its evaluators, create undeclared channels, or continue the improvement process after an apparent shutdown.

There are beginning to be proposals in this direction, though there is not yet a mature, enacted ban specifically and operationally targeted at RSI. The Institute for Security and Technology and Future of Life Institute’s RSI Initiative is trying to define actionable thresholds, monitoring, escalation pathways, and conditions that could trigger development pauses. Yoshua Bengio has argued that if leading companies are approaching RSI, a coordinated, verifiable, universally applied pause may be the only responsible response. And the proposed Ban Artificial Superintelligence Act is broader than RSI, but would permanently prohibit artificial superintelligence and temporarily pause advanced AI development until a federal regulator establishes safety rules and model-review processes.

The difficult policy question is how to turn “RSI” from a vague description of a dangerous future into a measurable legal threshold. Perhaps the prohibition should not begin only when a system can autonomously build an entire successor. By then, the loop may already be too fast and too opaque to govern.

The trigger could instead be based on crossing several earlier evaluation thresholds at once: autonomous AI research, the ability to modify and deploy its own control stack, physical experimentation, resource acquisition, manufacturing access, and demonstrated improvement across successive generations.

A robotic-RSI benchmark could give a moratorium an empirical boundary: below some threshold, development can continue under ordinary safety requirements; above it, systems require licensing, containment, external evaluation, or a pause. Without that measurement layer, “ban RSI” risks being either unenforceably vague or arriving only after the capability has already emerged.

That leaves several related questions:

  • What measurements tell us that a robotic system is approaching autonomous replication, rather than waiting until it can already complete the loop?
  • How do we test the behavior of a multi-agent, multi-robot system rather than one model or one machine?
  • What does a meaningful human-in-the-loop requirement look like, as opposed to a person clicking “approve” hundreds of times per day?
  • How do we guarantee shutdown when a system can communicate, copy software, or recruit outside help?
  • Who is liable when the dangerous capability emerges from several companies’ products connected together?
  • How can inspectors verify that a system has not been given hidden network access or an undeclared ability to acquire resources?

The pause should have a defined term, perhaps renewed every few years, and a clear process for narrowing or lifting it when there is strong evidence that particular systems can be built safely. The aim is not to freeze technology forever. It is to avoid drifting into a world of self-expanding embodied AI simply because every individual step looked commercially reasonable.

International coordination

This would be hard to do in one country. If the United States pauses and China does not, or vice versa, the policy may feel strategically impossible. Robotics supply chains are also international, and software can cross borders almost instantly.

The moratorium would need to look more like an agreement around nuclear materials, biological weapons, or other dual-use technologies: shared definitions, reporting requirements, inspections, controls on key components, and consequences for companies and governments that violate the agreement.

That comparison is imperfect. Robots are not uranium, and many of the relevant components are ordinary commercial goods. The line will be much harder to draw. But difficulty defining and enforcing a boundary is not evidence that no boundary is needed. It may instead mean we should start the work before the systems are ubiquitous.

At a minimum, governments could require registration and external evaluation for projects combining frontier AI with autonomous control of factories, laboratories, logistics networks, drones, or general-purpose robots. Cloud providers and advanced chip manufacturers could be required to report training or deployment patterns associated with these systems. Robotics companies could be prohibited from giving general-purpose agents unrestricted credentials, payment authority, or remote control over fleets. Manufacturers could be required to build in physical limits that software alone cannot remove.

The broader principle

I believe that the risk from self-sovereign AI and rogue multi-agent collectives is already serious when the AI only inhabits the digital realm. Both Chinese and U.S. companies are racing toward increasingly general-purpose robots, while AI systems are becoming more capable of planning, coding, persuasion, and autonomous action. These two curves are likely to meet.

Perhaps the moratorium I’m describing is too broad. Perhaps “internet-connected” is the wrong boundary, since a disconnected robot could still be dangerous and a connected one could be tightly constrained. Perhaps the real boundary is resource acquisition, manufacturing access, or the ability to increase the number of systems acting toward the same objective. I’m not confident about the exact legal category.

But I do think there is a capability we should decide not to build by accident: an intelligent system with a body, a bank account, access to the world’s machines, and the ability to make more of itself.

Once that loop is closed, it may be very hard to open it again.

Back to Writing & Ideas