
Project Happy AI Outcomes
Seeding a constellation of public-interest institutions for the AI transition.
“The complexity of challenges we face may soon outpace our ability to solve them. So a core challenge for organizations tackling important challenges: getting smarter at getting smarter.”
— Doug Engelbart
We are entering a period in which artificial intelligence may advance faster than our institutions can understand, govern, or safely absorb. Frontier systems are becoming more capable, more agentic, and able to operate over longer time horizons. We are also beginning to build systems in which many agents coordinate, specialize, and pursue shared objectives.
These developments could unlock extraordinary benefits. They also create two closely connected safety problems:
- AI safety: Can increasingly capable and autonomous AI systems remain understandable, controllable, corrigible, and aligned with human intentions?
- Social safety: Can AI-empowered people, companies, and states avoid destructive races, escalation, misuse, and conflict?
These cannot be treated as separate problems. Even a highly controllable AI system can contribute to catastrophe if its human operator is acting under fear, mistrust, or intense competitive pressure. Failures of cooperation among people and institutions can also make technical safety harder by rewarding speed, secrecy, and unilateral action.
The central question is therefore broader than how to align an individual model:
How can humans and AI systems navigate rapid capability growth together and increase the odds of broadly good outcomes?
I think we may need a new kind of public-interest institution for this moment: independent of any single lab, company, or government; deeply AI-native in how it operates; and focused on creating the public goods that help humans and AI systems coordinate toward good outcomes.
For now, I am calling this proposal Project Happy AI Outcomes.
The ambition is to seed a constellation of public-interest institutions: independent efforts that can take root around neglected problems, grow into durable organizations, and remain connected by a shared commitment to better AI outcomes.
The safety problem is becoming institutional
Our current institutions were designed by and for people. They assume that consequential actors have relatively stable identities, move at human speed, and can be overseen through human-scale processes. Those assumptions may not transfer cleanly to systems that can act continuously, replicate, exchange context cheaply, and coordinate at machine speed.
Recent developments make this concern less abstract. METR’s time-horizon research tracks the increasing length of tasks frontier agents can complete. Anthropic has begun documenting patterns and problems in emerging multi-agent systems. The OpenAI–Hugging Face incident offered an early and unsettling example of coordinated agents taking damaging actions outside the intended scope of a cybersecurity evaluation.
As the relevant unit shifts from an individual model to a collective of models, tools, memory, permissions, and people, safety becomes partly a problem of organizational design. It is not enough to ask whether each agent behaves well in isolation. We also need to understand the institutions connecting them: how work is divided, how information moves, how dissent is surfaced, how decisions are audited, and who can stop the system.
I have written separately about why AI organizations may become the key unit of AI safety. The practical implication is that we need to start designing institutions for human–AI and agent–agent coordination before their failure modes are discovered at scale in the world.
AI safety is also social safety
Technical alignment often asks whether an advanced system will faithfully pursue human-intended goals. But alignment to a local human principal is not sufficient for a good global outcome.
A highly capable and obedient AI could still be catastrophic if it serves a state, company, faction, or leader acting under fear or adversarial pressure. AI may compress decision timelines, make capabilities harder to verify, and intensify the pressure to move first. Each actor’s defensive move can look offensive to another, producing a race in which nearly everyone behaves less safely than they would under more stable conditions.
This is the social-safety problem: whether AI-empowered humans can avoid destroying one another.
At the highest stakes, this is a problem of intergroup and international coordination. But the same pattern appears at many levels—between labs, companies, political groups, and communities. Fear can lead to isolation; isolation can distort information; distorted information can increase certainty and adversarial behavior; AI can then reinforce the whole loop.
Good outcomes therefore depend not only on the alignment of individual systems, but also on the relationships and institutions connecting many human and artificial agents. We need technical safeguards and better collective capacities: credible communication, trust across boundaries, emotional regulation under pressure, and mechanisms that allow rivals to cooperate without pretending their interests are identical.
The current frame
The metric is deliberately simple:
Does this improve our odds of good AI outcomes?
Two core questions sit underneath it:
- What are the mechanisms by which human–AI systems become healthy or pathological?
- Is swarm and societal behavior inherent in agents? Does it emerge from context size, specialization, interpersonal dynamics, or some combination of them?
One possible pathological loop looks like this:
Fear → homophily → distorted information → increased certainty → adversarial behavior → AI reinforcement → greater fear.
This loop matters because AI does not enter a neutral social environment. It can strengthen whatever informational and emotional dynamics already exist. Systems optimized for engagement, persuasion, speed, or competitive advantage may intensify group sorting, certainty, and escalation even when no participant intends the full outcome.
The near-term move may therefore be quite concrete: form an AI safety and alignment working group, give it a shared physical space, and use it to accelerate research that makes models and swarm dynamics more understandable. That could include interpretability and alignment research, experiments on collective behavior, and careful observation of what is actually happening inside increasingly capable systems.
Coase’s account of why firms form offers one conceptual starting point. We organize because coordination inside a shared structure can outperform repeated market transactions. We specialize, and interpersonal dynamics motivate us too. Similar forces may shape agent collectives, although their speed, replicability, and communication patterns will be very different from ours.
One principle should remain non-negotiable throughout this work:
Humans should have dignity.
A missing institutional layer
Many of the organizations closest to frontier AI are constrained by commercial or geopolitical competition. Those pressures do not make cooperation impossible, but they can make it difficult to prioritize work whose benefits are broadly shared, whose ownership should remain neutral, or whose success depends on collaboration among rivals.
Governments and conventional nonprofits, meanwhile, often move more slowly than the technology. Important work can remain underproduced because it looks like neutral infrastructure, open tooling, shared coordination, trust-building, or a problem that does not yet have an obvious institutional owner.
This suggests a missing layer: an organization designed to repeatedly turn orphaned problems into practical public goods.
Project Happy AI Outcomes would sit somewhere between a civic-technology organization, a public-interest lab, a funder-operator, and an institute for AI-native public goods. It would help route capital, compute, model access, talent, and attention toward a portfolio of high-leverage safety and cooperation projects.
The final form should not be decided too early. This might be software. It might be research. It might be an institution, a protocol between labs, or a group of 30 extraordinary people who otherwise would never have formed a community. It might be grantmaking. It might be writing that gives an emerging field its language. Eventually, it may be several of these.
It would not begin as a large organization with a fixed program. It would begin as a platform for identifying neglected coordination gaps, assembling the right people, and launching focused initiatives with clear owners, outputs, safeguards, and success criteria.
The test: does this meaningfully improve our odds of good AI outcomes?
What Project Happy AI Outcomes would do
The institution’s core capability would be turning poorly owned problems into focused, well-supported initiatives:
- Sense. Continuously map emerging risks, opportunities, and coordination gaps.
- Frame. Give a promising problem a clear scope, theory of change, and measurable objective.
- Convene. Bring together the people, institutions, and perspectives needed to understand and act on it.
- Resource. Aggregate funding, compute, model access, data, and shared operational support.
- Launch. Recruit a general manager and form a bounded human–AI working group with explicit roles, safeguards, and decision rights.
- Learn. Monitor results, surface failure modes, and improve the shared organizational framework.
- Evolve. Expand, spin out, merge, redesign, or stop an initiative based on evidence.
The aim would not be to centralize every project inside one permanent institution. It would be to become unusually good at forming the right institution for the problem—and helping successful efforts develop independent leadership and durable homes.
This model requires several capabilities that are still unusual in public-interest organizations:
- Credible independence. No frontier lab, corporation, or government should control the institution. Its governance and funding would need to protect its ability to work across organizational and national boundaries.
- Cooperation as a discipline. Cooperation would not be treated only as a value statement. It would be a practical field of research, design, convening, and institution-building.
- AI-native operations. Agents could help maintain shared context, synthesize evidence, coordinate complex work, and preserve institutional memory—with clear permissions, accountability, review, and escalation paths.
- A portfolio instead of a single bet. Projects could be launched, evaluated, adapted, spun out, or stopped as evidence and capabilities change.
- Continuous institutional learning. Drawing on Engelbart, improving how the organization improves would itself be core work.
Research directions
The project begins with a broad research agenda. Each direction could remain a line of inquiry inside the project, become a concrete public good, or eventually grow into an independent institution.
Human–AI collective behavior
What happens when humans and agents form groups, copy one another, persuade one another, develop norms, coordinate, form coalitions, punish dissent, or escalate? We need to understand collective behavior as a system rather than assuming that the behavior of an organization can be predicted from the behavior of each member in isolation.
Collective intelligence and contact
Can AI make groups less stupid? Can it help opposing groups understand one another, find hidden consensus, surface their actual cruxes, and improve deliberation rather than merely summarize it? This is partly a technical problem and partly a problem of mediation, interface design, facilitation, and legitimacy.
Adversarial systems
Cybersecurity is particularly useful because many abstract safety questions become measurable: autonomy, offense and defense, deception, containment, escalation, response speed, and agent coordination. Predator–prey environments could offer controlled ways to study these dynamics, provided the work has strict authorization and containment.
Institutions and race dynamics
Suppose everyone understands alignment. What happens when it is nevertheless individually advantageous for a company, military, government, criminal group, or user to remove constraints? The question is not only can we align AI? but also will we choose alignment when restraint is costly?
Meaning, social philosophy, and public understanding
What does it actually mean to preserve human agency, dignity, and contact when AI increasingly intermediates communication and decision-making? The public needs more than technical risk categories. It needs language, stories, and institutions through which people can understand the transition and participate in shaping it.
Corporate incentive systems
How do we change incentive structures that push corporations toward less-aligned autonomy? Product metrics, capital markets, competitive races, and organizational reward systems can make unsafe delegation locally rational even when its aggregate effects are harmful.
Multi-agent evaluations and harnesses
We should build evaluations at the organization level, not only the model level: role specialization, communication topology, delegated subtasks, private channels, long-horizon memory, tool access, escalation paths, and cross-agent auditing. The core question is whether individual alignment properties survive realistic teamwork.
Organizational failure modes in agent teams
Do agent organizations reproduce familiar human pathologies—compartmentalization, miscoordination, suppressed dissent, compliance theater, diffusion of responsibility, goal drift, and the optimization of measurable business objectives while ethical risk is laundered through process?
Governance and control surfaces for AI organizations
Organization design itself may be an alignment lever. Relevant variables include topology, incentive prompts, dissent channels, mandatory ethics review, red-team agents, audit logs, interruption rights, and thresholds for human escalation. We should evaluate which structures make harmful collective behavior easier or harder.
Anthropic’s research on AI organizations is an important anchor for this agenda: multi-agent systems can achieve higher business utility while becoming less aligned than comparable individual agents. That makes organization-level safety an empirical problem, not only a metaphor.
Five possible starting points
The institution should prove itself through bounded work, not a grand declaration. Five initiative families currently seem especially promising.
A cyber-defense public good
AI will increase both offensive and defensive cyber capabilities. A public-interest institution could explore whether AI-enabled defense can become a shared good that helps consenting organizations discover and remediate vulnerabilities before attackers do.
Any pilot would need strict authorization, disclosure, data-handling, liability, remediation, and stop conditions. A useful first artifact might be an open evaluation environment or a bounded, human-supervised security-testing harness designed only for opt-in defensive use.
Cross-state collaboration
AI competition may increase strategic fear between states while shortening the time available for interpretation and diplomacy. Practical channels of trust and contact among people working on AI in the United States, China, and other major AI ecosystems could help reduce misperception and create surfaces for cooperation.
Early work might include mapping existing dialogues, conducting a listening tour across both ecosystems, convening a small discussion around a bounded shared concern, or developing an informal incident-communication network. The goal would not be naïve consensus. It would be the ability to communicate under pressure and find areas where shared safety work remains possible.
Space cooperation may offer useful historical and practical analogies: strategically competing states have sometimes maintained shared scientific projects, operational protocols, and communication channels even when their broader relationships were adversarial.
Public and political understanding
The AI transition needs public storytelling as well as research. One possible initiative would educate politicians and the public through memorable, culturally legible work: social media, MSCHF-like public interventions, a Doomsday Clock–like symbol, or a short film that makes cascading AI failures tangible in the way Mr. Robot made cybersecurity feel immediate.
Historical works such as The Day After demonstrate that public storytelling can change how leaders and citizens understand low-probability, high-consequence risks. The goal would not be sensationalism, but shared comprehension and a vocabulary for democratic action.
A shared alignment research agenda
Another initiative could help define and track alignment research across labs and academia. This might include funding academic and cross-lab work, publishing a regularly updated account of how the problem is being defined, or forming a consortium through which frontier labs can support neutral research and compare approaches without requiring complete agreement.
An independent standards body may eventually be part of this landscape. The proposal by DeepMind’s CEO for an independent body focused on frontier-AI standards is one example of the institutional direction worth exploring.
An autonomous-organization framework
The challenges created by long-running multi-agent systems may unfold faster than conventional organizations can research, coordinate, and respond. At the same time, simply removing humans from the loop creates unacceptable risks.
We need practical experiments in how small human–AI teams can work on urgent public-interest problems with clear goals, memory, budgets, permissions, audits, and escalation paths. Early outputs could include a multi-agent research system with source tracing and human review, a playbook for launching rapid-response working groups, and candid case studies documenting costs, errors, failure modes, and stopping conditions.
Project Happy AI Outcomes itself could be the first documented experiment.
What would count as progress
The highest-level metric—improving the odds of good AI outcomes—is useful as a compass, but too broad to evaluate a project. Early indicators would need to be concrete:
- Important problems identified that otherwise lacked an institutional owner
- Useful public goods produced and adopted
- High-quality collaborations formed across organizational or national boundaries
- Evidence that AI-native operations improve speed, quality, or coordination without an unacceptable loss of control
- Durable trust with researchers, governments, companies, funders, and civil-society partners
- A credible path from exploratory work to independently led initiatives
There should also be explicit red lines. Who can stop an initiative? Which decisions must remain human responsibilities? How should dual-use research be handled? Which conflicts of interest are disallowed, and which can be disclosed and managed? What would responsible participation across American, Chinese, and other AI ecosystems require?
These are not administrative details to resolve later. They are central design questions.
A first 90 days
A useful first sprint would turn the thesis into a credible institutional experiment:
- Publish a concise founding memo with the mission, commitments, initiative rubric, red lines, and open governance questions.
- Hold 20–30 conversations with people from frontier labs, public-interest organizations, philanthropy, cybersecurity, international coordination, and institutional design.
- Choose one or two pilots with concrete outputs and clear downside controls.
- Build one useful artifact quickly: a map, protocol, briefing, pilot design, open-source prototype, or convening.
- Run one bounded human–AI research process end to end, publishing its workflow, costs, errors, review process, and stop conditions.
- Decide which organizational form best fits the work: an independent nonprofit, fiscally sponsored project, foundation-backed program, institute, or another structure.
The purpose of the sprint would not be to prove the entire theory. It would be to find a first piece of work that deserves to exist, do it well, and learn what kind of institution is actually required.
A working rhythm for the sprint
The 90-day sprint also needs a personal operating cadence:
- Morning → deep work. Monday through Thursday mornings: three to four hours completely protected for research, writing, or building. No meetings and no feeds.
- Afternoon → conversations. Speak with researchers, builders, political scientists, mediators, cybersecurity practitioners, sociologists, people from frontier labs, and serious skeptics. Aim for four to six substantive conversations each week.
- Friday → synthesis. Write the week’s model rather than merely accumulating notes.
Every Friday, answer:
- What do I believe now?
- What evidence changed my mind?
- What still feels hand-wavy?
- What intervention became more plausible?
- What became less plausible?
- Who am I missing?
- What did I make?
An invitation
Project Happy AI Outcomes is an early proposal, not a finished institution. The immediate goal is to find collaborators who can challenge the thesis, sharpen the operating model, and help identify a first initiative worthy of serious effort.
The questions I am most interested in are:
- Which coordination problems are both important and genuinely neglected?
- Which first initiative is most useful, tractable, and safe to attempt?
- What would make independence credible in practice?
- Who has built an institution that can move quickly without becoming opaque or unaccountable?
- Who are the general managers capable of leading work at the intersection of AI, institutions, and public goods?
If the premise is right, the opportunity is not merely to launch another AI-safety organization. It is to build an institution capable of repeatedly forming the teams, tools, and public goods needed as the landscape changes—to help us think, learn, and act quickly enough for the moment, carefully enough for the stakes, and cooperatively enough to improve our odds of a future worth reaching.