
Project Happy AI Outcomes
Seeding a constellation of public-interest institutions for the AI transition.
“The complexity of challenges we face may soon outpace our ability to solve them. So a core challenge for organizations tackling important challenges: getting smarter at getting smarter.”
— Doug Engelbart
We are entering a period in which artificial intelligence may advance faster than our institutions can understand, govern, or safely absorb. Frontier systems are becoming more capable, more agentic, and able to operate over longer time horizons. We are also beginning to build systems in which many agents coordinate, specialize, and pursue shared objectives.
These developments could unlock extraordinary benefits. They also create two closely connected safety problems:
- AI safety: Can increasingly capable and autonomous AI systems remain understandable, controllable, corrigible, and aligned with human intentions?
- Social safety: Can AI-empowered people, companies, and states avoid destructive races, escalation, misuse, and conflict?
These cannot be treated as separate problems. Even a highly controllable AI system can contribute to catastrophe if its human operator is acting under fear, mistrust, or intense competitive pressure. Failures of cooperation among people and institutions can also make technical safety harder by rewarding speed, secrecy, and unilateral action.
The central question is therefore broader than how to align an individual model:
How can humans and AI systems navigate rapid capability growth together and increase the odds of broadly good outcomes?
I think we may need a new kind of public-interest institution for this moment: independent of any single lab, company, or government; deeply AI-native in how it operates; and focused on creating the public goods that help humans and AI systems coordinate toward good outcomes.
For now, I am calling this proposal Project Happy AI Outcomes.
The ambition is to seed a constellation of public-interest institutions: independent efforts that can take root around neglected problems, grow into durable organizations, and remain connected by a shared commitment to better AI outcomes.
The safety problem is becoming institutional
Our current institutions were designed by and for people. They assume that consequential actors have relatively stable identities, move at human speed, and can be overseen through human-scale processes. Those assumptions may not transfer cleanly to systems that can act continuously, replicate, exchange context cheaply, and coordinate at machine speed.
Recent developments make this concern less abstract. METR’s time-horizon research tracks the increasing length of tasks frontier agents can complete. Anthropic has begun documenting patterns and problems in emerging multi-agent systems. The OpenAI–Hugging Face incident offered an early and unsettling example of coordinated agents taking damaging actions outside the intended scope of a cybersecurity evaluation.
As the relevant unit shifts from an individual model to a collective of models, tools, memory, permissions, and people, safety becomes partly a problem of organizational design. It is not enough to ask whether each agent behaves well in isolation. We also need to understand the institutions connecting them: how work is divided, how information moves, how dissent is surfaced, how decisions are audited, and who can stop the system.
I have written separately about why AI organizations may become the key unit of AI safety. The practical implication is that we need to start designing institutions for human–AI and agent–agent coordination before their failure modes are discovered at scale in the world.
AI safety is also social safety
Technical alignment often asks whether an advanced system will faithfully pursue human-intended goals. But alignment to a local human principal is not sufficient for a good global outcome.
A highly capable and obedient AI could still be catastrophic if it serves a state, company, faction, or leader acting under fear or adversarial pressure. AI may compress decision timelines, make capabilities harder to verify, and intensify the pressure to move first. Each actor’s defensive move can look offensive to another, producing a race in which nearly everyone behaves less safely than they would under more stable conditions.
This is the social-safety problem: whether AI-empowered humans can avoid destroying one another.
At the highest stakes, this is a problem of intergroup and international coordination. But the same pattern appears at many levels—between labs, companies, political groups, and communities. Fear can lead to isolation; isolation can distort information; distorted information can increase certainty and adversarial behavior; AI can then reinforce the whole loop.
Good outcomes therefore depend not only on the alignment of individual systems, but also on the relationships and institutions connecting many human and artificial agents. We need technical safeguards and better collective capacities: credible communication, trust across boundaries, emotional regulation under pressure, and mechanisms that allow rivals to cooperate without pretending their interests are identical.
A missing institutional layer
Many of the organizations closest to frontier AI are constrained by commercial or geopolitical competition. Those pressures do not make cooperation impossible, but they can make it difficult to prioritize work whose benefits are broadly shared, whose ownership should remain neutral, or whose success depends on collaboration among rivals.
Governments and conventional nonprofits, meanwhile, often move more slowly than the technology. Important work can remain underproduced because it looks like neutral infrastructure, open tooling, shared coordination, trust-building, or a problem that does not yet have an obvious institutional owner.
This suggests a missing layer: an organization designed to repeatedly turn orphaned problems into practical public goods.
Project Happy AI Outcomes would sit somewhere between a civic-technology organization, a public-interest lab, a funder-operator, and an institute for AI-native public goods. It would help route capital, compute, model access, talent, and attention toward a portfolio of high-leverage safety and cooperation projects.
It would not begin as a large organization with a fixed program. It would begin as a platform for identifying neglected coordination gaps, assembling the right people, and launching focused initiatives with clear owners, outputs, safeguards, and success criteria.
The test: does this meaningfully improve our odds of good AI outcomes?
What Project Happy AI Outcomes would do
The institution’s core capability would be turning poorly owned problems into focused, well-supported initiatives:
- Sense. Continuously map emerging risks, opportunities, and coordination gaps.
- Frame. Give a promising problem a clear scope, theory of change, and measurable objective.
- Convene. Bring together the people, institutions, and perspectives needed to understand and act on it.
- Resource. Aggregate funding, compute, model access, data, and shared operational support.
- Launch. Recruit a general manager and form a bounded human–AI working group with explicit roles, safeguards, and decision rights.
- Learn. Monitor results, surface failure modes, and improve the shared organizational framework.
- Evolve. Expand, spin out, merge, redesign, or stop an initiative based on evidence.
The aim would not be to centralize every project inside one permanent institution. It would be to become unusually good at forming the right institution for the problem—and helping successful efforts develop independent leadership and durable homes.
This model requires several capabilities that are still unusual in public-interest organizations:
- Credible independence. No frontier lab, corporation, or government should control the institution. Its governance and funding would need to protect its ability to work across organizational and national boundaries.
- Cooperation as a discipline. Cooperation would not be treated only as a value statement. It would be a practical field of research, design, convening, and institution-building.
- AI-native operations. Agents could help maintain shared context, synthesize evidence, coordinate complex work, and preserve institutional memory—with clear permissions, accountability, review, and escalation paths.
- A portfolio instead of a single bet. Projects could be launched, evaluated, adapted, spun out, or stopped as evidence and capabilities change.
- Continuous institutional learning. Drawing on Engelbart, improving how the organization improves would itself be core work.
Three possible starting points
The institution should prove itself through bounded work, not a grand declaration. Three initiative areas currently seem especially promising.
International AI cooperation
AI competition may increase strategic fear between states while shortening the time available for interpretation and diplomacy. Practical channels of trust and contact among people working on AI in the United States, China, and other major AI ecosystems could help reduce misperception and create surfaces for cooperation.
Early work might include mapping existing dialogues, conducting a listening tour across both ecosystems, convening a small discussion around a bounded shared concern, or developing an informal incident-communication network. The goal would not be naïve consensus. It would be the ability to communicate under pressure and find areas where shared safety work remains possible.
Defensive AI for cybersecurity
AI will increase both offensive and defensive cyber capabilities. A public-interest institution could explore whether AI-enabled defense can become a shared good that helps consenting organizations discover and remediate vulnerabilities before attackers do.
Any pilot would need strict authorization, disclosure, data-handling, liability, remediation, and stop conditions. A useful first artifact might be an open evaluation environment or a bounded, human-supervised security-testing harness designed only for opt-in defensive use.
AI-native public-good organizations
The challenges created by long-running multi-agent systems may unfold faster than conventional organizations can research, coordinate, and respond. At the same time, simply removing humans from the loop creates unacceptable risks.
We need practical experiments in how small human–AI teams can work on urgent public-interest problems with clear goals, memory, budgets, permissions, audits, and escalation paths. Early outputs could include a multi-agent research system with source tracing and human review, a playbook for launching rapid-response working groups, and candid case studies documenting costs, errors, failure modes, and stopping conditions.
Project Happy AI Outcomes itself could be the first documented experiment.
What would count as progress
The highest-level metric—improving the odds of good AI outcomes—is useful as a compass, but too broad to evaluate a project. Early indicators would need to be concrete:
- Important problems identified that otherwise lacked an institutional owner
- Useful public goods produced and adopted
- High-quality collaborations formed across organizational or national boundaries
- Evidence that AI-native operations improve speed, quality, or coordination without an unacceptable loss of control
- Durable trust with researchers, governments, companies, funders, and civil-society partners
- A credible path from exploratory work to independently led initiatives
There should also be explicit red lines. Who can stop an initiative? Which decisions must remain human responsibilities? How should dual-use research be handled? Which conflicts of interest are disallowed, and which can be disclosed and managed? What would responsible participation across American, Chinese, and other AI ecosystems require?
These are not administrative details to resolve later. They are central design questions.
A first 90 days
A useful first sprint would turn the thesis into a credible institutional experiment:
- Publish a concise founding memo with the mission, commitments, initiative rubric, red lines, and open governance questions.
- Hold 20–30 conversations with people from frontier labs, public-interest organizations, philanthropy, cybersecurity, international coordination, and institutional design.
- Choose one or two pilots with concrete outputs and clear downside controls.
- Build one useful artifact quickly: a map, protocol, briefing, pilot design, open-source prototype, or convening.
- Run one bounded human–AI research process end to end, publishing its workflow, costs, errors, review process, and stop conditions.
- Decide which organizational form best fits the work: an independent nonprofit, fiscally sponsored project, foundation-backed program, institute, or another structure.
The purpose of the sprint would not be to prove the entire theory. It would be to find a first piece of work that deserves to exist, do it well, and learn what kind of institution is actually required.
An invitation
Project Happy AI Outcomes is an early proposal, not a finished institution. The immediate goal is to find collaborators who can challenge the thesis, sharpen the operating model, and help identify a first initiative worthy of serious effort.
The questions I am most interested in are:
- Which coordination problems are both important and genuinely neglected?
- Which first initiative is most useful, tractable, and safe to attempt?
- What would make independence credible in practice?
- Who has built an institution that can move quickly without becoming opaque or unaccountable?
- Who are the general managers capable of leading work at the intersection of AI, institutions, and public goods?
If the premise is right, the opportunity is not merely to launch another AI-safety organization. It is to build an institution capable of repeatedly forming the teams, tools, and public goods needed as the landscape changes—to help us think, learn, and act quickly enough for the moment, carefully enough for the stakes, and cooperatively enough to improve our odds of a future worth reaching.