← All Paramerge writing

Paramerge Writing

EssaySave

Steering Together

How Organizations, Professions, and Nations Can Coordinate on AI

Stephen Lieberman

Paramerge, Real-World AI Governance Center

August 2026

stephen@paramerge.com · paramerge.com

The hard problem is not the one we are funding

Somewhere right now, a coalition is signing an agreement about AI. Picture one of its rooms. A hospital system, a county agency, a professional body, and a ministry have spent a year negotiating a shared framework, and the program director who carried it, a woman who has spent twenty years getting institutions to work together, has the document open to the definitions section she rewrote four times, and watches the signatures with the relief of someone who believes the hard part is over. The partners have planned and budgeted on the belief that they mean the same things by the words in the document: safety, evidence, oversight, harm. That belief can fail, and when it fails it fails late, at the moment the words are asked to bear weight. A term everyone had been using turns out to have carried three incompatible meanings. What one partner counts as evidence another calls anecdote. What one calls safety another experiences as harm. The program stalls, the funding pauses, and everyone at the table is certain the others changed the deal.

Nobody changed the deal. The deal was never shared.

It only looked shared, and no instrument existed that could have shown the director otherwise in time. Building capable AI has turned out to be easier than agreeing on where it should take us. I wrote that sentence at the close of an earlier Paramerge paper, and this essay takes it seriously as an engineering problem rather than a lament. The obstacle it names is not model capability, and it is not model alignment, because it would remain even if alignment succeeded. It is coordination: the capacity of organizations, professions, nations, and international bodies to reach agreements about AI that are actually shared, durable, and carried out.

The AI safety literature itself points here. Anwar and colleagues (2024), surveying the foundational challenges in assuring the alignment and safety of large language models, identify a class of problems that training procedures alone cannot solve: no consensus on which values these systems should encode, no adequate scheme for culpability when they cause harm, no sufficient framework for governing their use once they are embedded in institutions. I read those as constitutively social problems. They require deliberation among the people and institutions affected, and they stay unsolved while that deliberation stays impossible at scale. Bengio and colleagues (2024), writing in Science with two dozen senior researchers across AI and adjacent fields, argue that society's response to AI risk is not commensurate with the pace of progress, and that what is needed is governance capable of adapting as the systems change rather than fixed rules the systems will outrun. Adaptive governance, on my reading, is coordination sustained over time. Neither paper takes the next step. This one does. The field has invested enormously in aligning models. The investment in instruments that would let institutions align with each other about the models is a fraction of that. Structured elicitation traditions have worked on pieces of it for decades, from Q methodology and the policy Delphi to group model building, concept mapping, and deliberative polling, and they work in facilitated sessions, on stated positions, at workshop scale. What I have not found among the nearest neighbors to this program, surveyed in its companion papers, is the combination: operative meanings read from the records of actual practice, with per-claim provenance, coupled to a population testbed. That is a survey of neighbors and not a universal negative.

This is a wicked problem in the exact sense Rittel and Webber (1973) gave the phrase: problem definitions entangled with proposed solutions, stakeholders holding contested values, knowledge incomplete and shifting. Wicked problems do not have solutions the way equations do. They have better and worse ways of being navigated. The next claim is mine, and Rittel and Webber bear no responsibility for it: in the AI case the difference will be decided substantially by the quality of the coordination among the people doing the navigating.

Why coordination fails

If coordination were merely difficult, exhortation would fix it. It fails for structural reasons, and a serious program has to name them before it builds anything.

The first reason is the failure in the opening scene. I have called these problematic latent epistemological differences, or PLEDs (Lieberman, 2023). They are epistemological because they concern what participants will accept as grounds for a claim. They are latent because naive realism, the tendency to experience one's own perception as the way things are, keeps them invisible until circumstances force them into the open (Gilovich, Griffin, and Kahneman, 2002). And they are problematic because when they surface late they do not merely slow a collaboration. They end it, and the ending is borne by people: the director who spent a year building trust that the collapse spends in a week, and the people downstream whom the program existed to serve. You cannot ask people to list assumptions they do not know they hold. The pattern runs from a single agency to the top of the international order. Brincat (2015) reads the failures of global climate justice as an instance of the same shape: nations and international bodies holding conceptions of justice that do not connect. AI governance, a coordination problem at least as value-laden and at least as global, is unlikely to be exempt.

The shape is familiar from medicine. A patient transfers between two hospitals, and both charts say she is on an anticoagulant. Both teams are competent, both records are current, and both are reading the same word. Underneath the word the dose is different and the indication is different, because each institution wrote its protocol for its own population and never had reason to write down what it was assuming. Nobody is wrong. Nobody is concealing anything. The discrepancy stays invisible to everyone who could correct it until the morning of the procedure, when it becomes a bleed. Hospitals answer that with medication reconciliation, a step built to compare what each side means by the same prescription before anyone acts on it. A signed framework is the same document held by four institutions, and it has no reconciliation step. The bleed comes in month fourteen.

The second reason is that our standard picture of agreement ignores exit. The most visible deliberation and consensus systems, including the new AI-mediated ones, take a fixed set of participants as given. Candidate statements are generated, preferences are aggregated, a winner is declared, and everyone is presumed bound. In the real world participants have outside options, and an agreement endorsed under a forced choice can fail the moment people are free to pursue a smaller coalition that fits them better. Hirschman (1970) named the choice between exit and voice, and his own warning cuts both ways here, because making exit salient can atrophy voice. Bargaining theory sharpened the other half. In the two-party models collected by Osborne and Rubinstein (1990), an outside option affects a bargain only when it binds, and when it binds it sets the terms, though when every side holds an outside option, which is the governance case, the theory grows less sharp. The formal ancestry runs to Nash (1950), whose bargaining problem prices every deal against what each party would get without one. I should also state what is not known. No field measurement of exit from a deployed deliberation system appears to exist, so this second reason rests on bargaining and coalition theory, not on measured failures. An agreement about AI that has not priced anyone's willingness to walk away may hold anyway. But its durability has been tested against nothing. It is a photograph of a room.

The third reason is scale. Surfacing every party's operative meaning of every load-bearing concept has always exceeded any facilitation budget. Twenty organizations and a few hundred load-bearing concepts mean thousands of cells of skilled interpretive work, so the standard response has been to not try. Organizations gamble instead. Gambling worked, more or less, when the systems being governed changed slowly. Nothing about AI changes slowly.

And there is a fourth reason, which is timing, and here I am reporting my own reading of the landscape rather than a measurement, because complex systems announce their thresholds mostly in hindsight. Complex adaptive systems approach tipping points, thresholds at which gradual change produces sudden reorganization that is hard to reverse. Three pressures in AI governance look to me like the approach to one. Ungoverned adoption is normalizing toward the point where governance feels optional. Crisis-driven regulation tends to arrive faster than deliberate design can inform it. And the governance conversation is being framed, by default, by the builders of the systems being governed. If that reading is right, then what the landscape reorganizes into depends in part on which instruments exist when the reorganization comes.

What would make coordination possible

What would have to exist for the director to have seen this before she signed? The argument of this essay is that the instruments can now be built, and that they are AI instruments. Not AI that decides for us. AI that lets us see what we are actually disagreeing about, rehearse what we are about to commit to, and reach agreements designed to survive the freedom of the people who make them. Four components make up the program. They are at different stages of maturity, and I will state the stage of each exactly, because a coordination instrument that overstates its own readiness would be a small example of the problem it exists to solve. No component of this program has been deployed for the use described here. That is the honest first line of the ledger.

Map the disagreement before it strikes. The hypothesis, and it is a hypothesis this program commits to testing rather than a demonstrated capability, is that a modern language model can read every party's documentary record, ingested only with that party's consent, and compare how each party actually uses the concepts an agreement will stand on, with every finding tied to inspectable source passages. The output is not a ranking and not an arbiter's verdict. It is a map of where operative meanings align and diverge, built from a concept list the parties set rather than whoever commissions the map, and delivered before the first joint decision instead of at month fourteen. The full architecture, its failure modes, and its build order are specified in a companion Paramerge paper, which also records the caveats that matter most. An announced instrument changes the record it reads, because a sophisticated party then knows which sentences to rewrite. And the premise beneath the whole layer, that surfacing a latent divergence early improves the outcome, has documented counterexamples, because some collaborations are held together precisely by the plasticity of a shared term (Star and Griesemer, 1989; Sunstein, 1995). So the decision about a surfaced divergence belongs to the parties, including the decision to leave it alone. What the component changes is the moment of choice. The director in the opening scene reads the divergence in month one, on the record, and decides then what it means, instead of discovering it in month fourteen, in the wreckage.

Rehearse before committing. Once the divergences are visible, partners can propose a shared framework and stress-test it against a simulated population grounded in real data, before committing real people and real resources. This is the in silico testbed. Its foundation is the CASS framework, the Complex Adaptive Social System agent-based simulation program I founded and led at the Naval Postgraduate School, which builds artificial societies from survey data (Lieberman, 2012) and whose original military intended use went through the Department of Defense verification, validation, and accreditation discipline (Alt, Jackson, Hudak, and Lieberman, 2009; Alt, Lieberman, and Blais, 2010). That accreditation attaches to the original model and purpose. It does not transfer to the research reimplementation now in use, and it would not extend to rehearsing coordination frameworks anyway. The testbed's fitness for this use is exactly what has to be established, in a stated order, against stated benchmarks.

Language-model agents can give these rehearsals the texture of deliberation, an integration that is designed and largely untested, and the evidence about such agents is real but bounded, so it is worth stating precisely. In a study in Science, group statements drafted by an AI mediator were preferred by human deliberators over those of trained lay human mediators 56 percent of the time (Tessler et al., 2024). That is evidence for AI-assisted mediation of human deliberation, not for agents standing in for people. In a preprint, language-model agents built from in-depth interviews reproduced participants' held-out survey answers at 83 percent of those participants' own two-week test-retest consistency. The comparison is to the person and not to a right answer: people asked the same questions a second time do not fully agree with themselves, and the agents reached about five sixths of that. Agents given demographics alone reached 74 percent. The authors state that whether such agents reproduce how attitudes covary across a population, which is what a rehearsal leans on, remains unvalidated (Park et al., 2026). And the documented failure modes of simulated people all push in one direction. Variance collapse erases minority positions. Sycophancy manufactures agreement. Caricature distorts exactly the groups a coalition most needs to see accurately. Every one of these failures makes a simulated population look more agreeable than the real one, which is precisely what a coordination instrument must not do. So the discipline is fixed. A rehearsal output is a hypothesis, never a forecast. A rehearsed fracture is a lead worth investigating. A rehearsed consensus is not evidence that the framework holds.

Make exit a declared input. The third component treats willingness to walk away as data rather than as betrayal. Let each participant state what they would contribute to a candidate policy at each setting of its main parameter, and the level past which they contribute nothing. The aggregate is a priced curve, and beneath it sit the individual schedules where the group's divisions are actually located: who would leave at which setting, and where those exit points cluster. The mechanism is specified as a thought experiment in a companion paper and has not been built, and its hardest limits are recorded there in full. A declared walk-away point is free to misreport and impossible to verify, which makes the exit input itself an attack surface. And the mechanism's reach is bounded by excludability: it has force where a coalition's work is a club good that non-signers can be excluded from, and much of AI governance, especially at the national and international scale, is not excludable in that way. Within that scope, the point is this. Exit moves from the unspoken thing that ends agreements to a declared input the agreement was shaped against.

Keep every decision human. This one is not a component to be built. It is a rule, and it is in force from the first line of the program. The outputs of the first three components are inputs to human deliberation, not substitutes for it. This is a safety property, not a courtesy. Language-model text persuades people on policy issues about as effectively as text written by other lay people (Bai et al., 2025), so a system that can surface differences can also move them. A human decision layer is necessary but not sufficient, because what the humans are shown still shapes what they decide, and no rule in this program removes that residual influence. What the rule fixes is accountability. No output of this program is a recommendation. No agreement is valid because an instrument endorsed it. The accountability stays with the people, which is where it was always going to return when something breaks.

Three guiding principles

Underneath the components sit three commitments. They are what the instruments are for. Each is stated with the same maturity discipline as the components, because principles can overstate readiness too.

The first is a durability standard, and it comes from game theory. Nash (1951) formalized the equilibrium in which no participant improves their outcome by changing course alone. That standard bounds only solitary defection. The threat a Nash standard does not bound is the breakaway group, the subset that discovers it does better on its own, and ruling that out is a stronger demand, the one game theorists study under the name of the core. The durability standard this program aims at is stability against both: an agreement each party keeps because no alternative available to them, alone or with others, does better. Two honesty clauses bind the aim. The companion mechanism computes no equilibrium and inherits no core property, so a candidate agreement that survives its search has survived one partial search for breakaway coalitions, which is as much evidence as one search can give. And for some decisions no such agreement exists at all, in which case the honest output is the smallest relaxation the search could find, together with a statement of who bears it. Even so bounded, the standard changes what signing means. Endorsement in a room can be produced by any process that offers no alternative to endorsing. Durability outside the room cannot, and durability is the currency the signing itself cannot supply.

The second is an ethical standard, and it comes from political philosophy. Rawls (1971) built the original position, a device in which principles are chosen without knowledge of one's own place in society, and he built it to discipline which reasons count, not to describe what anyone knows. Rehearsal cannot build that device, and this program does not claim to. What it can build is one thing that argument concludes we should care about: a view of how a candidate framework lands across a grounded population before anyone signs. The rule this program takes from that tradition is modest and operational. A rehearsed harm concentrated at the margins is a first-class finding, and the burden of explaining it falls on the framework, not on the people it lands on. A rehearsal that finds no harm at the margins tells the coalition only that this rehearsal did not find it. The margins matter twice over, because the simulation failure modes above degrade accuracy exactly where the stakes concentrate, and because the people least represented at the table and the people worst off in the population are not always the same people, so both must be looked for. Where a subgroup is too thin to audit, the system must say so rather than report anyway. The director in the opening scene will not meet most of these people. They are downstream of every institution in her signing room, and a rehearsal is the only room they are in before she signs. The people hardest to simulate are the people governance most often fails, and an instrument that cannot see them honestly must at least know that it cannot.

The third is a method commitment: iterate in silico against the real world. No rehearsal is an oracle, and no agreement survives contact with a changing world unchanged. The systems being governed are non-stationary, and so are the coalitions governing them. The loop that works runs continuously. No iteration of that loop has been run for this purpose, and no scored record of rehearsal against outcome yet exists. Model, rehearse, decide, deploy, measure what actually happened, and feed the divergence between rehearsal and reality back into the next iteration of both the models and the agreement. A governance system built this way is designed to get better at its job by doing its job, which is the kind of governance a technology that will not hold still requires.

What the three principles add up to is a single test: judge an agreement by its durability outside the room, by what it does to the people at its margins, and by whether it keeps learning from the world, never by the ceremony of its signing. That test is also a personal statement. This is the work I have chosen to spend the coming years on, the program of the practice I run.

How it scales

None of this matters if it works once, in one room. The scaling mechanism is demonstration and cross-recognition.

A single rigorously documented demonstration changes what an adjacent institution can point to. A real coalition that mapped its divergences, rehearsed its framework, priced its exits, and reached an agreement that held gives professional bodies, standards organizations, public agencies, and international bodies something concrete to evaluate against their own coordination problems. What lets separate institutions coordinate without merging is cross-recognition: each adopting, adapting, or referencing work the others can inspect and defend. That mechanism is not my invention, and the polycentric governance tradition has described it for decades (Ostrom, 2010). Four conditions look to me necessary, and each is a design target. There must be a coherent shared reference the actors can recognize as authoritative. It must address the concerns each actor actually faces. It must be produced through a process each actor can defend inside its own institution, which is what genuine practitioner and community participation buys and why codesign is a structural requirement rather than a virtue (Costanza-Chock, 2020; Israel et al., 2005). And coordinating must be visibly better than acting alone.

The same logic runs to the top of the stack. Nations negotiating AI commitments face PLEDs in their purest form: conceptions of fairness, sovereignty, evidence, and risk that do not connect, and no international process that surfaces the divergence before it is load-bearing. Diplomacy, standards harmonization, and the growing network of safety institutes all work on aligning stated positions, and that work matters. What I have not found anywhere in that landscape is the combination this program builds: operative meanings mapped from records rather than statements, frameworks rehearsed against grounded populations before commitment, and exit priced as an input. The first two instruments do not become less applicable as the stakes rise. The third does, for the excludability reason stated above, and at that scale it degrades to a disagreement report rather than a stability test.

The third state

A coalition trying to steer AI can be in one of three states. In the first, nobody knows what the parties actually disagree about until the disagreement ends the effort. In the second, skilled facilitation and structured deliberation have surfaced the most visible differences, and everyone hopes the rest are small. The second state is the best that current practice reaches, and it is genuinely valuable. What facilitation cannot reach, because it depends on someone raising the question, is the divergence nobody knows to raise. In the third state, every participant holds, before the first joint decision, a sourced map of where their operative meanings align and diverge, a rehearsal record showing how the candidate framework lands across a grounded population including its margins, and an agreement tested against every party's declared willingness to leave. That last search is partial, and stated as partial.

Only the third state has ever been out of reach. For the first time the components that could reach it can all be built, in the mixture this essay has tried to report exactly: one built for another purpose, others designed and untested, and one specified but not yet built. The program director in the opening scene does not need any of it to be magic. She needs to see what her partners actually mean, in time to decide what to do about it.

On the reading of the pressures I gave above, the window in which coordination instruments can shape AI governance, rather than merely document its failures, is open now. Organizations, professions, and nations do not need to agree about everything to steer this technology well. They need to see what they are actually disagreeing about, and they need agreements built to survive the people who make them. That is a buildable thing.

References

Note on sources. Every reference below is carried over from the verified reference records of the Paramerge white papers and the author's capstone research corpus, where each was verified against a live publisher page. The Nash and Rawls entries are owner-sanctioned additions, and the Ostrom entry was added in this revision; all three were verified against live publisher records in August 2026.

Alt, J. K., Jackson, L. A., Hudak, D., and Lieberman, S. (2009). The Cultural Geography Model: Evaluating the impact of tactical operational outcomes on a civilian population in an irregular warfare environment. The Journal of Defense Modeling and Simulation, 6(4), 185 to 199. https://doi.org/10.1177/1548512909355000

Alt, J. K., Lieberman, S., and Blais, C. (2010). A use-case approach to the validation of social modeling and simulation. In Proceedings of the 2010 Spring Simulation Multiconference (pp. 1 to 7). Society for Computer Simulation International. https://doi.org/10.1145/1878537.1878545

Anwar, U., et al. (2024). Foundational challenges in assuring alignment and safety of large language models. Transactions on Machine Learning Research. https://doi.org/10.48550/arXiv.2404.09932

Bai, H., Voelkel, J. G., Muldowney, S., Eichstaedt, J. C., and Willer, R. (2025). LLM-generated messages can persuade humans on policy issues. Nature Communications, 16, article 6037. https://doi.org/10.1038/s41467-025-61345-5

Bengio, Y., et al. (2024). Managing extreme AI risks amid rapid progress. Science, 384(6698), 842 to 845. https://doi.org/10.1126/science.adn0117

Brincat, S. (2015). Global climate change justice: From Rawls' Law of Peoples to Honneth's conditions of freedom. Environmental Ethics, 37(3), 277 to 305. https://doi.org/10.5840/enviroethics201537329

Costanza-Chock, S. (2020). Design justice: Community-led practices to build the worlds we need. MIT Press.

Gilovich, T., Griffin, D. W., and Kahneman, D. (Eds.). (2002). Heuristics and biases: The psychology of intuitive judgment. Cambridge University Press.

Hirschman, A. O. (1970). Exit, voice, and loyalty: Responses to decline in firms, organizations, and states. Harvard University Press.

Israel, B. A., Eng, E., Schulz, A. J., and Parker, E. A. (Eds.). (2005). Methods in community-based participatory research for health. Jossey-Bass.

Lieberman, S. (2012). Extensible software for whole of society modeling: Framework and preliminary results. SIMULATION, 88(5), 557 to 564. https://doi.org/10.1177/0037549711404918

Lieberman, S. (2023). Latent epistemological intersubjectivity undermines social justice; the science and practice of social good can change that. Unpublished manuscript, University of Southern California, Suzanne Dworak-Peck School of Social Work.

Nash, J. F. (1950). The bargaining problem. Econometrica, 18(2), 155 to 162.

Nash, J. F. (1951). Non-cooperative games. Annals of Mathematics, 54(2), 286 to 295.

Osborne, M. J., and Rubinstein, A. (1990). Bargaining and markets. Academic Press.

Ostrom, E. (2010). Beyond markets and states: Polycentric governance of complex economic systems. American Economic Review, 100(3), 641 to 672.

Park, J. S., Zou, C. Q., Kamphorst, J., Egan, N., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Liang, P., Willer, R., and Bernstein, M. S. (2026). LLM agents grounded in self-reports enable general-purpose simulation of individuals. arXiv:2411.10109, revised 28 June 2026. [Preprint.]

Rawls, J. (1971). A theory of justice. Belknap Press of Harvard University Press.

Rittel, H. W. J., and Webber, M. M. (1973). Dilemmas in a general theory of planning. Policy Sciences, 4(2), 155 to 169.

Star, S. L., and Griesemer, J. R. (1989). Institutional ecology, "translations" and boundary objects: Amateurs and professionals in Berkeley's Museum of Vertebrate Zoology, 1907 to 39. Social Studies of Science, 19(3), 387 to 420. https://doi.org/10.1177/030631289019003001

Sunstein, C. R. (1995). Incompletely theorized agreements. Harvard Law Review, 108(7), 1733 to 1772.

Tessler, M. H., Bakker, M. A., Jarrett, D., Sheahan, H., Chadwick, M. J., Koster, R., Evans, G., Campbell-Gillingham, L., Collins, T., Parkes, D. C., Botvinick, M., and Summerfield, C. (2024). AI can help humans find common ground in democratic deliberation. Science, 386(6719), eadq2852. https://doi.org/10.1126/science.adq2852

About the Author

Stephen Lieberman is the founder of Paramerge, an AI safety and governance practice, and home of the Real-World AI Governance Center, the PAN Lab, and the Oversight simulation. At the Naval Postgraduate School he founded and led the Complex Adaptive Social System (CASS) framework, an agent-based social simulation validated to Department of Defense VV&A standards for its original intended use. Also at NPS he founded and led GlobalECCO, the Combating Terrorism Fellowship Program's international education and collaboration network, serving as Principal Investigator of its founding projects and grounding its serious games and collaboration systems in simulation science coupled with real-world data and in silico testing. His work draws on more than a decade of research applying complex adaptive systems methods to live derivatives markets, and he is a Doctor of Social Work candidate at the University of Southern California. He can be reached at stephen@paramerge.com.

Paramerge builds the coordination and governance instruments this essay describes, and works with teams that want their AI governed where the behavior actually lives.

All Paramerge writing

Contact Paramerge