← All Paramerge white papers

Paramerge White Paper

Getting Everyone on the Same Page

Latent Epistemological Differences and the New Case for AI-Mediated Coordination

Stephen Lieberman

Paramerge, Real-World AI Governance Center

August 2026

stephen@paramerge.com · paramerge.com

Executive summary

Large collaborative efforts fail for a reason that is rarely named. Participants assume they mean the same thing by words like evidence, fairness, risk, and safety, and discover only at a critical moment that they do not. I have called these failures problematic latent epistemological differences, or PLEDs. They are relational, multilevel, and hard to see from the inside, and no instrument has yet surfaced them at the scale of a real coalition, agency, or international body. This paper argues that large language models, deployed inside empirically grounded social simulation and disciplined by provenance tracking and human deliberation, are the first plausible candidate for that instrument, and it lays out what building and testing one would take.

The evidence has arrived quickly and unevenly. A peer-reviewed study with 5,734 UK participants, most recruited by convenience sampling, found that AI-drafted group statements were preferred over statements written by incentivized participant mediators 56 percent of the time, a finding replicated in a separate, demographically representative 200-person virtual citizens' assembly. Language-model agents built from in-depth interviews with 1,052 Americans reproduced held-out survey responses at a large fraction of the participants' own test-retest consistency. Simulation platforms now run tens of thousands to a million interacting agents, a scale claim rather than a fidelity claim. This validation covers constrained tasks, chiefly survey and experiment replication, not open-ended deliberation. And the failure modes are documented, not speculative. Variance collapse, sycophancy, caricature, and degraded subgroup accuracy appear repeatedly across the recent literature, and every one biases such a system toward understating divergence, which is exactly what a PLED instrument must not do.

The paper is explicit about its central premise. That surfacing a latent difference early improves the outcome is a hypothesis, not a theorem. There are documented settings in which ambiguity is load-bearing and surfacing can harm, and the architecture treats the decision to surface as a human decision, with a stated test of the premise in the build order.

The architecture has four layers. An epistemic mapping layer reads each party's documentary record and maps where operative meanings align and diverge. A provenance layer ties every finding to inspectable sources. An in silico testbed, built on the published equations of an agent-based framework that went through Department of Defense verification, validation, and accreditation for its original military intended use, lets partners rehearse a proposed shared framework against a simulated population. That accreditation attaches to the original model and purpose and does not transfer to the research reimplementation used here or to any language-model coupling, and the paper says so plainly. A deliberation layer keeps every decision, and all accountability, with the humans. The mapping layer and the language-model integration are designed and largely untested. This paper is the argument for building and testing them, in a stated order, against stated benchmarks.

1. The problem nobody puts on the agenda

Every large collaborative undertaking begins with the unexamined assumption that the people at the table understand one another. Partners plan and spend on the belief that their conceptions of the shared goal are compatible, and at some critical point that assumption fails. A term that everyone had been using turns out to have carried three incompatible meanings. What one partner counts as evidence another regards as anecdote; what one calls fairness another experiences as harm.

I have called these failures problematic latent epistemological differences, or PLEDs (Lieberman, 2023). They are epistemological because they concern what participants will accept as grounds for a claim, latent because they stay hidden until circumstances force them into the open, and problematic because when they surface late they do not merely slow a program; they can end it, taking relationships and funding with them. The social good paradigm has made cross-sector collaboration the central path to social impact (Mor Barak, 2020). Three features make PLEDs a wicked problem rather than a communication problem.

First, they are relational and multilevel, never a property of one person. They live between individuals, organizations, communities, and nations, and across those levels, where differences at the top get baked into programs at the bottom (Levin, 2020; Brincat, 2015; Lieberman, 2023).

Second, they are invisible from the inside. Naive realism, the tendency to experience one's own perception as simply the way things are, means participants do not notice that a concept could be construed differently until someone else construes it differently (Gilovich, Griffin, and Kahneman, 2002; Lopez-Rodriguez et al., 2022). You cannot ask people to list assumptions they do not know they hold.

Third, they scale badly, and the binding cost is elicitation, a burden growing with parties times concepts; twenty organizations and a few hundred load-bearing concepts mean thousands of cells of skilled interpretive work, so the standard response has been to not try. Even a complete pairwise map does not certify coalition-level shared meaning, because tolerable pairwise differences can chain into an endpoint gap nobody would tolerate.

Brincat (2015) argues that international climate justice founders on crises of intersubjectivity, nations and international bodies holding conceptions of justice that do not connect (Lieberman, 2023). Within a single district, restorative justice programs built to run identically across schools have fractured where individual schools, advised by different consultants, understood the principles differently (Song et al., 2020, as read in Lieberman, 2023). The scale changes. The mechanism does not.

2. Why the problem has been unsolvable

The failure is not for lack of trying; structured traditions have attacked pieces of this problem for decades. Q methodology surfaces distinct viewpoints from how participants sort statements, the policy Delphi surfaces dissensus rather than forcing convergence, group model building, concept mapping, and repertory grids elicit divergent stakeholder constructions of a shared problem, and deliberative polling puts balanced information before representative microcosms of the public (Fishkin, 2011). What has been missing is capacity, not understanding. Four constraints have held.

First, the cost of surfacing has exceeded the cost of failure. The traditions above elicit in facilitated sessions, capping them at workshop scale, and eliciting every stakeholder's operative definition of every load-bearing concept exceeds any program budget. So organizations gamble.

Second, no instrument has read the record. Surveys and sorts capture stated positions, and interviews capture depth for a few. No existing method reads the documents, decisions, and arguments in which operative meanings live, or attaches to each finding a source a disputing party could inspect.

Third, empiricism itself became contested. The postmodern turn brought essential gains, above all the centering of marginalized voices and lived experience, but in its stronger forms it made the meaning of evidence a matter of dispute (Caputo et al., 2015). A coalition cannot resolve disagreements by appeal to evidence when what counts as evidence is one of the disagreements.

Fourth, there was no way to test an agreement before committing. Partners who did the hard work of alignment found out in deployment what happened when the framework met a real population.

Any solution must address all four constraints, and they define the requirements.

3. Requirements for a solution

The 2023 analysis established what PLEDs are (Lieberman, 2023). The requirements below are my derivation from it and from the constraints above, stated as this paper's design commitments so the architecture in Section 5 can be judged against them.

1. Surface differences before the critical point, during formation and design, not only in retrospect.

2. Operate across all levels, individual through national, and between them.

3. Work from the record of what people write and decide, not only what they say when asked, knowing that documents are audience-tailored performances as well as traces of use.

4. Remain value-neutral with respect to individuals and groups. The system shows that two parties use a concept differently; which use is correct is the people's decision.

5. Preserve the gains of participatory and critical approaches; marginalized voices are the population whose epistemologies most need to be visible rather than assumed away.

6. Track provenance. Every claim, inference, and rehearsal outcome must trace to its sources, turning the fight about evidence into a conversation about sources.

7. Permit rehearsal. Partners must be able to stress-test a proposed framework against a realistic simulated population before committing real resources and people.

8. Augment rather than replace human judgment. The agreement, and the accountability for it, remain human.

Requirements 3 and 5 pull against each other, since the least documented parties are the ones requirement 5 most protects, a tension to manage rather than assume away.

3.1 The premise, stated as a hypothesis

Everything that follows rests on a premise that must be named as a hypothesis, because that is what it is. The premise is that surfacing a latent divergence before commitment improves the outcome. Hirshleifer (1971) showed that publicly revealed information can destroy value among strategic parties, and treaty craft calls the protective version constructive ambiguity. Sunstein (1995) documented that plural institutions function through incompletely theorized agreements, agreeing on particulars precisely by not forcing abstract principles into the open. Star and Griesemer (1989) showed that cooperation across communities is often carried by boundary objects, shared terms that coordinate exactly because each side reads them locally. There are collaborations in which the plasticity of a shared term is the coordination mechanism, and printing its incompatible readings, with sources, could prevent an agreement that practice would have completed.

The psychology cuts the same way. The volume that documents naive realism also documents that awareness-based debiasing is difficult and often fails (Gilovich, Griffin, and Kahneman, 2002), and its bias blind spot findings predict that each party will accept the map's account of others and resist its account of themselves. The experimental evidence that awareness of naive realism increases acceptance is real but modest, a small effect on self-reported measures in laboratory samples (Lopez-Rodriguez et al., 2022).

Three design commitments follow. Plural operative use of a term is reported as a first-class finding that may be protective, never as a defect awaiting resolution. The parties, not the system, decide whether a surfaced divergence goes on the agenda, including leaving it alone. And the premise enters the Section 8 build order as an early, falsifiable test, with the asymmetric-acceptance prediction registered as the adversarial hypothesis. The claim this paper defends is therefore conditional; visibility is necessary for informed agreement, not sufficient, and an instrument that can surface must also know when not to.

4. What has changed

4.1 The capability, as hypothesis

Large language models are, among other things, instruments for reading and comparing meaning at scale. A modern model can ingest every party's full documentary record, from policy manuals and grant applications to meeting notes. The hypothesis this paper advances, and Section 8 commits to testing, is that it can identify how each party uses each load-bearing term in practice, faster than any facilitation budget allows, and attach to every inference the passages that led it there. The model is no arbiter. Language models are fluent, not authoritative; they confabulate, inherit their corpora's biases, and will confidently misread a document if allowed to. The claim is narrower. Disciplined by the components below, they may at last make the space of latent difference visible, under the conditions of Section 3.1.

4.2 The evidence, and its boundaries

The strongest evidence comes from AI-mediated deliberation. Tessler et al. (2024), in a peer-reviewed study with 5,734 adult UK residents, most recruited by convenience sampling, found that participants preferred an AI mediator's group statements over those written by human mediators 56 percent of the time versus 44 percent. External judges rated the machine statements higher for quality, clarity, informativeness, and perceived fairness, and the findings were replicated in a virtual citizens' assembly with a separate, demographically representative sample of 200 UK residents. A follow-on essay frames this work against the trilemma among participation, deliberation, and political equality identified by Fishkin (2011) (Tessler et al., 2026). It is the closest prior art, foundation and contrast at once; the Habermas Machine line optimizes toward a single consensus statement without mapping where operative meanings diverge, and a consensus statement built over an unsurfaced PLED inherits the PLED.

The second thread is simulation of people. Park et al. (2026), in a preprint, built language-model agents from interviews and surveys with a diverse national sample of 1,052 Americans. On held-out General Social Survey items, interview-based, survey-based, and combined agents reached 83, 82, and 86 percent of the participants' own two-week test-retest consistency, against 74 percent for demographics-only agents, narrowing though not eliminating accuracy disparities across racial and ideological groups. At larger scales, AgentSociety reports more than ten thousand agents and five million interactions (Piao et al., 2025), and OASIS reports simulations of up to a million agents reproducing information spread, polarization, and herd effects (Yang et al., 2025).

The boundary matters as much as the evidence. Validation to date covers constrained tasks, chiefly survey and experiment replication, and reproducing observed aggregates is a weaker warrant than answering intervention questions, since different models can fit the same aggregates yet diverge under interventions absent from the fitting data. Park et al. themselves note that their evaluation is individual-level and does not establish whether agents reproduce how attitudes covary across a sample; a PLED map is precisely a claim about how meanings covary across parties. Anthis et al. (2025) conclude these simulations can already be used for pilot and exploratory studies, and a peer-reviewed critical review finds that language models may exacerbate rather than alleviate the long-standing validation challenges of agent-based modeling (Larooij and Tornberg, 2026). The honest conclusion is that these systems have earned a role as rehearsal instruments, not as oracles about what a real population will do.

4.3 The documented failure modes

The same literature documents how these systems fail. Four failure modes matter most here.

Variance collapse. Persona-conditioned language-model survey respondents recover human averages while showing less variation than real surveys, most severely on feelings toward racial and religious groups (Bisbee et al., 2024). On standardized creativity tests, Wenger and Kenett (2025) measured population-level response variability of 0.459 for unpersonified models against 0.738 for humans, an effect halved under a response-format control; that measures creative output diversity, not simulation agents, and its authors caution against extrapolation. For a PLED instrument this failure is the most dangerous of all, because it erases minority positions and hides exactly the divergences the tool exists to surface.

Sycophancy and false consensus. Models tuned to be agreeable over-agree, refuse realistic conflict, and can manufacture consensus that does not exist (Anthis et al., 2025; Taillandier et al., 2026).

Caricature, in tension with this paper's own mechanism. Pushed away from the average, simulated agents drift into stereotype, unevenly across groups, and caricature is worst for general, uncontroversial topics (Cheng, Piccardi, and Yang, 2023). Coalition abstractions like safety and fairness sit at the general end. Worse, agents built from 1,186 real users showed that richer contextualization improves internal consistency while amplifying polarization, stylized signals, and toxic language (Nudo et al., 2026). Richer context is this architecture's central mechanism, so context volume must be a tuned variable with caricature audits, not an unalloyed good.

Degraded subgroup accuracy. In the studies Anthis et al. (2025) review, performance was substantially reduced in group-level rather than population-level prediction; the parties most likely to be misrepresented have the least documentary presence, and an aggregate-only audit can pass while the margins fail.

These failure modes share a direction, pushing measured divergence down and rehearsed consensus durability up, toward the same decision error, a coalition committing to a framework that will fracture. The architecture therefore adopts a standing rule. Measured divergence is a lower bound, rehearsed durability an upper bound, and a no-fracture rehearsal is uninformative rather than evidence that the framework holds.

The mitigations are hypotheses too. Taillandier et al. (2026) propose, as a conceptual research direction, hybrids that stratify classical agent-based models with language models, while flagging explanatory and predictive uses as the epistemically risky ones. Anthis et al. (2025) recommend context-rich prompting, fine-tuning on social science data, and predicting as experts rather than role-playing personas. Constraining language-model agents with structure the model does not supply is a family the field has proposed and not yet evaluated; this architecture is one member of it, a weaker and more honest warrant than consensus.

5. An architecture that meets the requirements

The design has four layers, each layer's maturity stated in Section 8. I am aware of no published system that combines them. None of the nearest relatives, disagreement detection, viewpoint-clustering civic tools, or the traditions of Section 2, builds a cross-party map of how the same load-bearing term carries different operative meanings from documentary records, with per-claim provenance, coupled to a population testbed. That survey of nearest neighbors, not a universal negative, is the gap claim.

5.1 The epistemic mapping layer

This is the language-model layer, serving requirements 1 through 5. It reads each party's record, extracts operative meanings for a working set of load-bearing concepts, compares them across parties and levels, and renders a relational map of alignment and divergence. Three constraints matter. The layer reports differences, never rankings, so the community partner's conception of fairness is described rather than graded against the funder's. There is one canonical map, identical for every party; translations into each party's register are an additive layer anyone can inspect, because a map existing only in private versions would manufacture a second-order divergence about the map itself. And the concept list is set participatorily, not by whoever commissions the map, because selecting which concepts appear is where neutrality is won or lost.

Disagreement detection now aligns a model's comprehension with expert sub-concept taxonomies (Liu et al., 2025), alignment between model and expert rather than a map of how the disputing parties use a term. Civic tools such as Talk to the City cluster free text into themes grounded in participants' actual statements (AI Objectives Institute, n.d.). Both are ingredients. Neither produces the thing a coalition needs.

5.2 The provenance and evidence layer

This layer satisfies requirement 6. Every claim the mapping layer produces is tied to a registered source. Provenance does not settle the inference. A partner who disputes a finding is shown the passages, can say what the model got wrong, and can amend the record, the amendment standing alongside the finding rather than beneath it. What provenance buys is that the dispute becomes inspectable and bounded. It also creates a manipulation gradient, because showing a party which passages are load-bearing tells it what to rewrite; Section 7 treats that as a failure mode, not a footnote.

5.3 The in silico testbed

This layer satisfies requirement 7, and it could not be built by a language model alone. Its foundation is CASS, the Complex Adaptive Social System framework developed at the Naval Postgraduate School, an agent-based population simulation that builds artificial societies from survey data (Alt and Lieberman, 2010; Lieberman and Alt, 2010; Lieberman, 2012). Its published intended use was evaluating civilian population response in an irregular warfare environment (Alt, Jackson, Hudak, and Lieberman, 2009), and its validation followed a documented use-case approach within the Department of Defense verification, validation, and accreditation discipline (Alt, Lieberman, and Blais, 2010). Validation and accreditation under that discipline attach to a specific model, intended use, and time; they are not portable. The research implementation now in use reimplements the framework's published, open equations; it has not inherited that accreditation, which would not extend to rehearsing coalition frameworks anyway. It is a research tool, not a product, and its fitness for this use is what Section 8's build order must establish.

Once the mapping layer has surfaced where meanings diverge, the partners can propose a shared framework and rehearse it against a simulated population grounded in the real one. Language-model agents can be seeded from the grounded model to give the rehearsal the texture of deliberation, an integration designed and largely untested. Seeding fixes the starting distribution, not the dynamics, so the coupling must be specified and gated, stating which state variables the language layer may write and under what bounds, with an acceptance test that observed dispersion survives over simulated time and per subgroup, and a null run with the language layer disabled reporting what it changed. Rehearsals run as ensembles across a declared range of population specifications; only ensemble-invariant conclusions are reported, the rest listed as unresolved. Every rehearsal output is a hypothesis, never a forecast; under the failure profile of Section 4.3, a fracture is a lead worth investigating while a hold is not evidence of durability. Until a scored track record exists, a rehearsal may reorder the questions the partners investigate, never decide adoption or rejection.

5.4 The deliberation layer

This layer satisfies requirement 8. The outputs of the first three layers are inputs to human decision, not substitutes for it. In Tessler et al. (2024) the machine drafted and humans decided; in the Community Notes extension, language models draft candidate notes while a diverse human community remains the sole evaluator (Li et al., 2025). The principle is a safety property. Language-model text persuades people on policy issues about as effectively as text written by other people (Bai et al., 2025), so a system that can surface differences can also move them, and a human decision layer is necessary though not sufficient, since what the humans are shown still shapes what they decide. So the common map, the participatory concept list, and the right of amendment are commitments of this layer too, and no party's record is ingested without its consent.

The institutional side draws on Paramerge's Policy Actor Network system, which models how decisions and errors propagate through organizations under pressure, showing how an agreement reached at the top will behave inside the institutions that must carry it out. Whether an endorsed agreement is durable when participants retain the option to walk away is left open here; a companion Paramerge paper takes it up directly.

6. A worked example

Imagine a multi-agency effort to deploy a shared risk-assessment tool across a state's child welfare, housing, and behavioral health systems. Six agencies, two community partners, an evaluation team, and a funder sign a memorandum of understanding built around three words: safety, equity, and evidence.

Under current practice, those words would be assumed compatible until, months into implementation, the community partners discover that the evaluation team's definition of evidence excludes practitioner observations, the housing agency discovers that the child welfare agency's operational definition of safety would flag a large share of its own clients, the program stalls, and the funder pauses disbursement.

Under the architecture described here, the mapping layer would have read all ten parties' records before signature, and returned a map showing that safety carried two incompatible operational definitions and evidence at least three, with each party able to inspect and amend the passages behind those findings. The testbed, once validated for this use, would have let the group rehearse a candidate shared definition of safety against a population grounded in the state's own data, surfacing fracture hypotheses for the partners to check, with subgroup readings gated on the audits Section 8 requires. The deliberation layer would have let the parties negotiate a definition with their eyes open, and a Policy Actor Network model would have shown where in the six agencies that definition was most likely to drift once deployed.

There is another branch. Shown a sourced three-way disagreement about evidence before disbursement, the funder might walk, and the memorandum might never be signed. Whether that is a save or a loss is the parties' judgment, which is why surfacing is an input, not a verdict. What the architecture removes is the failure in which nobody could see the difference until it was too late.

7. Principles and failure modes

Because this architecture puts a powerful and imperfect instrument at the center of human agreement, its principles matter as much as its components.

Augment, never replace. The system surfaces and rehearses; people decide. No output is a recommendation, and no agreement is valid because the system endorsed it.

Value neutrality is enforced, not assumed. Neutral description is an engineering goal model behavior does not deliver by default; homogenization flattens real differences and caricature distorts them (Bisbee et al., 2024; Cheng, Piccardi, and Yang, 2023; Nudo et al., 2026). Output audits are not enough, since neutrality is also decided upstream, in which concepts get mapped, whose documents count, and how parties are individuated; the participatory concept list and the common map keep those choices visible and shared.

Provenance is mandatory, and the corpus is strategic. Any finding that cannot be traced to a source is discarded. But a deployed, announced instrument changes the record it reads. Parties who know their manuals will be mapped will write them for the mapper, load-bearing vagueness is sometimes deliberate, and passage-level provenance tells a sophisticated party exactly which sentences to rewrite. The partial responses are to weight corpus components by how costly they are to restate, prefer records tied to conduct such as budgets and case dispositions, compare disclosed against held-out portions of a record, and treat a corpus that changes sharply after announcement as a finding. A residual gaming risk remains open.

Validation before trust. A simulated population is only as trustworthy as its grounding, and this testbed is not yet validated for this use. Rehearsals are stress-tests, never forecasts, no-fracture results are uninformative, and the gates of Section 5.3 must pass before any rehearsal informs a decision.

The margins are the test. The parties least represented in the record are the most likely to be misread, group-level prediction degrades relative to population-level prediction (Anthis et al., 2025), and caricature concentrates on marginalized personas (Cheng, Piccardi, and Yang, 2023). Participatory design is how the mapping layer is corrected, per-subgroup audit is how the correction is checked, and where a subgroup is too thin to audit, the system must say so rather than report anyway.

The most dangerous failure mode is not that the system will be wrong. It is that it will be fluent, and that fluency will be mistaken for authority. Every layer above exists to prevent that.

8. What exists and what remains

The CASS framework was developed, published, and taken through the Department of Defense verification, validation, and accreditation discipline for its original intended use beginning more than fifteen years ago, and its equations are published and open. A research implementation of those equations for building artificial societies and testing interventions in silico is in active use, with an iterative pipeline for grounding populations in real-world data; it is a research tool, not a product, with no accreditation for the use proposed here. The Policy Actor Network system and the PAN Lab are deployed and public, and the provenance discipline of Section 5.2 governs the Governance Center's documented case files.

The epistemic mapping layer and the language-model integration are designed and largely untested, and Section 4 says their evaluation cannot be waved through. The build order follows.

1. Build the mapping layer first, as a provenance-tracked map of operative-meaning divergence over real documentary records, the highest-novelty, lowest-simulation-risk component.

2. Test the premise. Run a preregistered comparison in which real planning teams receive either a divergence map with provenance or a matched control briefing, with outcomes measured on the agreement artifact and with asymmetric acceptance, each party's endorsement of map claims about others minus claims about themselves, registered as the bias blind spot's prediction. If self-directed acceptance collapses, redesign the layer, for instance by eliciting self-authored definitions first.

3. Validate the map with two gates. A recognition gate, whether the mapped parties accept and can amend the account of their own operative meanings, sits above an accuracy gate, precision and recall on divergence detection, reported per subgroup with a stated prevalence and false-discovery policy, since unmanaged multiplicity at coalition scale would bury the divergences that matter. Build the retrospective corpus early, real collaborations that fractured over a latent conceptual difference, with a matched control arm that did not fracture, so the detector is scored against latent divergence rather than manifest disagreement.

4. Only then couple language-model agents to the grounded testbed, under the gates of Section 5.3, with hard thresholds for variance collapse, sycophancy, and caricature that a rehearsal must pass before informing any decision.

5. Build the missing benchmark. No public benchmark tests whether a tool surfaces latent conceptual divergence before a collaboration fractures. It must score both whether the divergence was surfaceable in advance from the pre-fracture record and whether surfacing was associated with better outcomes, with the sample including collaborations that surfaced a divergence early and never formed.

9. Conclusion

Consider three states of a coalition. In the first, nobody knows a PLED exists until month fourteen. In the second, a heroic facilitator has surfaced the most obvious few in a two-day workshop. In the third, every participant holds, before the first joint decision, a sourced map of where their operative meanings align and diverge with everyone else's, in their own language. Only the third state has ever been out of reach. The components that could reach it exist now, and they become trustworthy only inside grounded simulation, disciplined by provenance, audited at the margins, gated by tests no one has yet run, and subordinate to human deliberation.

Building capable AI has turned out to be easier than agreeing on where it should take us. What may now be changing is that, for the first time, the people who must agree can see what they are actually disagreeing about, in time to decide what to do about it. Whether that visibility improves how we coordinate is a testable hypothesis, and this paper has tried to state it precisely enough to be tested. That is a smaller sentence than the ambitions in this field usually carry, and it is the one the evidence supports.

References

Note on sources. Preprints are marked as such, with the arXiv version number and revision date recorded, and are not peer reviewed. Every entry below was verified against a live page on 27 and 28 August 2026; where a publisher page blocked automated access, the entry was verified against the publisher's Crossref registration record, a public index record, or the published article file, as recorded in the paper's verification record.

AI Objectives Institute. (n.d.). Talk to the City. https://talktothe.city (accessed 27 August 2026). [Deployed open-source tool; organizational attribution via the site's linked repository; the site carries no publication date.]

Alt, J. K., Jackson, L. A., Hudak, D., and Lieberman, S. (2009). The Cultural Geography Model: Evaluating the impact of tactical operational outcomes on a civilian population in an irregular warfare environment. The Journal of Defense Modeling and Simulation, 6(4), 185 to 199. https://doi.org/10.1177/1548512909355000

Alt, J. K., and Lieberman, S. (2010). Developing cognitive models for social simulation from survey data. In S.-K. Chai, J. J. Salerno, and P. L. Mabry (Eds.), Advances in Social Computing (Lecture Notes in Computer Science, Vol. 6007, pp. 323 to 329). Springer. https://doi.org/10.1007/978-3-642-12079-4_40

Alt, J. K., Lieberman, S., and Blais, C. (2010). A use-case approach to the validation of social modeling and simulation. In Proceedings of the 2010 Spring Simulation Multiconference (pp. 1 to 7). Society for Computer Simulation International. https://doi.org/10.1145/1878537.1878545

Anthis, J. R., Liu, R., Richardson, S. M., Kozlowski, A. C., Koch, B., Evans, J., Brynjolfsson, E., and Bernstein, M. (2025). LLM social simulations are a promising research method. Position paper, International Conference on Machine Learning (ICML 2025). arXiv:2504.02234, version 2, revised 5 June 2025; first posted 3 April 2025.

Bai, H., Voelkel, J. G., Muldowney, S., Eichstaedt, J. C., and Willer, R. (2025). LLM-generated messages can persuade humans on policy issues. Nature Communications, 16, article 6037. https://doi.org/10.1038/s41467-025-61345-5

Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., and Larson, J. M. (2024). Synthetic replacements for human survey data? The perils of large language models. Political Analysis, 32(4), 401 to 416. https://doi.org/10.1017/pan.2024.5

Brincat, S. (2015). Global climate change justice: From Rawls' Law of Peoples to Honneth's conditions of freedom. Environmental Ethics, 37(3), 277 to 305. https://doi.org/10.5840/enviroethics201537329

Caputo, R., Epstein, W., Stoesz, D., and Thyer, B. (2015). Postmodernism: A dead end in social work epistemology. Journal of Social Work Education, 51(4), 638 to 647. https://doi.org/10.1080/10437797.2015.1076260

Cheng, M., Piccardi, T., and Yang, D. (2023). CoMPosT: Characterizing and evaluating caricature in LLM simulations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 10853 to 10875). Association for Computational Linguistics.

Fishkin, J. S. (2011). When the people speak: Deliberative democracy and public consultation. Oxford University Press.

Gilovich, T., Griffin, D. W., and Kahneman, D. (Eds.). (2002). Heuristics and biases: The psychology of intuitive judgment. Cambridge University Press.

Hirshleifer, J. (1971). The private and social value of information and the reward to inventive activity. American Economic Review, 61(4), 561 to 574.

Larooij, M., and Tornberg, P. (2026). Validation is the central challenge for generative social simulation: A critical review of LLMs in agent-based modeling. Artificial Intelligence Review, 59(1), article 15. https://doi.org/10.1007/s10462-025-11412-6 [Published online 18 November 2025.]

Levin, L. (2020). Rethinking social justice: A contemporary challenge for social good. Research on Social Work Practice, 30(2), 186 to 195. https://doi.org/10.1177/1049731519854161

Li, H., De, S., Revel, M., Haupt, A., Miller, B., Coleman, K., Baxter, J., Saveski, M., and Bakker, M. A. (2025). Scaling human judgment in Community Notes with LLMs. Journal of Online Trust and Safety, 3(1), article 255. https://doi.org/10.54501/jots.v3i1.255 [Also arXiv:2506.24118, version 1, 30 June 2025.]

Lieberman, S. (2012). Extensible software for whole of society modeling: Framework and preliminary results. SIMULATION, 88(5), 557 to 564. https://doi.org/10.1177/0037549711404918

Lieberman, S. (2023). Latent epistemological intersubjectivity undermines social justice; the science and practice of social good can change that. Unpublished manuscript, University of Southern California, Suzanne Dworak-Peck School of Social Work.

Lieberman, S., and Alt, J. K. (2010). Developing social networks for artificial societies from survey data. In S.-K. Chai, J. J. Salerno, and P. L. Mabry (Eds.), Advances in Social Computing (Lecture Notes in Computer Science, Vol. 6007, pp. 159 to 168). Springer. https://doi.org/10.1007/978-3-642-12079-4_21

Liu, J., Song, J., Pang, Y., Shen, Z., and Rao, Y. (2025). CARE: A disagreement detection framework with concept alignment and reasoning enhancement. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 13264 to 13279). Association for Computational Linguistics.

Lopez-Rodriguez, L., Halperin, E., Vazquez, A., Cuadrado, I., Navas, M., and Gomez, A. (2022). Awareness of the psychological bias of naive realism can increase acceptance of cultural differences. Personality and Social Psychology Bulletin, 48(6), 888 to 900. https://doi.org/10.1177/01461672211027034

Mor Barak, M. E. (2020). The practice and science of social good: Emerging paths to positive social impact. Research on Social Work Practice, 30(2), 139 to 150. https://doi.org/10.1177/1049731517745600

Nudo, J., Pandolfo, M. E., Loru, E., Samory, M., Cinelli, M., and Quattrociocchi, W. (2026). Generative exaggeration in LLM social agents: Consistency, bias, and toxicity. Online Social Networks and Media, 51, article 100344. https://doi.org/10.1016/j.osnem.2025.100344 [Also arXiv:2507.00657, version 1, 1 July 2025.]

Park, J. S., Zou, C. Q., Kamphorst, J., Egan, N., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Liang, P., Willer, R., and Bernstein, M. S. (2026). LLM agents grounded in self-reports enable general-purpose simulation of individuals. arXiv:2411.10109, version 3, revised 28 June 2026; version 2, 22 April 2026; first posted 15 November 2024. [Preprint.]

Piao, J., Yan, Y., Zhang, J., Li, N., Yan, J., Lan, X., Lu, Z., Zheng, Z., Wang, J. Y., Zhou, D., Gao, C., Xu, F., Zhang, F., Rong, K., Su, J., and Li, Y. (2025). AgentSociety: Large-scale simulation of LLM-driven generative agents advances understanding of human behaviors and society. arXiv:2502.08691, version 2, revised 10 April 2026; first posted 12 February 2025. [Preprint.]

Song, S. Y., Eddy, J. M., Thompson, H. M., Adams, B., and Beskow, J. (2020). Restorative consultation in schools: A systematic review and call for restorative justice science to promote anti-racism and social justice. Journal of Educational and Psychological Consultation, 30(4), 462 to 476. https://doi.org/10.1080/10474412.2020.1819298

Star, S. L., and Griesemer, J. R. (1989). Institutional ecology, "translations" and boundary objects: Amateurs and professionals in Berkeley's Museum of Vertebrate Zoology, 1907 to 39. Social Studies of Science, 19(3), 387 to 420. https://doi.org/10.1177/030631289019003001

Sunstein, C. R. (1995). Incompletely theorized agreements. Harvard Law Review, 108(7), 1733 to 1772.

Taillandier, P., Zucker, J. D., Grignard, A., Gaudou, B., Huynh, N. Q., and Drogoul, A. (2026). Integrating LLM in agent-based social simulation: Opportunities and challenges. arXiv:2507.19364, version 2, revised 27 February 2026; first posted 25 July 2025. [Preprint position paper.]

Tessler, M. H., Bakker, M. A., Jarrett, D., Sheahan, H., Chadwick, M. J., Koster, R., Evans, G., Campbell-Gillingham, L., Collins, T., Parkes, D. C., Botvinick, M., and Summerfield, C. (2024). AI can help humans find common ground in democratic deliberation. Science, 386(6719), eadq2852. https://doi.org/10.1126/science.adq2852

Tessler, M. H., Evans, G., Bakker, M. A., Gabriel, I., Bridgers, S., Jain, R., Koster, R., Rieser, V., Dragan, A., Botvinick, M., and Summerfield, C. (2026). Can AI mediation improve democratic deliberation? arXiv:2601.05904, version 1, 9 January 2026. [Preprint; contribution to the Knight First Amendment Institute Symposium on AI and Democratic Freedoms, 10 to 11 April 2025.]

Wenger, E., and Kenett, Y. (2025). We're different, we're the same: Creative homogeneity across LLMs. arXiv:2501.19361, version 1, 31 January 2025. [Preprint.]

Yang, Z., Zhang, Z., Zheng, Z., Jiang, Y., Gan, Z., Wang, Z., Ling, Z., Chen, J., Ma, M., Dong, B., Gupta, P., Hu, S., Yin, Z., Li, G., Jia, X., Wang, L., Ghanem, B., Lu, H., Lu, C., Ouyang, W., Qiao, Y., Torr, P., and Shao, J. (2025). OASIS: Open agent social interaction simulations with one million agents. arXiv:2411.11581, version 5, revised 23 March 2025; first posted 18 November 2024. [Preprint.]

Appendix A. Evidence table

Each row records a load-bearing source, what it shows, and what it does not show. Quantitative figures reported only in preprints are attributed as such and remain subject to revision.

SourceWhat it showsWhat it does not show
Tessler et al. (2024), Science, peer reviewedWith 5,734 UK participants, most convenience-sampled, AI-drafted group statements were preferred over statements by trained, incentivized participant mediators 56 percent of the time versus 44 percent (predicted probability 0.575, interval 0.535 to 0.614); external judges rated them higher for quality, clarity, informativeness, and perceived fairness; replicated in a separate, demographically representative 200-person virtual citizens' assembly.No professional-mediator benchmark; the human-mediator arm had no critique phase; the assembly mimicked the deliberation phase without fact-finding or expert testimony; does not map latent conceptual divergence; optimizes toward one consensus statement; not validation of open-ended LLM agent deliberation.
Tessler et al. (2026), preprint, Knight Institute symposium contributionFrames AI mediation against the participation, deliberation, and political equality trilemma identified by Fishkin (2011), and surveys open challenges, including good-faith participation under real political stakes.Not new large-scale validation.
Park et al. (2026), preprint (arXiv:2411.10109 v3, revised June 2026)Agents grounded in self-reports from 1,052 Americans reached 83, 82, and 86 percent (interview-only, survey-only, combined) of participants' own two-week test-retest consistency on held-out General Social Survey items, versus 74 percent for demographics-only agents; disparities across racial and ideological groups narrowed but were not eliminated.Not peer reviewed; validation is individual-level and constrained-task; the authors state it does not establish whether agents reproduce how attitudes covary across a sample; economic games were a boundary case where no agent type clearly beat the demographic baseline; the five-study experimental replication is underpowered for strong conclusions; figures changed across versions.
Anthis et al. (2025), ICML position paperSystematizes five challenges for LLM social simulation (diversity, bias, sycophancy, alienness, generalization); states these simulations can already be used for pilot and exploratory studies; reports substantially reduced performance in group-level versus population-level prediction in reviewed studies; recommends context-rich prompting, fine-tuning on social science data, and predict-as-expert framing.A position paper, not new empirical validation; does not prescribe agent-based-model grounding or stratified hybrids.
Taillandier et al. (2026), preprintNames the Replicant Effect (collapse of persona diversity), generative exaggeration and toxic caricature, and sycophancy bias as risks; proposes Hybrid Constitutional Architectures, stratifying classical agent-based models with language models, as a conceptual research direction; distinguishes operational value in interactive simulation and serious games from epistemic concerns in explanatory and predictive modeling.Not peer reviewed; a synthesis and design argument, not an empirical measurement; the hybrid direction is proposed, not evaluated.
Larooij and Tornberg (2026), Artificial Intelligence Review, peer reviewedA critical review finding that LLMs may exacerbate rather than alleviate the validation challenges of agent-based modeling, with reported validation practice often resting on face validity or loosely mechanism-tied outcome measures.Does not evaluate this paper's architecture; a review of practice, not a new method.
Bisbee et al. (2024), Political Analysis, peer reviewedPersona-conditioned language-model survey responses recover human averages while showing less variation than real surveys, most severely for feelings toward racial and religious groups, with results sensitive to prompt wording and unstable over months.One model family and one survey tradition; measures attitude simulation, not deliberation.
Wenger and Kenett (2025), preprintPopulation-level response variability of 0.459 for language models versus 0.738 for humans on the Alternative Uses Task (effect size 2.2), shrinking to 0.708 versus 0.850 (effect size 1.1) under a one-word response-format control; models resemble each other across three creativity tests.Not peer reviewed; unpersonified models on creativity tests, not simulation agents; authors caution against extrapolating human psychological tests to models.
Piao et al. (2025), AgentSociety, preprintReports simulation of social lives for more than 10,000 agents and about 5 million interactions, applied to polarization, message spread, universal basic income, and disaster scenarios, with claimed alignment to real-world experimental results.Not peer reviewed; alignment claims are the authors' own; aggregate-pattern reproduction, not individual fidelity or deliberation validity.
Yang et al. (2025), OASIS, preprintSocial-platform simulations up to one million agents reproduce information spread, group polarization, and herd effects; the full text reports agents more inclined to herd behavior than humans; also reports that larger agent groups produced more diverse opinions.Not peer reviewed; itself documents an over-conformity failure mode; not validation for deliberative settings.
Liu et al. (2025), CARE, EMNLP, peer reviewedDisagreement detection improved by extracting a sub-concept taxonomy that aligns the model's comprehension with expert annotation.Aligns model to expert, not party to party; builds no cross-party map of operative meanings; not coupled to any population testbed.
Li et al. (2025), Journal of Online Trust and Safety, peer reviewedA working large-platform design in which language models draft candidate Community Notes while a diverse human community remains the sole evaluator.A design and pipeline description from the platform team, not an independent evaluation; includes a feedback loop in which community ratings train the drafting models.
Bai et al. (2025), Nature Communications, peer reviewedAcross three preregistered experiments (N = 4,829), language-model-generated messages measurably persuaded people on policy issues, about as effectively as messages written by lay humans.Does not concern mapping or simulation fidelity; does not show machine persuasion exceeding human persuasion; its relevance is as a documented risk motivating the human deliberation layer.
Cheng, Piccardi, and Yang (2023), EMNLP, peer reviewedDefines and measures caricature in LLM simulations; caricature is highest for general, uncontroversial topics and decreases with topic specificity; marginalized and political personas are most susceptible.Predates the 2025 to 2026 systems; characterizes the failure, does not remove it.
Nudo et al. (2026), Online Social Networks and Media, peer reviewedAgents built from 1,186 real users and 21 million interactions show that richer contextualization improves internal consistency while amplifying polarization, stylized signals, and toxic language.A social-platform reply setting, not coalition documents; does not test mitigation designs.
AI Objectives Institute, Talk to the City, deployed toolA deployed open-source tool that clusters perspectives into themes and grounds reports in interviewees' actual statements, used with government partners.Not peer reviewed; no publication date on the site; maps viewpoints and claims, not divergent operative meanings of shared terms.
Alt, Jackson, Hudak, and Lieberman (2009); Alt and Lieberman (2010); Lieberman and Alt (2010); Alt, Lieberman, and Blais (2010); Lieberman (2012), peer reviewed and refereed venuesCASS builds artificial societies, including cognitive models and social networks, from survey data; its published intended use was civilian population response in irregular warfare; a use-case validation approach within the DoD verification, validation, and accreditation discipline is documented; the equations are published.The accreditation attaches to the original model and intended use; it does not transfer to the research reimplementation, to language-model coupling, or to coalition rehearsal; none of these papers validates the architecture proposed here.
Lieberman (2023), unpublished manuscriptDefines PLEDs and establishes their relational, multilevel, latent character, and frames value-neutral description via modern data methods as the goal.Unpublished; conceptual analysis, not an empirical study; the requirements and architecture in this paper are derived here, not stated there.
Hirshleifer (1971); Sunstein (1995); Star and Griesemer (1989), peer reviewedThe case against unconditional surfacing. Public information can destroy value among strategic parties; institutions function through incompletely theorized agreements; shared terms coordinate across communities precisely by carrying different local meanings.None concerns AI systems; they bound the premise, not the implementation.
Fishkin (2011), Oxford University PressThe trilemma among participation, deliberation, and political equality, and deliberative polling as a structured response.Not about AI mediation or simulation.
Brincat (2015), peer reviewedA political-theory argument that climate justice founders on disconnected conceptions of justice among nations and international bodies.Interpretive analysis; does not quantify PLED prevalence or causation.
Song et al. (2020), peer reviewedA systematic review of restorative justice consultation in schools, read in Lieberman (2023) as showing school-by-school divergence in understanding a shared program's principles under different consultants.Does not measure latent divergence directly.
Lopez-Rodriguez et al. (2022), peer reviewedExperimental evidence that raising awareness of naive realism increases acceptance of cultural differences, a small effect on self-report measures.Laboratory studies of individuals; not evidence about organizational coordination tools.
Gilovich, Griffin, and Kahneman (2002), peer-reviewed volumeThe foundational treatment of heuristics, biases, and naive realism, including the debiasing and bias blind spot literature that predicts asymmetric acceptance of a divergence map.Not about coordination or simulation.
Caputo et al. (2015), peer reviewedDocuments that the meaning of evidence itself became contested ground in social work epistemology.A position piece within a live dispute; cited for the existence of the dispute, not its resolution.
Levin (2020); Mor Barak (2020), peer reviewedEstablish that social justice language across organizations is often ill-defined with unexplored embedded assumptions, and that the social good paradigm makes cross-sector collaboration central.Conceptual and programmatic papers; not empirical measurements of divergence.

Appendix B. Adversarial review log, condensed

This paper was reviewed in draft by ten independent expert reviewers writing from distinct disciplines: behavioral economics and heuristics; decision making under uncertainty and ambiguity; market microstructure and options pricing; applied mathematics (convexity and aggregation); mechanism design and social choice; complex adaptive systems and simulation validation; large language model failure modes and evaluation; multi-agent systems and simulation validation practice; deliberative democracy and political theory; and a hostile foundation program lead. The reviews were conducted independently and without sight of one another, followed by a supervising adjudication. The full log, with each objection stated at full strength, accompanies the paper's working record. The principal objections and their dispositions are summarized here.

1. Citation and evidence corrections, accepted and made. The variance-collapse source was misattributed to an author not on the paper; it is now cited to Wenger and Kenett (2025) with its construct, elicitation sensitivity, and the authors' own caution stated, and the on-construct persona result of Bisbee et al. (2024) added. The Tessler et al. (2024) description conflated the 5,734 mostly convenience-sampled participants with the separate 200-person representative citizens' assembly, reported a share of comparisons as a share of participants, and omitted the lay-mediator comparator and protocol asymmetry; all corrected. Claims wrongly attributed to Anthis et al. (2025), that it prescribes agent-based grounding and documents aggregate-audit failure, were removed or re-grounded in what that paper states. A caution paraphrase attributed to Park et al. was replaced with the authors' actual stated limitations, including the covariance-structure limitation and the economic-games boundary case. The CARE system's concept alignment was recharacterized as model-to-expert alignment. Li et al. was recognized as published in the Journal of Online Trust and Safety. Every preprint now carries its version and revision date.

2. The field-convergence claim, accepted. "The field's own prescription" overstated two position papers, one of which does not make the attributed recommendation. The mitigation is now stated as one proposed, unevaluated conceptual direction among several, with a peer-reviewed critical review (Larooij and Tornberg, 2026) cited for how underdeveloped validation remains.

3. The visibility premise, accepted. Multiple reviewers showed the paper assumed rather than argued that surfacing divergence early is beneficial, against standing results on the negative value of information, constructive ambiguity, incompletely theorized agreements, and boundary objects. Section 3.1 now states the premise as a hypothesis, cites the countervailing literature, commits to reporting plural use as possibly protective, and adds a preregistered premise test with the bias blind spot's asymmetric-acceptance prediction to the build order.

4. The validation transfer claim, accepted. Accreditation attaches to a model, an intended use, and a time; it does not transfer to a reimplementation, a new domain, or a language-model coupling. Section 5.3 and Section 8 now state the original intended use with its published sources, state that no accreditation carries over, and add a coupling specification, dispersion acceptance test, ensemble reporting, and null comparison as gates.

5. The signed-bias structure, accepted. The documented failure modes all push measured divergence down and rehearsed durability up. The paper now adopts the reserve rule that measured divergence is a lower bound, rehearsed durability an upper bound, and a no-fracture rehearsal is uninformative.

6. The strategic corpus, accepted with residue. An announced instrument makes the documentary record it reads an object of optimization, and provenance shows parties which passages to rewrite. Section 7 now treats this as a failure mode with partial countermeasures and records the residual risk as open.

7. Prior art and the universal negative, accepted. Structured elicitation traditions are now acknowledged, the novelty claim is restated as a survey of nearest neighbors, and Fishkin is credited for the trilemma.

8. Aggregation and neutrality, accepted in part. The paper now commits to one canonical common map, a participatory concept list, a right of amendment, and consent to ingestion, and concedes that output-level audits do not cover upstream selection choices. A formal statement of the map's aggregation rule and weighting axioms remains future work.

9. Untraceable figures and provenance of the requirements, accepted. The unsourced case-file count was removed, and the requirements are now presented as this paper's derivation rather than as conclusions of the 2023 manuscript.

10. Objections standing unresolved, recorded rather than resolved, are listed in full in the review log and are: a complete operating characteristic for the mapping layer (assumed prevalence, decision thresholds, loss asymmetry, false-discovery procedure, and calibration, stated per subgroup, with numeric gate thresholds); a formal specification of the divergence measure and its geometry, including symmetry, rendering distortion budget, and a testable equivariance statement of neutrality; a calibration and prospective scoring regime for rehearsal outputs with a published reliability record; a welfare analysis of coalition non-formation, including the unmeasured base rate of PLED-caused failure and who bears the cost when surfacing prevents an agreement; a mechanism-design treatment of strategic misreporting with tested countermeasures; a docking exercise against an independent model and a prospective sealed-envelope study; treatment of higher-order, more-than-pairwise divergence structure; a round-trip fidelity check for per-party register translations of the common map; formal aggregation axioms for the map, including the documentary-volume weighting rule and a published agenda procedure; engagement with the published normative critiques of the Habermas Machine line; a publicly citable accreditation record for the original CASS framework (authority, date, intended-use statement, acceptability criteria); and a measured cost and capacity model for the mapping layer at the claimed scales.

About the author

Stephen Lieberman is the founder of Paramerge, an AI safety and governance practice for high-stakes human systems, and home of the Real-World AI Governance Center, the PAN Lab, and the Oversight simulation. At the Naval Postgraduate School he founded and led the CASS framework, the Complex Adaptive Social System agent-based simulation program, which was validated to Department of Defense verification, validation, and accreditation standards for its original intended use. At Northrop Grumman he led federal program work including the Global Force Management Data Initiative. He has more than a decade of live derivatives research applying complex-systems methods to equity index futures, has founded and led organizations in the nonprofit sector, and is a Doctor of Social Work candidate at the University of Southern California. He can be reached at stephen@paramerge.com.

The companion paper, Pricing Consensus, asks what a group is actually prepared to do once its members are free to leave.

All Paramerge white papers

Contact Paramerge