Society Is the Deployment
Why There Is No Test Environment for AI at the Scale of Everyone
Stephen Lieberman
Paramerge, Real-World AI Governance Center
August 2026
stephen@paramerge.com · paramerge.com
The study that already ran
In 2018, a study in Science reported a longitudinal look at how true and false news moves through one online network: eleven years of Twitter, about 126,000 fact-checked rumor cascades, roughly three million people passing them along (Vosoughi, Roy, and Aral, 2018). On every measure of cascade diffusion the authors took, false news outran the truth. It traveled farther, faster, deeper, and more broadly. The top one percent of false cascades spread to between a thousand and a hundred thousand people, while the truth rarely spread to more than a thousand. The advantage survived removing the bots their detection tools could find, and adding the bot traffic back accelerated true and false alike. False stories were measurably more novel than the tweets their spreaders had recently seen, and novel information was more likely to be passed on, an association the authors were careful not to call a cause. What remained was people.
That is what one network did with falsehood produced at human speed, one lie at a time. Machines that produce fluent, confident text at industrial scale have now been wired into that network and thousands like it, into the search results, the feeds, the inboxes, the drafting tools, the classroom, the clinic. Whether machine-authored content inherits falsehood's advantage is genuinely open, and I want to say so before building anything on it. Language models are built to produce the probable next word, which cuts toward typicality rather than novelty; the systems around them are tuned for engagement, which cuts the other way; moderation and provenance tooling push back; and nobody has run the eleven-year study on the machine-authored network. What the study establishes is not a prediction about AI. It is a property of the deployment environment: a network of human attention that amplifies some content classes over others for reasons that have nothing to do with truth, mapped at scale, before the new machines arrived in it.
When an engineering team ships software, it ships to a test environment first, what engineers call staging: a copy of the real system where failures are cheap, where the blast radius is contained, where the thing can break and nobody's life changes. The fact that organizes everything else about AI at its current scale is that there is no staging environment. There is no copy of society to try this on. The feeds are live, the institutions are live, the people are real, and every capability and every failure mode is being tried for the first time in the only environment there is. Society is the deployment.
This essay closes the set of pieces published here this month. Two of its companions work one level down, inside the organization and its governing philosophy; the other two already reach for this altitude, one naming coordination among organizations, professions, and nations as substantially deciding the outcome, one stating what we stand to gain, conditional on governance earning it. This one is about the deployment environment itself: what kind of system a society is, what the record says about intervening in systems like it, and what steering could mean where there is no test copy of the world.
The physics of the deployment environment
To reason about a deployment environment, one has to know its physics. A society is a complex adaptive system, and the properties that matter here have been mapped by decades of work across network science, epidemiology, and social research. Five do most of the work in what follows.
A society is nested. A person lives inside a family, inside a neighborhood, inside institutions, inside an economy, and the levels interpenetrate rather than stack. Effects cross them along network paths whose small-world shape lets simple contagions, a story or a virus, reach system scale with almost no warning (Watts and Strogatz, 1998). A father's benefit sanction becomes a child's missed meals becomes a school's problem becomes, years later, a caseload; the illustration is mine, but the pathway it draws is the one the nested structure keeps open. The 2018 study is that structure carrying content on one platform: a false story, injected at one node, spreading in the rare extreme case to a hundred thousand people through nothing but the connectivity of human attention. The authors were explicit that network structure and the characteristics of spreaders favored the truth, and that what drove the differential was people choosing to pass falsity on.
A society runs on feedback. Reinforcing loops drive its booms, panics, and cascades; balancing loops give it the stubbornness that absorbs and defeats well-intended interventions. Which loop catches a policy decides its fate, and the loops are mostly invisible from above.
A society amplifies advantage. Money, connections, and attention accrue disproportionately to those who already hold them, cumulative advantage compounding small initial differences into structural disparities. Equity initiatives built on linear assumptions can be outrun by the system's own arithmetic.
A society does not seek equilibrium. It exists in flux, far from any steady state, and responses to intervention are often convex rather than proportional, impacts accelerating where a linear reading expects them to accumulate (Taleb and West, 2023). Attempts to stabilize a system like that can be futile or counterproductive, and enforcing an equilibrium can eventually produce the collapse it was meant to prevent (Taleb, 2012). Pressure held down is not pressure removed.
And a society tips. Gradual pressure can produce sudden reorganization when thresholds are crossed (Gladwell, 2000). Epidemiology gives the sharpest picture: below a percolation threshold, outbreaks stay small and die out; above it, a giant connected component forms and a large outbreak becomes likely. Some of these reorganizations reverse when the pressure does. The consequential ones often do not, and the ones that matter most are typically visible only in hindsight.
None of this is exotic theory. It is the documented behavior of the systems we all live inside, and every clause of it bears on what it means to weave a new class of information-producing system through that fabric. We are weaving one now.
What the money taught us
Before turning to AI, it is worth looking hard at the cleanest evidence we have about how a simple intervention meets a complex system, because the lesson transfers whole. It comes from the unconditional cash transfer pilots of the last decade, programs whose theory could not be simpler: give people money, no strings.
The interesting results were never the averages. In Stockton, guaranteed income reduced income volatility and improved well-being, and it also shifted employment risk-taking and personal agency in ways the program's designers had hypothesized only loosely (West, Castro Baker, Samra, and Coltrera, 2021). In Chicago, participants in the process evaluation described the money absorbing groceries, toiletries, a bill or two, and the gap from one month to the next, and one of them called it a sense of freedom in deciding where it went. The impact evaluation has not yet reported (Stapleton et al., 2024). The baseline survey of applicants to that same pilot, taken before any money moved, found a third of them scoring in the range for severe psychological distress, against six percent of the comparable national population below 250 percent of the federal poverty line (Brady et al., 2023). In Kenya, cash infusions produced general-equilibrium effects on whole local economies, and the researchers deserve exact credit here: they anticipated those spillovers and built the saturation-randomized design that could measure them, which is why we know the multiplier at all (Egger et al., 2022). Meanwhile equity outcomes turned on delivery details whose distributional consequences no program's theory had sized: how identity verification interacted with housing instability, how the mechanics of getting money to people interacted with financial exclusion (Stapleton et al., 2024). The FDIC's national survey finds 4.2 percent of United States households unbanked and 14.2 percent underbanked, with sharply higher rates among Black, Hispanic, and American Indian or Alaska Native households (FDIC, 2024). Those percentages are a person at an application desk with no bank account to receive the money, or no address a verification system will accept, and no program's theory of change said anything about either.
Read those results as one finding and it says this. Even the simplest intervention, studied by careful teams with randomized designs and independent evaluation, produced its most consequential effects through nesting, feedback, and initial conditions, and the effects that were caught were caught by evaluations built to look for them. What none of those teams had was a way to rehearse the interactions before the money moved. The surprises were discovered in the world, on the participants, because the world was the only place they could be discovered.
Now hold that finding next to the intervention actually underway. AI is not a bounded pilot designed by careful people with an evaluation plan. It is a general capability, introduced everywhere at once, by thousands of uncoordinated actors, with evaluation designs aimed almost entirely at models and their outputs rather than at the level where the effects land. The cash pilots' surprises were bounded by a budget and a city. This deployment is bounded by nothing comparable.
The new agents in the network
What exactly has been added? A new class of actor: systems that produce and route information, participate in decisions, and shape attention, without being stakeholders in any human sense. Societies have always harbored actors without moral agency; the corporation and the bureaucracy are the canonical cases, and the companion essay on high-stakes organizations opens with an automated system that ran roughly forty thousand fraud adjudications with no person inside the loop. What is new is narrower and worth stating exactly. These actors are not composed of persons. They can be instantiated at population scale at almost no marginal cost. They generate content and judgments rather than executing fixed rules. And the accountability structures that reach a corporation through its officers reach these systems mainly through the humans who deploy them, who are usually the people with the least say in their design, and only unevenly through the people who build them.
The failure I documented up close, in the essay that started this body of work, "What is Coulthart?", is fluent output whose correctness is not visible on its face, produced by systems that cascade, one fabrication becoming the premise for the next (Lieberman, 2026). That essay read the machine's inheritance of falsehood's advantage as settled. This one does not, and the correction is mine to make. At the scale of one practitioner and one afternoon, that is a professional hazard. In a network with a documented preference for what is novel and engaging, the same property has to be reasoned about as a contagion question, with the honest caveats already given: whether machine-authored content spreads like human falsehood is unmeasured, and the counter-mechanisms are real. What is not in doubt is the change in supply. False and fabricated content used to cost human effort; it is now effectively unbounded relative to the attention that consumes it.
The subtler entry is not the false story. It is the everyday integration: the drafting tool inside the agency, the summary inside the search result, the recommendation inside the caseload system. Each one small, each one an overloaded person's rational adaptation, the pattern the companion essays document from the inside. Norms are how it compounds. A practice spreads through professional networks by reinforcement, colleagues seeing colleagues use it, organizations reading other organizations' silence as permission, and the threshold picture from the physics section applies here as an analogy I am choosing, not a fitted model: scattered ungoverned use stays reversible; densely normalized ungoverned use becomes self-sustaining, held in place by expectation rather than decision. The companion essay on high-stakes organizations is the record of what governance that arrives after such normalization costs. Nothing about that dynamic stops at an organization's walls.
The window
Here I am reporting my reading of the pressures, not a measurement, and the discipline this set holds to applies: complex systems announce their thresholds mostly in hindsight. What I can name is direction. Three of the pressures I named in the companion essay on coordination have not changed: adoption normalizing toward the point where governance feels optional, crisis-driven regulation arriving faster than deliberate design can inform it, and the governance conversation framed, by default, by the builders of the systems being governed. The fourth belongs to this altitude, and the companions do not reach it. The new actors are being wired deeper into the loops that decide what a population sees: which stories reach it, which summaries stand in for documents nobody opens, which recommendation arrives already inside the tool a caseworker has open. Each of those wirings is a small decision, made by a vendor, a procurement officer, or a default nobody chose. That is a threshold being approached one integration at a time.
If that reading is right, the system does not drift; it reorganizes, and not in a fated direction. The same threshold dynamics that would make ungoverned normalization self-sustaining would do the same for governed practice, if governed practice spreads first: demonstrations moving through professional networks, standards recognized across institutions, norms of disclosure and verification becoming what competence looks like, and what organizations are expected to provide. Complex systems are not reliably steered by direct control alone. The recognition the systems-safety and adaptive-governance traditions have built on for decades is that the properties that matter are emergent over the whole rather than resident in the parts. Much of the remaining leverage, and this reading is my own, is in shaping which of a system's own self-reinforcing patterns wins. Ordinary direct instruments still do real work: in the companion essay on high-stakes organizations, the thing that actually stopped one of the harms was a compliance requirement, applied. The window for that shaping is the period before a threshold is crossed, and nobody will be able to certify the date.
Two errors are possible, and honesty requires pricing both. Acting as if the window is closing has real costs: premature standards, entrenched incumbents, and the companion essay on high-stakes organizations is a record of interventions that were themselves the harm. Both errors are lived before revision reaches them. The claimants in that essay's opening case are the proof, because the determinations against them were reversed in their tens of thousands and the garnished paychecks had already happened. So the difference is not that one error is cheap. A wrong standard is a bounded object with an author, a docket, and a party to petition. A crossed normalization threshold has none of those: nothing to appeal, nobody to name, no docket to enter. That narrower asymmetry, contestability rather than cheapness, is the whole of my case for urgency, and it is an argument, not a theorem.
Steering at the scale of everyone
So what could steering mean, where there is no test copy of the world? The companion essays lay out the program, and none of it is deployed for this purpose yet. The theory of reality stays constant at every scale: govern mechanisms and interactions, not components on a calendar. The coordination machinery is designed to scale because the problem's structure repeats, though whether it survives the level change is exactly what has to be established: at the scale of nations, the divergences in what parties mean by safety, evidence, and harm are as deep as any. And rehearsal matters most here, and claims the least. Artificial societies grounded in real data are the line of work I have pursued since building agent-based social simulations whose original military use was accredited under the Department of Defense's verification, validation, and accreditation discipline, an accreditation that attaches to that model and purpose, not to the research reimplementation now in use, and that does not transfer to this one. What such systems are being built to offer is a place for a policy or a governance framework to fail cheaply before it meets the only real society we have: rehearsal outputs held as hypotheses, never forecasts, checked against the world, corrected by it, continuously. Rehearsal will not surface the variable nobody thought to model. And the documented failure modes of simulated populations all push one way. Variance collapse erases minority positions, sycophancy manufactures agreement, and caricature distorts exactly the groups a program most needs to see accurately, which makes a modeled society look more agreeable than the real one, and thinnest where a subgroup is smallest. That is where the harm would land hardest.
What rehearsal can do, if those failures are measured rather than assumed away, is widen the set of futures a decision is forced to take seriously, and catch the interactions that live among the variables it does hold. That is the kind of surprise the cash pilots met in the world instead.
And the steering has to include the people the system's tails land on. The pilots' hardest lessons lived with unbanked families and housing-unstable applicants, with people like the father from the physics section, whose sanction is still cascading years later, and those people held knowledge about thresholds and feedback loops that no average could surface in time. A society-scale program that models mechanisms but excludes those voices will be wrong precisely where the stakes concentrate, and it will not know it. Participation is the requirement the practice building these instruments has furthest to go on. At this scale, participation is the method. That is a statement about what my own program does not yet meet.
The deployment is us
Here is the shape of the whole argument, small enough to carry. A society is a complex adaptive system: nested, looped, amplifying, far from equilibrium, capable of sudden reorganization. Into that system, at every level at once, a new class of actor is arriving through an adoption dynamic that behaves like a contagion of norms. The best evidence we have about intervening in systems like this says the decisive effects emerge from interactions nobody designed, land first on the people with the least say, and get discovered in the world, on the people, unless something rehearses them first. There is no staging environment. There is only the live system, and everyone is inside it, you included: the feed you read this morning, the tools your colleagues quietly use, the integration that arrived without anyone's decision.
The ask, at the end of a set of essays that has ranged from a fabricated name to the scale of everyone, is smaller than the scale suggests. Not control of the system; the physics offers no reliable version of it, and the record punishes those who claim it. Not a pause; the deployment is running. The ask is that the steering already happening by default, by vendors' defaults, by exhausted adaptations, by whatever crosses its threshold first, become steering on purpose: coordinated, rehearsed, participatory, corrected by the world it acts on. At its smallest, and this part is available on Monday, it is that any organization putting these tools into work that lands on people states what it is steering toward and measures whether it got there. A society does not un-cross a consequential threshold on request, and the harms of the interval are never refunded. What it can still do is choose what shapes it.
References
Note on sources. References are carried from the verified reference records of the author's capstone research corpus, where each was verified against a live publisher page. The Vosoughi, Roy, and Aral entry is a sanctioned addition verified against the publisher's record in August 2026.
Brady, F., Croes, M., Robinson, S., Schexnider, M., Stapleton, S., and Wallace, N. (2023). A first look: Chicago Resilient Communities Pilot. University of Chicago Inclusive Economy Lab.
Egger, D., Haushofer, J., Miguel, E., Niehaus, P., and Walker, M. (2022). General equilibrium effects of cash transfers: Experimental evidence from Kenya. Econometrica, 90(6), 2603 to 2643. https://doi.org/10.3982/ECTA17945
FDIC. (2024). 2023 FDIC national survey of unbanked and underbanked households. Federal Deposit Insurance Corporation.
Gladwell, M. (2000). The tipping point: How little things can make a big difference. Little, Brown.
Lieberman, S. (2026). What is Coulthart? Paramerge. https://paramerge.com/writing/what-is-coulthart
Stapleton, S., Abdul-Razzak, N., Brady, F., Croes, M., Leader-Smith, A., Schexnider, M., and Wallace, N. (2024). Big shoulders: Implementing the Chicago Resilient Communities Pilot. University of Chicago Inclusive Economy Lab.
Taleb, N. N. (2012). Antifragile: Things that gain from disorder. Random House.
Taleb, N. N., and West, J. (2023). Working with convex responses: Antifragility from finance to oncology. Entropy, 25(2), 343. https://doi.org/10.3390/e25020343
Vosoughi, S., Roy, D., and Aral, S. (2018). The spread of true and false news online. Science, 359(6380), 1146 to 1151. https://doi.org/10.1126/science.aap9559
Watts, D. J., and Strogatz, S. H. (1998). Collective dynamics of "small-world" networks. Nature, 393(6684), 440 to 442. https://doi.org/10.1038/30918
West, S., Castro Baker, A., Samra, S., and Coltrera, E. (2021). Preliminary analysis: SEED's first year. Stockton Economic Empowerment Demonstration. https://www.stocktondemonstration.org/
Paramerge builds the coordination and governance instruments this essay describes, and works with teams that want their AI governed where the behavior actually lives.
Contact Paramerge