The Ontological and Epistemic Foundations of Complexity Governance
Why Governing AI Requires a Different Theory of Reality, Not a Better Checklist
Stephen Lieberman
Paramerge, Real-World AI Governance Center
August 2026
stephen@paramerge.com · paramerge.com
Picture a compliance director at a hospital system, eight months after she closed out the annual AI risk review. She did everything right. The NIST AI Risk Management Framework was worked function by function, the ISO/IEC 42001 improvement cycle ran on schedule, the vendor's model documentation was current and attached, and the impact assessment named the coupling that would matter, an ambient documentation tool writing into the system that schedules clinic load. These instruments are not naive. They are iterative by design, built by serious people who know the technology changes, and she worked them in good faith. Between two quarterly reviews the coupling changed under her, because the tool began drafting faster than clinicians could correct it and the scheduler started reading uncorrected drafts as settled, and patients were booked against notes no human had confirmed. Every component involved passes the next review too.
The process did not fail. The process worked exactly as designed. What failed is something underneath the process, something nobody in her building was hired to examine. The instruments are not blind to context: the framework's MAP function exists to establish the setting a system lands in, and the management standard's Annex A requires assessment of an AI system's impact on individuals, on groups, and on society. What neither carries is a method for modeling an interaction. They point at context and then leave the unit of analysis, the system boundary, and the review cadence to organizational discretion, and where conformance is assessed at all, it is assessed by asking whether a process ran, not whether the organization's picture of its own system was adequate. So the reviewed object settles into what is reviewable on a schedule, which is the components, and the coupling that moved was documented once and governed never. The failure lived between the components, and it moved between the reviews. A list of parts held without their couplings is a commitment to the idea that the whole is the sum of its parts. That idea is a theory of reality, and for the systems now being governed it is the wrong one. Every governance instrument is a philosophical commitment wearing work clothes.
This essay is about that commitment: where it came from, why it worked as long as it did, why its effects fail against complex systems in general and AI in particular, and what has to be added to it. The ontology you hold determines what you can see. What you can see determines what you can govern. What you cannot see governs you anyway.
* * *
Start with the dominant orientation, the one so ordinary it does not look like a position at all. I call it Incremental Positivism. Incremental, because its method is the small adjusting step. Positivism, because its warrant is the observed datum. It is the working philosophy of the dominant governance instruments of the modern administrative world, the risk register, the retrospective evaluation, the component review, and almost nobody who uses those instruments would claim it, which is the point. The ontology is not stated. It is presupposed, and it has to be, for the instruments to make sense. The picture is of a closed system, a world where problems decompose into isolated variables and the relationships among them hold still long enough to measure (Bhaskar, 1975, on the open and closed system distinction; Byrne and Callaghan, 2014, on the decomposition critique). From that picture, a picture of knowledge follows. Causal claims are admitted only as observed regularities among measured variables. Knowledge is administrative data, separable metrics, the numbers a system produces about itself. Poverty rates, benefit take-up, recidivism. Test scores, incident counts. Watch the metrics, adjust the levers, watch the metrics again.
The method's best defender deserves to be met at full strength, because he was not a positivist and he was not wrong about his own problem. Charles Lindblom (1959) argued that comprehensive rational analysis of policy is impossible, that values conflict irreducibly, and that the honest response is successive limited comparison, the muddling through he named. That is a pragmatist's answer to the limits of knowledge, and in his hands it carried a genuine reply to everything in this essay: serial adjustment never commits to a model of the whole, it re-anchors on the present each round, so it is arguably the right method for a moving world. The reply fails on speed, not on principle. Incremental comparison corrects after the system has answered, so it steers safely only when the system changes slowly relative to the adjustment cycle. In a nonlinear, interconnected system, the consequential moves are the cascades, and a cascade finishes before the next comparison begins. What the administrative world took from Lindblom was in any case narrower than what he offered. It kept the small steps and married them to a measurement culture he did not endorse, and that marriage is the paradigm this essay is naming. The incremental half carries no ontology at all. The closed-system picture enters with the other half, the moment the thing measured is a separately inspected component.
The commitment to observation also carries a hidden temporal one. If knowledge is observed outcomes, then knowledge is the past. Incremental Positivism looks backward by construction. Its primary instrument is retroactive evaluation, its highest confidence attaches to what has already happened, and its picture of the future is the past extended in a straight line. Human beings then do what we reliably do with the past, which is fit a story to it that makes it feel as if it could have been foreseen (Taleb, 2007). Institutions harden the habit into machinery, because evaluation practice rewards the account of last year's failure that is most defensible, not the model of next year's that is most likely true. The evaluation arrives, the story is written, the story is believed, and governance settles into responding to yesterday's problems with yesterday's causal model.
The paradigm's successes are real, and fairness requires naming them. Where outcomes genuinely are separable and quantifiable, incremental observation delivers, and the measured poverty reductions of the United States safety net expansions are the honest example (Greenstein, 2025). It is worth saying what that example rests on. The reductions are estimates under the anchored Supplemental Poverty Measure, and the baseline they are measured against, what households would have had without the programs, is a constructed counterfactual rather than an observation, so the paradigm's cleanest win already carries a model of a counterfactual world inside it. The trouble announces itself at wicked problems, the class Rittel and Webber (1973) gave its canonical formulation, the ones that refuse to stay decomposed.
Take homelessness, and take it slowly, because it is the clearest way to draw the pattern of a paradigm failing by working. A man loses a job, then an apartment. The emergency shelter gives him a bed, the housing office gives him a voucher, and the voucher expires unused, because the units that accept it sit two bus lines from the jobs he could get, and the stigma of the shelter address filters his applications before a human reads them. Every agency that touched him can document a completed service. Now govern that with Incremental Positivism. Decompose it: a housing variable, a labor variable, a health variable. Fund the voucher, count the placements, evaluate retrospectively, adjust. Each component improves and the crisis continues, because the crisis was never in any component. It lives in the interactions, between housing supply and labor geography and health infrastructure and stigma, and the governing instruments in actual use, the funded program, the counted metric, the after-the-fact evaluation, have no line of sight to an interaction. Social science has interaction methods, and the governance machinery rarely reaches for them, almost never at the point where the funding is decided. It funds components, measures components, and evaluates components, and so the man cycles, and every file on him closes correctly. That is the map of the whole argument, drawn small.
Now hold that failure pattern up next to how the world currently evaluates AI. A model is tested component by component and behavior by behavior, including adversarially and before release, against elicitations someone has thought to try. That is real work, and it is prospective. What it is not is a review of the system the model will be part of: the scores are administrative data, and the review is scoped to the artifact and the session. The deployment assumption is that the system evaluated is the system that ships, and the frozen model file makes that assumption look safe. The assumption does not hold. What is being governed is not the artifact but the sociotechnical system the artifact lands in, the model coupled to its tools, its users, its accumulating context, and its institutional incentives, and that coupled system is neither decomposable nor stationary in the sense the instruments presuppose. I have shown elsewhere, in the essay "What is Coulthart?", what this looks like from close up: twelve documented failures, elicited across three frontier systems in a single hour of ordinary conversational use, not surfaced by any published component-level evaluation of those systems (Lieberman, 2026). Those were deliberate elicitations on a method built for the purpose, reproducible on that method and not yet independently replicated, and they are not a measured incidence rate. I am careful to claim only what they are, which is an existence proof that the failure class is real and reachable. That the same failure appears across three systems built by different teams on different architectures is what makes it structural, and that is an inference from the recurrence rather than something the existence proof delivers on its own. None of this is an argument for discarding the instruments. Component review and lifecycle documentation are necessary, in the way flawless bricks are necessary. They are not sufficient, because flawless bricks do not guarantee a stable building, and you cannot patch an ontology by adding rows to a checklist.
* * *
The alternative has a name, and it did not arrive with AI. Complex Realism, the term is David Byrne's, building on Reed and Harvey's earlier joining of critical realism to complexity science, begins from a different picture of what is real (Reed and Harvey, 1992; Byrne and Callaghan, 2014). Its root is Bhaskar's stratified ontology, which distinguishes three nested domains: the empirical, what is experienced; the actual, what happens whether or not anyone observes it, which contains the empirical; and the real, which contains both and adds the underlying structures and mechanisms whose powers produce what happens (Bhaskar, 1975). In social systems, where critical realism has been argued to be the meta-framework within which complexity is best understood (Gerrits and Verweij, 2013), those mechanisms show up as feedback loops, network structures, and path-dependent thresholds, often opaque, deeply interdependent, and they, not the surface events, are what a governing institution actually needs to reach. Social reality, on this reading, is not a static machine. It is a dynamic space of possible outcomes, and no snapshot of past outcomes reveals its shape.
Change the ontology and the epistemology has to change with it. All observation is theory-laden, in a laboratory as much as in a legislature; that much is old philosophy of science and no special property of complex systems. What complexity changes is the price of the default. Unframed observation does not arrive theory-free. It arrives carrying the linear, additive, decomposable theory as an unstated default, and in a complex system that default theory is not approximately right. It is wrong in the places where the stakes live. So interpretation must run through frameworks that explicitly account for emergence, path dependence, and interdependence, stated out loud. This is not a license for relativism, and critical realism is the tradition that explains why. Our theories are fallible and social. The mechanisms they are about are not, which is why frameworks can be tested against the world and some can be shown better than others. The tradition reached the human services well before AI did, and social work has its own critical realist literature making the case for practice (Houston and Swords, 2021). The frame is older than the technology it is now being asked to govern.
Then comes the payoff question, and here honesty costs something, so I will pay it. The classical critical realist position is that open systems permit explanation but not prediction, and every model of a complex system is a reduction that leaves something out. Both points are right, and the foresight this paradigm offers must be claimed inside them. Mechanism-based models do not forecast. What they do, and what component metrics cannot do even in principle, is expose the space of possibilities the system's own structure makes reachable: which cascades are wired to run, where burdens would concentrate if they ran, which configurations sit near thresholds. Outputs of that kind are hypotheses and rehearsals, disciplined imagination constrained by evidence, and they earn trust only by iteration against what the real system then does. That loop is post-hoc, and it inherits the objection this essay just made to Lindblom, so the distinction has to be drawn exactly. Credentialing a model runs at the pace of evidence and always will. What has to run ahead of the cascade is a different thing: the reachable failure paths laid out before the system is deployed, so that the first case is not also the first time anyone considered it. That is a humbler claim than prophecy and an enormously stronger one than hindsight, and it is the difference between explaining the last failure with precision and being positioned to see the next one while there is still time to act. This orientation is not mine alone, and it is not new. Systems-theoretic safety engineering has argued for decades that safety is an emergent property of a system rather than a property of its parts (Leveson, 2011), and the polycentric governance tradition has built on the same recognition (Ostrom, 2010). The argument here joins that family. What this essay adds is the philosophical floor under it, and a practice built on top.
One more consequence follows. In a complex human system, the mechanisms run through what people mean, intend, and do, so the people inside them hold knowledge of them that outside observation cannot reliably reach at the speed and resolution governance needs: where the feedback loops actually close, what the thresholds feel like from below, which rational policy fails in practice and why. External observers can model mechanisms, and epidemiologists and ecologists prove it every year. But social mechanisms are concept-dependent in a way rivers are not, so a model of a social mechanism that does not take the participants' own concepts as data is missing part of its object, and it will be wrong precisely where the man from the homelessness paragraph could have corrected it. That is what the ontology licenses, and it is narrower than it first looks. Concept-dependence makes people's understandings part of what has to be modeled. It does not make them infallible reporters of the structure they are inside, and critical realism, a tradition built partly to show that an agent's account of their own situation can be false, is the last one that would say otherwise. So the methodological requirement is to take participants' concepts as data. Participation in governance itself rests on the normative case, that people have standing in decisions that land on them, which was already sufficient. The system cannot be adequately known without the people who are part of it.
* * *
Everything Paramerge builds is built on this foundation, and the correspondence is worth stating plainly, with the same discipline the essay has demanded of everyone else. None of it is deployed for this purpose yet. A model built on Complex Realism is still a model. The PAN system decomposes an organization too, into people, machines, pressures, and the links among them. The difference is not an escape from simplification, which nothing gets. The difference is the unit of analysis: structure and interaction rather than isolated components, chosen because that is where the failures documented above actually live. Whether a given model retains the interactions that matter is a contingent, checkable question, and it can be wrong, which is why the whole practice is arranged around checking. The standard cuts both ways, and it should. A well-scoped implementation of the risk framework may retain the interactions that matter, and a badly scoped PAN model may lose them. What the instruments do not do is compel the retention, and that is the charge this essay makes against them. The PAN Lab is a research instrument under active development, a catalogue of modeled organizations, each grounded in the documented record of real systems working and failing, built so mechanisms can be stressed and understood before and during deployment rather than autopsied after. The evidence registry carries hundreds of claims, each tied to a published source or marked plainly as a model output. The Oversight simulation, in development, exists to make the dynamics graspable to the people who must govern them, because a mechanism you cannot see is a mechanism you cannot deliberate about. And the white papers published alongside this essay, "Getting Everyone on the Same Page" and "Pricing Consensus," carry the participatory requirement into method, each stating its own maturity plainly. The correspondence is real, and it is incomplete, which is what makes it a program rather than a slogan. The participatory requirement is the commitment the practice has furthest to go on, and the foresight machinery is designed and partly built, not proven. Those are the two places to check us first.
So the proposition underneath all of it should be stated as what it is, and I should say plainly that I sell the alternative, which is a reason to read this paragraph and the one before it harder than the rest. Responsible AI does not mean only AI accompanied by the right documents. It means AI governed by people who keep the documents and the component reviews, and who also hold a theory adequate to the kind of thing they are governing, who model the mechanisms, watch the interactions, include the people inside the system, and correct course as the system moves. The compliance director in the opening scene needed nothing taken away from her. She needed one more instrument that could model the coupling and keep it under review. That proposition is a working hypothesis, and it is falsifiable, so here is the test. Take the failure class this essay has been about, harm that lives in the coupling between an AI system and the human process around it, and compare, over a stated period, what mechanism-level modeling and rehearsal put on the table before deployment against what component-and-calendar review, honestly applied, put on the table. If the second list is as long as the first, the hypothesis is wrong. Paramerge will publish its models, its rehearsals, and their misses as they exist, on paramerge.com alongside these essays, so that someone other than me can reach that verdict. What the record already documents is narrower and sufficient: in the domains this essay has walked through, the component-and-hindsight paradigm keeps failing in the same structural way, and the failures land on people.
The ontology is not optional. Every act of governance already has one. The only choice is whether it is examined.
References
Note on sources. References are carried from the verified records of the Paramerge white papers and the author's capstone research corpus, and every entry below was checked again against a live publisher record in August 2026. The Reed and Harvey entry credits the earlier joining of critical realism and complexity science on which Byrne's term builds; the term Complex Realism is Byrne's own. Houston and Swords is cited by its online-first year of 2021, and the article appeared in the March 2022 print issue. The two governance instruments, Leveson, and Ostrom are entries added in this revision.
Bhaskar, R. (1975). A realist theory of science. Leeds Books.
Byrne, D., and Callaghan, G. (2014). Complexity theory and the social sciences: The state of the art. Routledge.
Gerrits, L., and Verweij, S. (2013). Critical realism as a meta-framework for understanding the relationships between complexity and qualitative comparative analysis. Journal of Critical Realism, 12(2), 166 to 182.
Greenstein, R. (2025). Changes in the safety net over recent decades and their impact. The Hamilton Project, Brookings Institution.
Houston, S., and Swords, C. (2021). Critical realism, mimetic theory and social work. Journal of Social Work, 22(2), 345 to 363.
International Organization for Standardization and International Electrotechnical Commission. (2023). ISO/IEC 42001, Information technology, Artificial intelligence, Management system. ISO.
Leveson, N. G. (2011). Engineering a safer world: Systems thinking applied to safety. MIT Press.
Lieberman, S. (2026). What is Coulthart? Paramerge. https://paramerge.com/writing/what-is-coulthart
Lindblom, C. E. (1959). The science of "muddling through." Public Administration Review, 19(2), 79 to 88.
National Institute of Standards and Technology. (2023). Artificial intelligence risk management framework. NIST AI 100-1.
Ostrom, E. (2010). Beyond markets and states: Polycentric governance of complex economic systems. American Economic Review, 100(3), 641 to 672.
Reed, M., and Harvey, D. L. (1992). The new science and the old: Complexity and realism in the social sciences. Journal for the Theory of Social Behaviour, 22(4), 353 to 380.
Rittel, H. W. J., and Webber, M. M. (1973). Dilemmas in a general theory of planning. Policy Sciences, 4(2), 155 to 169.
Taleb, N. N. (2007). The black swan: The impact of the highly improbable. Random House.
Paramerge builds the coordination and governance instruments this essay describes, and works with teams that want their AI governed where the behavior actually lives.
Contact Paramerge