The Wicked Problem of AI in High-Stakes Organizations
Why the Systems That Hold Human Lives Cannot Be Governed One Component at a Time
Stephen Lieberman
Paramerge, Real-World AI Governance Center
August 2026
stephen@paramerge.com · paramerge.com
The determination nobody saw
Between 2013 and 2015, Michigan's unemployment insurance agency accused tens of thousands of residents of fraud, and many of them never saw the accusation. An automated system called MiDAS compared benefit claims against employer data, treated discrepancies as fraud, and issued determinations without human review, posting them to online accounts that people who had stopped drawing benefits had stopped checking. For many, the first notice was the consequence: a garnished paycheck, a seized tax refund, a penalty of four times the disputed amount. In two years the system adjudicated roughly 40,000 fraud cases by algorithm alone, and the reviews that followed found the overwhelming majority wrong, 85 percent of the auto-adjudicated determinations in one review and 93 percent in another examining a different case set (Charette, 2018; University of Michigan STPP, 2024). Some filed for bankruptcy under the collections.
Hold on to two things about MiDAS, because the rest of this essay turns on them. The instrument that finally established the wreckage was the most ordinary one governance owns. It was a component-level accuracy audit, a stack of determinations checked one by one against human judgment. It worked. It was also years late, applied after deployment instead of before it, after the garnishments instead of ahead of them. And there was a second category of harm that no accuracy audit could have reached even in principle, what happened downstream of each false flag, the debt compounding at four times its false premise, the credit damage, the housing stress, the years of repayment, harm that lived not in the determination but in everything the determination touched. Keep both halves. Here the component instrument was late where it would have worked, and blind where it could not reach.
An organization becomes high-stakes when its decisions are hard to reverse and land on people. Child welfare agencies, benefits administrators, hospitals, courts, unemployment offices, crisis services. In these organizations a wrong output is not a defect rate. It is a family separated, a disabled person left without care, a false accusation with a penalty attached. AI is now being installed inside these organizations at speed, twice over: purpose-built decision systems procured from vendors, and general-purpose tools adopted one overloaded practitioner at a time. The record of what happens when the installation outruns the governance is no longer hypothetical, and it is worth reading closely. It does not say what either side of the AI argument wants it to say.
Four systems, one shape
The clearest way to see the problem is four documented deployments. They are chosen because they failed, which is a selection this essay should name rather than hide. They cannot establish how often such systems fail, only how they fail when they do. Selection on failure carries a second limit, and it is the sharper one: these cases cannot show that component review generally misses, because a deployment where component review caught the harm in time would not appear in the record as a failure at all. What they can show is where the harm sat in these four, and that is the structural point the rest of this essay argues. The argument rests on the structure, not on the count.
MiDAS is the first, and its decision rule was genuinely defective. It averaged income across periods instead of reading individual paychecks, so the ordinary irregularity of low-wage work read as fraud, and it adjudicated without a human in the loop. Call that a broken component, because it was one. The finding that matters is what the brokenness did not explain. A defective rule became a catastrophe because nothing in the surrounding system was positioned to catch it, and because every false output cascaded through collections machinery built to compound it.
The second is Arkansas, where in 2016 the state's Medicaid program began allocating home-care hours to disabled residents, the hours a person is granted for bathing, eating, and moving through their home, by algorithm, applied to nurses' assessment data in place of nurses' judgment about hours. Bradley Ledgerwood, who has cerebral palsy and needs attendant care to perform every activity of his day, had been assessed at fifty-six weekly hours. In February 2016 the state cut him to thirty-two. He and six other residents sued, alleging the switch had cut their care by an average of 43 percent, and the Arkansas Supreme Court affirmed a restraining order protecting those seven people's hours, holding that they had shown a likelihood of prevailing on their claim that the state adopted the rule without substantially complying with the rulemaking process the state's law requires (Arkansas Department of Human Services v. Ledgerwood, 2017). Read that holding closely. The courts never reached the algorithm. They reached the paperwork, and when the state promulgated the rule properly, the methodology came back and the injunction was dissolved (Arkansas Department of Human Services v. Ledgerwood, 2019). It did not stay. Within months the state replaced the RUGs allocation with a differently derived instrument, an independent assessment built by an outside firm (Benefits Tech Advocacy Hub). The rulemaking the state had skipped the first time ran properly the second, and what came out of it was not what went in. That is the one place in this record where a procedural instrument did substantive work, and it is the reason I am telling it. It is also worth finishing the story. The replacement left more than a quarter of the program's participants determined wholly ineligible, capped the care of many who remained, and drew a second federal suit that the state lost on appeal and settled for close to half a million dollars. Procedure moved the instrument. It did not make it adequate. The legal defect was procedural. The human defect, as I read the record, was that a formula applied to assessment data cannot see what a nurse weighing the whole situation can: the informal caregiving network, the particular body, the particular home. Those are what decide what a number of hours means for a life.
The third is the Allegheny Family Screening Tool, widely cited as the first predictive risk model used by a public child welfare agency in call-screening decisions, launched in 2016 to score referrals using administrative data, with early versions predicting re-referral and out-of-home placement within two years (Allegheny County DHS, 2024; Eubanks, 2018). Consistency in screening was the design goal, and one hope was that a model would narrow the racial disparities documented in human screening. This is the most contested of the four cases, and honesty requires reporting what the evaluation record actually shows. The county's own published evaluations report narrowed racial gaps in screening decisions, and the study that looked at the screeners found that the tool acting alone would have produced more racially disparate decisions than the workers using it did, which credits part of the narrowing to the humans who overrode it (Allegheny County DHS, 2024; Cheng et al., 2022). The tool remains deployed, and this essay does not claim its measured effects run one way. What survives the contest is the structural finding Eubanks documented. The model learns from records of prior contact with public systems, and prior contact with public systems is what poverty looks like in administrative data (Eubanks, 2018; Keddell, 2019). Poor families are watched more closely, closer watching produces more records, the records feed the scores, and the scores direct the watching. That loop is a mechanism for manufacturing more of the world the model was trained on, wearing the appearance of objectivity while it does. Whether it is in fact closing is measurable. As far as I can find, nobody has measured it.
The fourth is COMPAS, the recidivism score used across multiple states in pretrial and correctional decisions. ProPublica's analysis found that among people who did not go on to reoffend, Black defendants had been flagged as future risks at 44.9 percent against 23.5 percent for white defendants (Angwin, Larson, Mattu, and Kirchner, 2016; Larson, Mattu, Kirchner, and Angwin, 2016). The finding is contested in a way that is itself instructive. The vendor answered that the score is equally calibrated across races, and Alexandra Chouldechova then showed that when the underlying rates differ between groups, a score calibrated that way cannot also produce equal false-positive rates, a result Kleinberg, Mullainathan, and Raghavan proved in a more general form (Chouldechova, 2017; Kleinberg, Mullainathan, and Raghavan, 2017). The choice between the two standards is not a technical question. It is a governance question, and it was answered by default, by nobody, inside a proprietary product. The compounding is what makes it a high-stakes-organization problem rather than a statistics dispute. A score that skews who is detained skews who is convicted, and a drug felony conviction carries exclusions from food assistance, in states that have kept, in some form, the ban the 1996 welfare reform act imposed and left every state free to drop, and from housing at the broad discretion federal law gives housing authorities. The record makes that chain plausible rather than measured, and tracing it end to end is precisely the kind of cross-system accounting no single institution is positioned to do. The harm compounds across institutions that never exchanged a word.
Four domains, one shape. In each case a component-level decision rule was inserted into a complex human system, and the component instruments failed it in four distinct ways. Where the rule itself was defective, as in Michigan, the audit that would have caught it arrived only after years of harm. Where the instrument was aimed correctly and skipped, as in Arkansas, where a rulemaking process existed and the state went around it, the instrument would have worked and nobody ran it. Where component review ran on the calendar and reported, as in Allegheny, it answered the question it was asked, which is not the same as the question that matters, and the harms lived in the feedback between surveillance and the data that surveillance produces. And where the governing choice was a contested question of fairness, as in COMPAS, it was answered by default, by nobody, inside a proprietary product, and one system's output compounded through another system's rules. The deference of the humans who read a score belongs on that list as a documented general phenomenon, not as a finding of these four. Allegheny is the only one of them with evidence on the point, and its evidence runs the other way.
Medicine has an old word for injury administered by the treatment itself. Its history is full of therapies given faithfully, by the book, that were themselves the harm. Bloodletting is the textbook case, and the mechanism is worth stating, because the mechanism is the whole analogy. The physician drew off the blood volume the patient needed in order to fight the illness, and the sicker the patient then looked, the stronger the case for drawing more. The treatment's own signal of failure was read as a call for a larger dose. That is iatrogenic harm, and it locates the injury correctly. Not in malice, not in carelessness, but in the intervention doing what it does inside a body nobody fully modeled. Michigan is the same shape without the blood. The penalty multiplied the disputed amount fourfold on a flag no human had read, so the system's confidence in its own output was exactly what set the dose, and the debt that confidence created was collected before anyone asked whether the flag was right. The treatment can deepen the disease.
The shape of the problem
Rittel and Webber (1973) gave the term wicked problem a precise meaning, and four of their clauses do the work here. The extension to AI governance is my claim, not theirs. The problem's definition is entangled with its proposed solutions: decide the problem is fraud and you build MiDAS, decide it is inconsistent screening and you build a risk score, and each definition manufactures its own evidence. There is no stopping rule, no criterion that tells you the analysis is finished. You stop when time or money runs out, not when the problem is solved. Solutions are not true or false but better or worse, judged differently by the agency, the practitioner, and the family on the receiving end. And deployments are, in their phrase, one-shot operations. The consequences begin with the first case and cannot be taken back. Rehearsal in simulation, which this practice builds and the companion essays describe, does not dissolve that property. A rehearsed failure is a hypothesis you get to have before the first case instead of after it. The first real case is still one-shot, which is why the rehearsal discipline holds that a rehearsed consensus is never evidence the framework holds. A rehearsal is not a result.
Underneath the wickedness is the structure this body of writing keeps returning to. A high-stakes organization is a complex adaptive system. Its outcomes emerge from interactions among caseloads, incentives, data systems, workarounds, and the lives of the people it serves. Feedback loops, like the surveillance loop in the Allegheny data. Sensitivity to small design choices, like the delivery details that shaped families' experience of the cash-transfer pilots in Stockton and Chicago (West, Castro Baker, Samra, and Coltrera, 2021; Stapleton et al., 2024), programs with no AI in them at all, which is the point. The complexity belongs to the environment, and it is waiting for whatever is deployed into it. Path dependence, where a false fraud flag cascades through garnishment, credit, and housing long after the flag is corrected. Social work has a name for this insight and formalized it generations ago, the person in environment: a person cannot be understood apart from the nested environments they live inside. Neither can a decision about them.
I have argued in the essay published alongside this one that every governance instrument carries a theory of reality. The instruments now aimed at AI are not naive. They are iterative by design. What they presuppose is narrower and more consequential. They observe components, one at a time, on a calendar, and they assume the relationships among the components hold still long enough to measure. Every harm in the four cases above lived either in the gap between review cycles or in an interaction no component review examines: between the model and its historical data, between one institution's output and another's eligibility rule, between a determination and the collections machinery waiting behind it. The claim is falsifiable, and here is its test. Take a set of deployment failures registered in advance and defined by their consequences, never by where in the system they sat. Run both instruments at them. If system-level analysis anticipates those failures no better than component review does, the argument of this essay is wrong.
The second wave is already inside the building
The four cases were purpose-built systems that arrived through procurement, visible to leadership, contestable, eventually, in court. The second wave is different. General-purpose AI tools are entering high-stakes organizations from the bottom up, one practitioner at a time, mostly without any organizational decision at all. In social work, the profession whose evidence I know best, three in ten of the 860 practitioners answering a recent national survey reported no AI adoption plan in their department at all. Nearly two-thirds were already using AI in the role, on the survey's own report and uncorroborated (Borah and Landers, 2026). The evidence review commissioned by the profession's English regulator reached the boundary that matters. The demonstrated benefits so far are administrative, and AI is not yet reliable for risk prediction or decision-making in children's and adults' services (Social Work England, 2025).
The adoption is rational adaptation to unsustainable caseloads, a case I have made in full in the companion essay on what we stand to gain, and the pattern of the adoption is the old one. The sequence is on the profession's own record, across every modern transition. Statewide case-management systems arrived through the 1990s under federal funding rules, with practitioners rarely at the table. Electronic records spread through the 2000s, shaped by billing logic that compressed clinical narrative into checkboxes. Telehealth exploded in 2020, far ahead of the reimbursement rules and organizational policies that would eventually govern it, with the profession's 2017 technology standards straining to cover pandemic-scale practice. Each time, governance arrived after deployment, and each time the populations most affected by the gap were the ones the professions were most obligated to protect. AI is running the same pattern with greater velocity, a wider surface, and systems that resist audit in ways none of the earlier technologies did.
And the next MiDAS will likely be less legible than the first one. A determination posted to an unwatched portal was at least a document, discoverable, auditable, actionable in court. A caseload summary quietly weighted by a drafting tool is not a document anybody can sue over. It is a nudge inside a workflow, invisible unless the organization built the instruments to see it.
One more feature of the second wave deserves plain statement. The professional duty of confidentiality sits on the licensed practitioner, and license sanction lands on the licensee. Regulatory fines for a breach land mostly on the organization. And the standard commercial terms I have read disclaim the vendor's warranties and cap the vendor's liability, a reading I offer as my own, since no systematic survey of these contracts appears to exist. The net of those three mechanisms is an allocation problem. The practitioner holds exposure that can end a career and the client bears the harm, both out of all proportion to their say in how the systems are built, while the organization's exposure is a cost and the organization is the party with design authority. Correcting that is governance work the professions can only do together, by deliberating their own conditions of practice, what must be disclosed, what uses are prohibited, what may never leave protected infrastructure. No exhausted practitioner should have to decide it alone.
What answers it
If the diagnosis is right, the response has to match the structure of the problem. Three lines of response follow from the record, and their maturity should be stated exactly. Parts are built, parts are designed, and none is proven for this use.
The first is the unit of analysis. Govern the sociotechnical system, and the component too. For automated adjudication, the component-level clauses are already legible in Michigan's record, in the negative. No adverse determination without human adjudication, and no penalty multiplier on a flag no human has reviewed. The system-level questions are the ones no instrument inside the governance of the four deployments was pointed at. What happens to a person falsely flagged, and how fast does the harm compound before correction? What does a score do to the humans who read it? What does the training data actually measure, and what loop closes when the system's outputs become its future inputs? Researchers did point at some of these, years after the fact, from outside, with no standing to stop anything. They are the iatrogenic questions, and they have to be asked while the dose can still be changed. Paramerge's research program exists to make questions of that kind askable before deployment, on modeled organizations grounded in the documented record, with structure and interaction as the unit of analysis, each claim tied to a published source or marked plainly as a model output. A model of an organization is still a model. It can be wrong, and no instrument, this one included, has yet been demonstrated to anticipate failures of this class in advance. That is the falsifiable wager the foundations essay states, not a capability claim.
The second is coordination, because no single organization can govern this alone. A hospital system, an agency, and a professional body adopting the same class of tool face the same failure modes and currently learn about them separately, through their own casualties. The companion essay lays out the program. Map where the parties' operative meanings of risk, evidence, and safety diverge before the divergences are load-bearing, a capability there stated as a hypothesis under test. Rehearse candidate frameworks against grounded simulated populations, with rehearsal outputs held as hypotheses, never forecasts. Treat willingness to walk away as data, in a mechanism specified but not built. And keep every decision human, a rule that fixes accountability without removing the influence of what the humans are shown. Coordination reaches the part of the COMPAS lesson where institutions adopting the same class of tool could see a failure mode at the same time, rather than one casualty at a time, and no body above them currently does that work. It does not reach the rest. The exclusions that turn a conviction into hunger and homelessness were written by legislatures. Only a legislature can unwrite them.
The third is specification discipline. In the essay on what we stand to gain, the governed version of an AI-assisted child welfare meeting is written explicitly as a specification, not a report of anything deployed: a tool prohibited from recommending removal, required to show in the record what it foregrounded and what it left out, its influence on the record inspectable by the parent it concerns, with the practitioner's judgment intact and the shaping of that judgment on the record, since no rule fully cures the shaping. Read against the four cases, each clause makes a documented harm inspectable rather than invisible. It is also worth being honest about the limits of transfer. That specification governs an assistive tool. Applied to a screening score like Allegheny's it would amount to prohibition, not governance, and what would govern a screening score, continuous examination of what its proxies measure and whether its feedback loop is closing, is designed work, not demonstrated work. Arkansas adds its own lesson twice over. The instrument that stopped the reductions for those seven people was an ordinary compliance requirement, applied, which is a reminder that the boring instruments matter. And what the process failure cost them was the ability to understand and contest the number that ran their lives. Bradley Ledgerwood's hours went from fifty-six to thirty-two. Per-decision inspectability is what would let a person in that position see the reasoning that moved them.
The determination, again
Michigan residents repaid debts for years, for frauds that never happened, assessed by a system with a defective rule, announced in a portal nobody checked, compounded by penalties nobody reviewed. Every instrument that finally established what happened, the audits, the litigation, the journalism, worked at the level of components and after the fact. Nothing stood at the level of the system, and nothing stood before deployment.
High-stakes organizations are now absorbing systems more capable, more general, and less auditable than MiDAS, faster than any technology they have absorbed before, under governance instruments aimed at components on a calendar. The failures documented here were established eventually, by courts, journalists, and researchers, years into the harm, and the systems mostly outlived the findings. The recidivism score is still in use. The screening tool is still deployed. In Arkansas the methodology came back as soon as the paperwork was right, and it took a full comment process to displace it. The tools now arriving will not necessarily fail as legibly as a portal full of false determinations.
That is the case for governing these systems as what they are: parts of complex human systems whose behavior lives in the interactions. The instruments for seeing at that level are the work, some built, some designed, none yet proven, and the honest version of this essay's promise is small. Seeing the system is not a guarantee of catching the next failure. It is the difference between a failure that had to wait years for an audit, and one that at least had something pointed at the interactions.
References
Note on sources. References are carried from the author's capstone research corpus and verified evidence ledgers, with the case-record entries checked against the published opinions, the county and regulator publications against the documents themselves, the state reviews as reported in the records cited, and the original analyses in August 2026.
Allegheny County Department of Human Services. (2024). Summarizing recent research on predictive risk models in child welfare (24-ACDHS-04). Allegheny County Department of Human Services.
Angwin, J., Larson, J., Mattu, S., and Kirchner, L. (2016). Machine bias. ProPublica.
Arkansas Department of Human Services v. Ledgerwood, 2017 Ark. 308, 530 S.W.3d 336 (Ark. 2017).
Arkansas Department of Human Services v. Ledgerwood, 2019 Ark. 100 (Ark. 2019).
Benefits Tech Advocacy Hub. Arkansas Medicaid home and community based services hours cuts. Upturn, National Health Law Program, and TechTonic Justice. https://www.btah.org/case-study/arkansas-medicaid-home-and-community-based-services-hours-cuts.html
Borah, E., and Landers, J. (2026). AI in social work: Survey reveals widespread adoption amid infrastructure gap. Steve Hicks School of Social Work, University of Texas at Austin.
Charette, R. N. (2018). Michigan's MiDAS unemployment system: Algorithm alchemy created lead, not gold. IEEE Spectrum.
Cheng, H.-F., Stapleton, L., Kawakami, A., Sivaraman, V., Cheng, Y., Qing, D., Perer, A., Holstein, K., Wu, Z. S., and Zhu, H. (2022). How child welfare workers reduce racial disparities in algorithmic decisions. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (pp. 1 to 22). Association for Computing Machinery. https://doi.org/10.1145/3491102.3501831
Chouldechova, A. (2017). Fair prediction with disparate impact: A study of bias in recidivism prediction instruments. Big Data, 5(2), 153 to 163. https://doi.org/10.1089/big.2016.0047
Eubanks, V. (2018). Automating inequality: How high-tech tools profile, police, and punish the poor. St. Martin's Press.
Keddell, E. (2019). Algorithmic justice in child protection: Statistical fairness, social justice and the implications for practice. Social Sciences, 8(10), 281. https://doi.org/10.3390/socsci8100281
Kleinberg, J., Mullainathan, S., and Raghavan, M. (2017). Inherent trade-offs in the fair determination of risk scores. In Proceedings of the 8th Innovations in Theoretical Computer Science Conference. Schloss Dagstuhl.
Larson, J., Mattu, S., Kirchner, L., and Angwin, J. (2016). How we analyzed the COMPAS recidivism algorithm. ProPublica.
Rittel, H. W. J., and Webber, M. M. (1973). Dilemmas in a general theory of planning. Policy Sciences, 4(2), 155 to 169.
Social Work England. (2025). The emerging use of artificial intelligence (AI) in social work. Social Work England commissioned evidence review.
Stapleton, S., Abdul-Razzak, N., Brady, F., Croes, M., Leader-Smith, A., Schexnider, M., and Wallace, N. (2024). Big shoulders: Implementing the Chicago Resilient Communities Pilot. University of Chicago Inclusive Economy Lab.
University of Michigan Science, Technology, and Public Policy Program. (2024). Case over the Michigan Unemployment Insurance Agency's faulty automated system finally settled. University of Michigan Ford School of Public Policy.
West, S., Castro Baker, A., Samra, S., and Coltrera, E. (2021). Preliminary analysis: SEED's first year. Stockton Economic Empowerment Demonstration. https://www.stocktondemonstration.org/
Paramerge builds the coordination and governance instruments this essay describes, and works with teams that want their AI governed where the behavior actually lives.
Contact Paramerge