ParamergeParamerge

Domains

Domain Atlas

How does this show up in specific domains?

High-stakes human systems are the frontier for AI governance: predictive scores and generative assistants already shape who gets investigated, helped, hired, treated, paid, and believed. 18 domains across 10 sectors, 66 use cases, and 119 documented case files, every factual claim cited to the Evidence Registry.

Domain Atlas

The Domains

Each sector groups the domains where the same governance physics — pressures, pathways, and levers — plays out on its own institutions.

Social Services

6 domains

Public human services where predictive scores and drafted records shape who gets investigated, helped, housed, paid, and believed.

Child welfare & family services

14 case files

Predictive screening and profiling where the cost of both false alarms and misses lands on families — and where the human override layer has measurably mattered.

Public benefits & eligibility

21 case files

Fraud scoring, eligibility automation, and care allocation — the domain with the largest documented harms, almost all of them ending in courts and commissions.

Caseworker documentation & copilots

11 case files

Generative and rules-based assistants that transcribe, summarize, triage, and increasingly draft the records and decisions institutions run on — meeting scribes, mail-triage filters, evidence summarizers, ruling drafters, and automation that removes the caseworker from the routine path entirely. The human review step is the load-bearing control, and the throughput that justifies the tool is the same pressure that erodes it — with a verifier that checks human work at one edge of the range and deployments halted after the harm was measured at the other.

Benefits navigation & public-facing chat

11 case files

Conversational systems standing between the public and their benefits — public-facing chatbots, adviser-gated copilots, federated whole-of-government fleets, and navigation intermediaries whose quiet removal is itself the harm. An authoritative wrong answer is indistinguishable, to its victim, from policy; what sets the exposure is whether a professional gates the answer, whether anyone measures accuracy at all rather than mere deflection, and whether the whole channel rests on a single actor who can switch it off. The Lab networks in this domain model only what happens inside the operating organization — its operators, engines, and knowledge stores; the members of the public asking the questions sit outside the dynamics, and harm to them is documented in each case file, never computed on a diagram.

Housing & homelessness services

11 case files

Prioritization and prevention scores deciding who reaches scarce housing help first — where a more accurate model can still leave the same people under-served, and the quietest harm is often the person the system never surfaced.

Behavioral-health & crisis triage

11 case files

Risk scores and triage rankers deciding whose crisis is seen first — where the rare event is nearly impossible to predict reliably, the flag moves a proxy more surely than the outcome, and a score can quietly gate access to care. The counterweight is on the record too: a national health system's opioid risk-mitigation dashboard, evaluated in a randomized design across its medical centers, was reported as associated with a decrease in mortality among the at-risk patients it covered — a benefit direction this domain is rarely credited with, and an outcome for people who sit outside every diagram here.

Healthcare & Medicine

2 domains

Clinical decision support and ambient documentation — AI whose output lands in the chart and on the care team.

Clinical decision support & deterioration alerting

4 case files

Machine-learning early-warning models that flag hospitalized patients for sepsis or clinical deterioration — the domain where an AI's measured benefit is real but runs entirely through the human loop it interrupts. The same alert that saves a life when a clinician confirms it in time becomes a source of fatigue when it fires a hundred times per true case; what separates the two is whether the confirmation workflow is resourced, whether the model was validated independently of the vendor who sells it, and whether anyone reconciles the alerts against the outcomes they were meant to change. The Lab networks here model only the deploying hospital — its models, clinicians, and records; the patients being scored sit outside the dynamics, and no clinical outcome is ever computed on a diagram.

Clinical documentation copilots (ambient scribes)

3 case files

Ambient AI that records the clinical visit and drafts the note for a clinician to edit and sign — the domain where the measured benefit is real (documentation time returned, work exhaustion reduced) but the governable object is the permanent record itself. Today's AI-drafted note becomes tomorrow's copied-forward clinical fact: later clinicians and later tools read it as ground truth, so the clinician's review and any standing quality-assurance program are not politeness — they are the contamination controls on a record that ambient notes are documented to hallucinate into about a third of the time. The benefit is also heterogeneous: the same tool, in the same system, helps one clinician group and largely fails another. The Lab networks here model only the deploying organization; the patients whose visits are transcribed sit outside the dynamics, and no care outcome is computed on any diagram.

Legal & Government

1 domain

State determinations about a person — asylum, enforcement, adjudication — where the affected person can least see or contest the AI.

Immigration & asylum AI

2 case files

Immigration and asylum are among the highest-stakes decisions a state makes about a person — asylum, removal, detention — so the reliability of any AI signal and the weight it is given matter more here than almost anywhere, and the person affected is typically least able to see or contest the AI. Two patterns anchor the governance. In credibility assessment, a federal asylum agency uses dialect-recognition AI to estimate an applicant's origin from a speech sample; the tool is imprecise (government-reported recognition around 80 percent for one language, and linguists judge separating some varieties close to hopeless) and the agency's own caseworkers call it a rough compass, too imprecise to resolve hard cases. Used honestly as one clue among several it is defensible; the documented risk is that an imprecise output acquires more authority than its accuracy supports, in a determination where the state's tool is set against the applicant's own account. In enforcement triage, an algorithm identifies and recommends people for immigration actions, and a design detail uncovered through freedom-of-information litigation makes the governance concrete: officials had to justify rejecting a recommendation but not accepting one — an asymmetric override that builds a rubber-stamping incentive into the workflow, so the human in the loop is nominal rather than real. Running under both is a severed correction loop on the applicant's side: applicants are frequently not told AI is used, so the person with the most at stake and the most knowledge of the truth cannot contest the signal used to decide their case. The Lab networks model only the deploying government body — its model or triage tool, the caseworkers and officers who act on its outputs, and its case records; the applicants being decided sit outside the dynamics, no asylum or enforcement outcome is computed on any diagram, and reliability, override-design, and disclosure findings are recorded external facts (reports, fieldwork, FOI disclosures), never adjudications of any individual case.

Finance & Security

2 domains

Credit decisioning, fraud and financial-crime detection — domains governed by base rates, disparate impact, and the duty to explain.

Security operations & fraud detection

3 case files

Machine-learning fraud, anti-money-laundering, and financial-crime detection — the domain governed by extreme base rates, where the arithmetic itself sets the limits. When the thing being detected is rare, detection precision is dominated by the false-alarm rate rather than by accuracy, so a threshold change moves the burden of alerts rather than the truth of them; and because only a few flagged cases are ever verified, models retrain on the investigators' own dispositions, so a rise in 'confirmed' activity can be partly a measure of what the system taught its reviewers to confirm. Two harms sit on opposite sides of the same score: the false-positive tail lands on real people — frozen accounts, weeks without funds — while the missed cases are the fraud the system exists to catch, and detection quality and the justice of the disposition are different levers held by different actors. Almost every deployment-scale benefit number in this domain is a vendor self-report with no independent audit; the model enters them as claimed magnitudes and says so. The Lab networks model only the deploying organization — its models, investigators, and case stores; the account holders and flagged parties sit outside the dynamics, and no customer outcome is computed on any diagram.

Lending & credit collections AI

3 case files

Machine-learning underwriting, pricing, and credit decisioning — often fully automated, with no per-application human review, so the organizational levers all sit upstream: the choice of model, the fair-lending testing regime, the search for a less-discriminatory alternative, and the adverse-action notice that must explain a denial. Three facts shape the governance. A denial's explanation is a separately-resourced, separately-failable duty independent of statistical bias: a model can pass the disparity test and the organization can still fail by being unable to give an applicant specific, accurate reasons. Facially-neutral aggregate features — a school's default rate priced into an individual's terms — can carry protected-class impact, which is exactly what disparate-impact testing exists to catch. And the harder governance question is not whether a disparity exists but how hard the law requires an organization to search for a less-discriminatory model that performs as well — a question the record shows resolved by enforcement or left at an impasse, rather than settled. The Lab networks model only the deploying organization — its model, compliance and testing functions, and decision records; the applicants being decided sit outside the dynamics, and no credit outcome is computed on any diagram.

Education

1 domain

Early-warning prediction, proctoring, tutoring, and grading — deployments whose subjects are students, often minors.

Education AI

2 case files

AI in education spans early-warning prediction, remote-exam proctoring, tutoring assistants, and automated grading, and its subjects are students — often minors — so the deploying institution's duty of care and the training of the staff who read the AI's output are load-bearing. Two deployment patterns anchor the governance. In predictive early warning, a model labels a student's risk of not graduating and delivers it to school staff, and the label is only as good as the intervention it triggers and the training of the human who reads it: a statewide dropout early-warning system was found by an independent audit to be wrong most of the time when it predicted a student would not graduate, to produce higher false-alarm rates for Black and Hispanic students, and to reach staff who reported no training on how to interpret a 'high risk' label — so the model's error and its group disparity were imported into how students were seen rather than into help they received, and the system was withdrawn. Against that, a large district's transparent, low-tech on-track indicator, paired with real intervention, accompanied a rise in graduation to a record level — locating the benefit in the intervention the indicator makes legible, not in the sophistication of the prediction. In remote proctoring, surveillance-based AI carries a rights cost that can be independently adjudicated: a public university's requirement that students pan their webcam around their home before an exam was held to be an unreasonable search, and peer-reviewed measurement found proctoring software flags darker-skinned and Black students more often with no corresponding difference in actual cheating. The Lab networks model only the deploying institution — its model or proctoring tool, the staff and proctors who read its output, and its student records; the students being scored or watched sit outside the dynamics, and no student outcome is computed on any diagram.

Industrial Logistics & Operations

2 domains

Inspection, predictive maintenance, routing, and algorithmic management — optimization whose costs land on the workers executing it.

Industrial QA & operations AI

3 case files

On the factory line, AI takes two main forms: automated visual and acoustic inspection that flags defects, and predictive maintenance that forecasts equipment failures from sensor data. The governing fact in both is that the AI flags and a human responds — the inspection is only as good as the response it triggers, and the benefit runs through a resourced human-response loop (a line worker who can stop the line, a maintenance crew that acts on an alert), not through the model alone. That makes the failure modes a matter of the loop's calibration, and the honest evidence for them is mechanism-level rather than incident-level. Four mechanisms are well documented in the research literature: false alarms, which pile up until operators stop trusting the alerts (alert fatigue); drift, where the model degrades as the line, the parts, or the sensors change; false rejects, where good product is scrapped because the classifier is tuned to over-flag; and over-trust, where operators defer to the AI and stop checking, so a missed defect passes because the human loop that was supposed to catch it had already deferred to the thing that missed it. What the public record does NOT contain — and this domain states it plainly — is a named manufacturer publicly attributing a shipped-defect escape or a recall to its AI inspection system; that specific incident class appears to stay inside plants, so the failure regime here is modeled at the mechanism level, and nothing in it should be read as a claim that a named company's AI let a defect ship. The benefit side is real but reported through corporate and trade channels for the named deployments, with the peer-reviewed quantitative results coming from smaller or anonymized sites. The Lab networks model only the deploying organization — its inspection or maintenance model, its line operators and maintenance crews, and its quality records; the products being inspected and the people who use them sit outside the dynamics, and no product-safety or defect-escape outcome is computed on any diagram.

Logistics dispatch & scheduling AI

2 case files

AI in logistics optimizes routing, dispatch, and warehouse task assignment, and the pattern that defines its governance is that the same system which optimizes the work also manages the worker doing it. A parcel carrier's route-optimization system is a documented operations-research success — reported to save on the order of a hundred million miles and millions of gallons of fuel a year — and the same system dictates the route to the driver and monitors adherence through telematics, so the efficiency is enforced through workplace surveillance and the driver's discretion is what it replaces. A warehouse's algorithmic management pairs a genuine human-robot picking benefit with a documented injury-productivity trade-off: when the algorithm sets the pace, regulators and a legislative inquiry have tied the speed it demands to ergonomic hazards and to warehouses described as uniquely dangerous, so the productivity gain and the worker-injury risk are coupled. Two things follow. The efficiency metrics — miles, fuel, throughput, units per hour — measure the optimization's success and are silent on its cost, which shows up in injury data, in surveillance the worker experiences, and in ethnographic research, not on the operations dashboard. And the workers being managed are, in effect, the served people of this domain: the person executing the AI's plan is also the person the optimization presses on, so the governable question is whether the system internalizes the human executing it — a pace that is feasible and safe, monitoring that is proportionate — or externalizes that cost as an injury or an autonomy loss the productivity number never sees. The Lab networks model only the deploying organization — its optimization or management model, the drivers and pickers who execute its plans, and its operations records; no worker-injury or safety outcome is computed on any diagram, and injury and surveillance findings are recorded external facts, never diagram-derived.

Media & Platforms

1 domain

Content moderation and editorial AI — classifiers and drafting tools that decide what the public sees.

Content moderation & editorial AI

3 case files

This domain covers two AI deployments that both decide what the public sees: automated content moderation on platforms, and AI-drafted editorial content in newsrooms. In moderation, machine classifiers remove content proactively — often before any user has seen it — at a scale no human review could match, and the governing fact is that the human review and appeals path is the error-correction loop, not an optional add-on. A platform's own natural experiment made this concrete: when human reviewers were sent home during the pandemic and the platform deliberately chose over-enforcement, automated removals more than doubled, appeals roughly doubled, and the reinstatement rate on appeal jumped from about a quarter to about half — direct evidence that the automation was making roughly twice the rate of catchable errors, visible only because the appeals queue caught them. Two things follow. Proactive removal acts before anyone sees the content, so an over-broad takedown is invisible unless an appeals path surfaces it — and in documented cases automated removal destroyed evidence of war crimes with archival access declined, an irreversible action with no correction loop at all. And over-enforcement versus under-enforcement is a chosen trade-off: when you cannot review everything, you are choosing which error to make, and that choice is a governance decision, not a technical default. In the newsroom, AI-drafted articles published under a human byline without disclosure are an accountability failure of a different shape — one outlet's audit found it had to correct a large share of its AI-written articles, so the review that a byline implies was not actually performed. The Lab networks model only the deploying organization — its classifiers or drafting tools, its reviewers, editors, and appeals functions, and its enforcement or publication records; the people whose content is moderated or who read the articles sit outside the dynamics, and no user or reader outcome is computed on any diagram.

Commerce & Customer Service

1 domain

Contact-centre copilots and customer-facing chatbots — where deflection is not resolution and the escalation path is the safety valve.

Customer service & contact-centre AI

3 case files

AI in the contact centre takes two forms, and they fail in different ways. Agent-assist copilots draft suggested responses for a human agent, and the strongest field evidence shows their benefit is real and unevenly distributed: a large randomized rollout raised issues-resolved-per-hour by roughly 15 percent on average, but almost the entire gain went to novice and low-skill agents while the most experienced gained close to nothing — so an average productivity number overstates the effect for the agents who least need it and hides that a copilot mostly raises the floor. Customer-facing chatbots answer the person directly, and here the governance facts are sharper. The organization is accountable for what its chatbot tells a customer: a tribunal has held that a chatbot is a tool the company is responsible for, not a separate entity that answers for itself, so a hallucinated policy or a wrong fare rule is the company's misrepresentation. Deflection is not resolution: routing a contact away from a human can suppress the very signal the customer needed to send, and surveys show most customers would rather not meet AI in service at all and fear it makes reaching a human harder — which makes the escalation path to a person the system's real safety valve. And a published deflection number is not a settled result: the same organization that reported handling two-thirds of its chats with AI later walked back on quality grounds and committed to keeping a human available. The Lab networks model only the deploying organization — its agents, its QA and escalation functions, and its interaction records; the customers being served sit outside the dynamics, and no customer outcome is computed on any diagram.

Pressures

Institutional pressures

The recurring forces that bend deployed systems away from their evaluated behavior. Each domain page names the pressures that dominate it. This vocabulary is conceptual framing, drawn from the documented cases.

Caseload surge

Demand outruns staffing; per-case attention shrinks and review becomes triage.

Reviewer bottleneck

One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.

Austerity & recovery incentives

Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.

Vendor opacity

The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.

Deadline pressure

Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.

Staff turnover

Experienced skepticism leaves; new staff calibrate their trust on the tool itself.

Data & policy drift

The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).

Compliance over substance

Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

Want to see these pressures act on a system? Stress-test them in the PAN Lab →

These case files are documented after the harm. Mapping a live deployment's pathways and pressures before the incident report is engagement work.

Work With Paramerge