PAN Lab: govern the network before failure spreads

Legend

Start here

Challenge of the Day

Challenge result

Round summary

Choose a starting point

Apply institutional pressure

Levers & components

System Readout

Edit this network

Explore's workshop — authored parts and pathways, saved in this browser

From Simulation to Lab — Sociotechnical Systems Modeling and Simulation

Modeling evidence and assumptions behind this network

The same AI running hands off: the agentic office

The AI here doesn't just draft casework — it acts on cases, and a stretched staff waves most of it through. Same model as the other two offices. In the sociotechnical simulation — a modeled office, not a real one — errors stick here about 75% of the time, versus 20% and 16% next door. Find the levers that change that.

Learn more about this example — its sources, evidence, and the concepts, pressures and levers behind it.

A narrated tour of the whole Lab — or today's ranked governance exercise.

Description

The same AI running hands off: the agentic office. No pressures applied. No governance levers in place.

Who and what is in the system

  • Drafts decisions and, as an agent, acts on cases with minimal review.
  • A small team supervising many automated actions at once.
  • The shared record system both people and agents read and write.
  • Pulls prior records into the model's context automatically.

What each pathway is doing right now

  • delivers the AI's work to staff: high, and failure flow high. Agent outputs pass to staff and are adopted into their work.
  • writes the AI's work into the records: medium, and failure flow medium. Agent writes case records directly.
  • staff frame the work for the AI, and failure flow medium. Hurried prompts frame the model toward confirmation.
  • staff file the work into the records, and failure flow high. Adopted outputs documented into the record.
  • grounds the AI's work in the records, and failure flow medium. Retrieved records re-enter drafts as fresh context.
  • staff work from the records, and failure flow medium. Staff read and rely on the record as written.
  • staff share working habits, and failure flow medium. Shortcuts and adopted claims spread between coworkers.
  • one AI hands the work to the next, and failure flow medium. Agent outputs chained into other agents' context, client context riding along.
  • checking off — a protective check, where a higher level is better. Second opinions between coworkers — rare under this load.
  • checking off — a protective check, where a higher level is better. An independent model checks the agents' outputs.

Where the gauges sit

The AI reads not helping — work getting done is strained against demand. How the AI is helping: ; ; ; .

The reads cascading failure. ; ; ; .

Under heavy incoming pressure: , , .

Sources & Evidence — what this Lab is and is not

Network ID: office-agentic-explore

Where each failure mode lives here

The Lab speaks in pathways, pressures, levers, and gauges. This map connects that vocabulary to the formal failure-mode names used in the research grounding it — including a 2026 national survey of 1,179 U.S.-based social workers.

  • Automation bias / overreliancecore

    The “Failures adopted by people or agents” pathway and the operator-deference-drift gauge. Staff turnover and autonomy expansion push it up; the deskilling-arrest lever caps it. Deference can also rise where accountability sits rather than where trust does: when following the tool is the defensible act, a worker who distrusts it may follow it anyway — and a training lever does not reach that driver.

    Evidence: In the 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers, 40.8% of respondents reported ethical concerns about relying on AI for decision-making, and overreliance on automated decision-making was among the most frequently cited concerns overall.

    Evidence: The same chapter reports that workers who distrust a screening tool may still follow it, because organizational and policy pressure makes following the tool the defensible act. Deference on this account is produced by where accountability sits, not only by how much the worker trusts the output.

    Evidence: The volume's child-welfare chapter reports that a widely deployed screening score predicts whether a child will be placed out of the home within two years, which is the system's own future response rather than the maltreatment the worker is deciding about. The chapter treats the gap between the modelled target and the decision's actual question as a design property of the deployment, not as a defect in the model's accuracy.

  • Deskilling / professional judgment erosioncore

    Overreliance in slow motion: the deference gauge drifting upward while correction capacity thins. The deskilling-arrest lever is its deliberate counter-schedule.

  • Sycophancy / agreement-seeking outputcore

    The pushback-heavy-usage pressure runs the “Operator framing biases the model” pathway hot and makes the agreeable answers easier to adopt; the framing-hygiene lever dampens the loop at its origin.

    Evidence: Research on AI sycophancy describes it as a fragmented construct — a family of distinct agreement-seeking behaviors that share a label but differ in form, mechanism, measurement, and required mitigation — and finds it intensifies under user pushback and across multi-turn interaction.

  • Hallucination / incorrect-output propagationcore

    The Lab's core premise: every pathway in the diagram carries incorrect output away from its source, and every lever is a way of governing that propagation rather than assuming a perfect model.

  • Unsafe data flow / privacy & confidentialitycore

    Modeled as pathways, gauged as exposure (Phase NP). Unsanctioned tool use opens a visible egress to the off-network sink; connector sprawl replicates an unverified cache between record systems; case-file-flagged pathways (MiDAS-class enforcement replication, records feeding vendor models) carry the same concern. While any such pathway runs, the Privacy gauge drains — and in Hard and Expert a full gauge is part of the win. “Vet connections” cuts the pathways structurally; “Store less data” shrinks what there is to expose.

    Evidence: In the 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers, concerns about data privacy and security were the most frequently reported challenge to using AI in practice (46.5% of respondents), and an increased focus on client privacy and confidentiality was the most requested improvement to AI tools for social work (50.4%).

  • Bias propagation / institutional workflow biascore

    Biased framings and contaminated records travel the same workflow pathways failures do — an institutional propagation question, and the documented cases show the workflow (not the model alone) carrying the equity outcome in both directions. The Lab models no demographics: differential harm to served people is recorded outside the network, never computed from its dynamics.

    Evidence: Evaluation evidence on the Allegheny Family Screening Tool found that screener overrides of the tool's recommendations reduced racial disparity in screen-in rates relative to the tool alone.

    Evidence: Independent scrutiny of Rotterdam's welfare-fraud risk model — a 2021 municipal audit followed by a 2023 journalistic investigation that obtained the model itself — documented scores skewed against already-vulnerable groups, and the city suspended the system's use.

  • Transparency / provenance failurecore

    The record-contamination gauge reads how much unlabeled machine content feeds back into decisions; “Mark AI-written records” discounts it and “Gate vendor updates” attacks opacity at procurement.

  • Weak human oversight / safeguard failurecore

    The scenario axis itself: one model, three oversight cultures, three very different outcomes. The correction and authority gauges track it; “Review the riskiest first”, “Pause AI on alarms”, “Require sign-off”, and “Review on schedule” govern it. A control drawn on one of these diagrams is drawn as present, which is a claim about structure — not a claim that anyone exercises it.

    Evidence: In the 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers, 42.1% of respondents reported having no role in decision-making about AI adoption in their workplace; the report concludes most respondents have limited or no control over how AI technologies are selected or implemented within their organizations.

    Evidence: The volume's ethics chapter names ethics washing as addressing ethical concerns superficially, to gain public trust, while making no substantive change to practice. It identifies three forms this takes: ethical statements that are vague or go unenforced, ethics boards constituted with limited authority, and adoption of frameworks that carry no accountability mechanism. The chapter offers this as a taxonomy of forms, not as a measurement of how often each occurs.

  • Low AI literacy / verification readinesscore

    The AI-literacy-gap pressure: verification skill (not time) drops and deference rises as trust calibrates on the tool itself. The correction-capacity gauge reads the result. The grounding literature proposes AI literacy as a core professional competency.

    Evidence: The 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers describes a gap between AI exposure and AI preparedness: 26.6% of respondents cited lack of training or understanding of AI technology as a challenge, 53.4% said training on AI tools and effective use would help, and clear guidelines on the ethical use of AI were the most-endorsed need (66.8%).

    Evidence: AI literacy — the knowledge and skills required to understand, use, and critically evaluate AI systems — has been proposed as a core competency for social work, relevant even to practitioners who never directly use AI tools.

  • AI iatrogenics / governance backfireadvanced

    First-class here: purging records without reading them backfires exactly as the sociotechnical simulation found, and the efficiency readout will call a lever stack counterproductive to its face. The quieter trap — oversight whose gains are bought by rising deference — is why the deskilling-arrest lever exists.

    Evidence: In the sociotechnical simulation, deleting records without reading them raised the contaminated share by stripping out benign entries; only content-aware cleanup reliably reduced it. (PAN governance-lever audit)

  • Monitoring failure / drift blindnessadvanced

    Two pressures carry it: the silent vendor update (drift arriving under controls tuned to old behavior) and monitoring going stale (dashboards nobody must act on — the authority gauge hollows while the regime worsens). “Review on schedule” and “Escalate checks” are the counters. Reported accuracy is where this mode hides: a held-out figure and a figure from another site are different quantities, and discrimination statistics are not the rates at which error spreads or gets caught.

    Evidence: How a model scores on data held back from its own training and how it scores at a different site are different quantities. The volume's disability chapter reports a named pair where the second is materially lower than the first. A figure quoted without saying which of the two it is does not tell a reader what the model will do in their setting.

    Evidence: The volume's substance-use chapter reports that nearly all models in the field it reviews are developed and evaluated on a single dataset. Where that holds, a reported ceiling describes that sample rather than a portable capability, and it should be read as the best case observed on one collection of records, not as what the tool will do elsewhere.

    Evidence: Discrimination statistics such as AUROC, precision, F1 and lift describe how a model separates cases on a labelled sample. They are not per-interaction rates at which an error is adopted, written into a record, or corrected. The volume's research chapter treats these as different quantities, and they must never be entered into a diagram as though one substitutes for the other.

  • Service starvation / over-throttlingadvanced

    The sixth iatrogenic, symmetric to deference-load: governance so tight the work stops. The “Work getting done” gauge reads strained and the Net AI benefit track sits on the hurting side while the failure regime reads contained — a breaker or write-gate has zeroed an assistive arm, or checking layers have throttled it to a satisficing region where the tool is compliant but not constructive. It is a Lab-only service quantity, distinct from PAN's harm-reduction: you paid budget, deference, and compute for a tool your governance won't let help.

    Evidence: Safety-only alignment establishes a behavioral floor without a ceiling: systems can be 'not-unsafe' yet directionless — compliant without being constructive — and benefit must be assessed as scaffold versus crutch.

    Evidence: In a two-year child-welfare ethnography, an ill-fitting algorithmic tool imposed ongoing repair work on caseworkers — anticipatorily editing the inputs they supplied so the tool would return a usable result, and bending or working around procedure to reconcile its output with the case in front of them — labor spent making a poorly-suited tool usable rather than on the casework itself, distinct from any deliberate checking of the output.

  • Repair work / tool-usability laboradvanced

    The drag the Work-quality component reads when a contaminated or ill-fitting tool imposes ongoing labor to make its output usable — anticipatory editing of inputs, bending procedure to reconcile the result — spent making the tool usable rather than on the casework itself, distinct from the deliberate checking priced as review latency.

    Evidence: In a two-year child-welfare ethnography, an ill-fitting algorithmic tool imposed ongoing repair work on caseworkers — anticipatorily editing the inputs they supplied so the tool would return a usable result, and bending or working around procedure to reconcile its output with the case in front of them — labor spent making a poorly-suited tool usable rather than on the casework itself, distinct from any deliberate checking of the output.

  • Authority-capacity mismatch / oversight assigned without capacityadvanced

    Adjacent to two core modes and distinct from both: automation bias and low AI literacy are about the checker's skill, while this is about the distance between where authority sits and where capacity sits. A sign-off requirement names who answers for a decision; it does not create the class of people who would have to catch the error. On these diagrams that reads as the authority gauge registering a control the correction-capacity gauge cannot staff, with the deference gauge free to drift underneath it. “Require sign-off” and “Review the riskiest first” place the duty; the training levers the catalogue already ships are what would move the capacity. This entry adds no lever and no pressure — it is a reading rule over gauges the Lab already has.

    Evidence: Two chapters of the volume describe the same gap from opposite ends. The governance chapter names an ethical capacity gap: many social workers have not been trained in data science or AI oversight, which it argues leaves them ill-prepared to question or interpret the algorithmic outputs they are nonetheless answerable for. The literacy chapter's professional-development framework assigns audiences by tier, placing sanctioned-tool lists, ethics review boards and vendor bias-mitigation terms with agency leadership while the skill to audit a decision and advocate for a misclassified client is taught at the tier below; it also reports professional-body guidance placing the duty to train on the employer rather than on the individual practitioner. Both are arguments about where authority sits relative to capacity. Neither reports a measured rate of either.

  • Purpose creep / function creepadvanced

    One family seen from two ends. A model leaves the scope it was published under, so an instrument built to study population-level risk gets read as a judgment about one person; and a store acquires a class of reader it was not collected for, so a school record kept for support is read for discipline or handed to an outside agency. The nearest surfaces today are honest about their reach: “Vet connections” closes pathways between systems structurally, “Store less data” shrinks what is there to reuse, and the scope-filter mediator ships as a starter default with no cited delta. None of them models a change in WHO may read a store that stays exactly where it is. Nothing here draws a perpetration or lethality predictor, or any risk threshold over students; the mode is crosswalked because a governance diagram blind to a scope change cannot govern one.

    Evidence: The volume's sexual and partner violence chapter draws a scope line around predictive risk modelling and restates it twice: such models are used in research contexts to study population-level risk factors, they are not intended for individual-level decision-making in clinical or legal settings, and they are not intended for practitioners to screen or label individuals directly in real-world settings. The chapter treats the distance between that declared scope and a case-level use as a central governance danger rather than as a modelling defect, and reports no measurement of how often the line is crossed.

    Evidence: The volume's school social work chapter names function creep, a term it takes from Koops (2021), as the long-term risk that data collected for a beneficial purpose such as identifying mental health needs is later repurposed for an entirely different function; it names student discipline and sharing with external law-enforcement agencies as its examples. The chapter's stated concern is the absence of specific, renewed consent for the new use and the erosion of trust that follows, not the volume of data held. It offers this as a risk argument and reports no incidence rate.

    Evidence: Under the GDPR (Regulation (EU) 2016/679, Article 5), personal data must be collected for specified purposes and not further processed in a way incompatible with them (purpose limitation), and kept adequate, relevant, and limited to what is necessary (data minimisation).

  • Environmental burden (external context)external context

    Deliberately outside the network: nothing in this Lab computes environmental cost, and no gauge claims to. Practitioner concern about AI's environmental impact is documented in the survey evidence and belongs in deployment governance as external context.

    Evidence: In the open-ended comments of the 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers, ethical concerns — prominently including the environmental impact of AI infrastructure — were the most common theme, and the report's first recommendation includes environmental impact among the topics profession-wide ethical guidance should address.

  • Coverage exclusion / exclusion error in targetingexternal context

    Deliberately outside the network, and the boundary is the point. When the output is a ranked worklist steering a scarce resource, the people that list does not name enter no pathway drawn here, no gauge reads them, and none should: these diagrams govern how a decision travels once a system has produced one, not who the system was able to consider in the first place. The targeting literature calls the two directions inclusion error and exclusion error, and only the first leaves a record inside the organization to correct against. Where a deployment has measured the second, the figure belongs to whoever measured it and is recorded in that case file — the Los Angeles homelessness-prevention unit's own equity audit is the shipped example. It is never computed from these dynamics, and no gauge here is a fairness metric.

    Evidence: The volume's poverty chapter uses the targeting literature's paired vocabulary for the two directions in which a system steering a scarce resource fails: an inclusion error reaches someone the program did not intend to reach, and an exclusion error leaves out someone it did intend to reach. The chapter reports reducing both as the stated aim of machine-learning-assisted eligibility and proxy-means targeting. It is a narrative review and measures neither rate itself; the accuracy results it summarises belong to its sources.

    Evidence: The Homelessness Prevention Unit's own November 2024 equity audit, on a test population of 47,582 individuals eligible to be scored, reported false-negative rates ranging from about 56% for Black individuals to roughly 63 to 65% for other groups: the model misses a majority of the people who later become homeless, while performing roughly consistently across race, ethnicity, and gender and identifying Black individuals slightly more strongly.

The Lab runs invented cases. Your organization runs on a real one.

We map your actual deployment — its pathways, pressures, and the levers your leadership can pull — through intake, diagnosis, prescription, and monitoring.

Sources & Evidence

Tap to expand

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalIndependent scrutiny of Rotterdam's welfare-fraud risk model — a 2021 municipal audit followed by a 2023 journ…

Independent scrutiny of Rotterdam's welfare-fraud risk model — a 2021 municipal audit followed by a 2023 journalistic investigation that obtained the model itself — documented scores skewed against already-vulnerable groups, and the city suspended the system's use.

rekenkamerrotterdam2021GroundingGovernmentSave

Rekenkamer Rotterdam, Gekleurde technologie: onderzoek naar het gebruik van algoritmes door de gemeente Rotterdam (2021) https://www.rekenkamers.nl/rapport/gekleurde-technologie/

https://www.rekenkamers.nl/rapport/gekleurde-technologie/

Appears in: Evidence reverification (2026)

Grounds: deployment audit: Rotterdam welfare-fraud algorithm

wiredlighthousereports2023GroundingInvestigativeSave

WIRED / Lighthouse Reports, Inside the suspicion machine (2023) https://www.wired.com/story/welfare-state-algorithms/

https://www.wired.com/story/welfare-state-algorithms/

Appears in: PAN framework development

Grounds: deployment audit: Rotterdam welfare-fraud algorithm

EmpiricalEvaluation evidence on the Allegheny Family Screening Tool found that screener overrides of the tool's recomme…

Evaluation evidence on the Allegheny Family Screening Tool found that screener overrides of the tool's recommendations reduced racial disparity in screen-in rates relative to the tool alone.

goldhaberfiebertprince2019GroundingGovernment evaluationSave

Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

Appears in: PAN framework development

Grounds: capability governance: at-node control; model org: allegheny_afst

Topics: child-welfare

ScenarioIn the sociotechnical simulation, over a supervised-plus-agent scenario, adding a verifier to the autonomous a…

In the sociotechnical simulation, over a supervised-plus-agent scenario, adding a verifier to the autonomous agent removed roughly 46% of the harm that persists and a coordinated governance package roughly 43%, while upgrading the model alone removed only about 6%.

From the sociotechnical simulation: PAN social-work governance guidance, lever-ranking comparison.

ScenarioIn the sociotechnical simulation, fixing the surrounding system out-leveraged an equal-effort model upgrade in…

In the sociotechnical simulation, fixing the surrounding system out-leveraged an equal-effort model upgrade in nearly every case tested, and by several times the margin - a better model helps least where the system, not the model, does the damage.

From the sociotechnical simulation: PAN baseline analysis.

ScenarioIn the sociotechnical simulation, deleting records without reading them raised the contaminated share by strip…

In the sociotechnical simulation, deleting records without reading them raised the contaminated share by stripping out benign entries; only content-aware cleanup reliably reduced it.

From the sociotechnical simulation: PAN governance-lever audit.

ScenarioIn the sociotechnical simulation, the same AI in three modeled office cultures - stylized, not real workplaces…

In the sociotechnical simulation, the same AI in three modeled office cultures - stylized, not real workplaces - let errors stick at very different rates: roughly 75% under low-oversight autonomy, 20% under human supervision, and 16% under high-governance professional controls.

From the sociotechnical simulation: PAN social-work governance guidance, three-office comparison.

EmpiricalIn a national survey of 1,179 U.S.-based social workers conducted from October 2025 to February 2026 by the Un…

In a national survey of 1,179 U.S.-based social workers conducted from October 2025 to February 2026 by the University of Texas at Austin in collaboration with NASW, 63.5% of respondents reported using AI tools or technologies in their current role.

isbanner2022AcademicSave

Isbanner, S., O'Shaughnessy, P., Steel, D., Wilcock, S., & Carter, S. (2022). The Adoption of Artificial Intelligence in Health Care and Social Services in Australia: Findings From a Methodologically Innovative National Survey of Values and Attitudes (the AVA-AI Study). Journal of Medical Internet Research, 24(8), e37611. https://doi.org/10.2196/37611

doi.org/10.2196/37611

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction, public-benefits

EmpiricalIn the 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers, concerns about d…

In the 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers, concerns about data privacy and security were the most frequently reported challenge to using AI in practice (46.5% of respondents), and an increased focus on client privacy and confidentiality was the most requested improvement to AI tools for social work (50.4%).

isbanner2022AcademicSave

Isbanner, S., O'Shaughnessy, P., Steel, D., Wilcock, S., & Carter, S. (2022). The Adoption of Artificial Intelligence in Health Care and Social Services in Australia: Findings From a Methodologically Innovative National Survey of Values and Attitudes (the AVA-AI Study). Journal of Medical Internet Research, 24(8), e37611. https://doi.org/10.2196/37611

doi.org/10.2196/37611

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction, public-benefits

EmpiricalIn the 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers, 40.8% of respond…

In the 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers, 40.8% of respondents reported ethical concerns about relying on AI for decision-making, and overreliance on automated decision-making was among the most frequently cited concerns overall.

pinazohernandis2026AcademicSave

Pinazo-Hernandis, S., & Carcavilla-Gonzalez, N. (2026). Are future social workers ready for AI? Fears, barriers, and learning needs in higher education. Social Work Education. https://doi.org/10.1080/02615479.2026.2631708

doi.org/10.1080/02615479.2026.2631708

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction, social-work

EmpiricalThe 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers describes a gap betw…

The 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers describes a gap between AI exposure and AI preparedness: 26.6% of respondents cited lack of training or understanding of AI technology as a challenge, 53.4% said training on AI tools and effective use would help, and clear guidelines on the ethical use of AI were the most-endorsed need (66.8%).

pinazohernandis2026AcademicSave

Pinazo-Hernandis, S., & Carcavilla-Gonzalez, N. (2026). Are future social workers ready for AI? Fears, barriers, and learning needs in higher education. Social Work Education. https://doi.org/10.1080/02615479.2026.2631708

doi.org/10.1080/02615479.2026.2631708

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction, social-work

EmpiricalIn the 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers, 42.1% of respond…

In the 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers, 42.1% of respondents reported having no role in decision-making about AI adoption in their workplace; the report concludes most respondents have limited or no control over how AI technologies are selected or implemented within their organizations.

EmpiricalIn the open-ended comments of the 2025–2026 University of Texas at Austin / NASW national survey of U.S. socia…

In the open-ended comments of the 2025–2026 University of Texas at Austin / NASW national survey of U.S. social workers, ethical concerns — prominently including the environmental impact of AI infrastructure — were the most common theme, and the report's first recommendation includes environmental impact among the topics profession-wide ethical guidance should address.

massey2026AcademicSave

Massey, M., Williams, I., Polistina, G., & Breaux, E. (2026). Artificial intelligence and environmental justice: A critical review of social work literature. Society for Social Work and Research Annual Conference. https://sswr.confex.com/sswr/2026/webprogram/Paper62664.html

https://sswr.confex.com/sswr/2026/webprogram/Paper62664.html

Appears in: National survey report (2026)

Topics: social-work

ConceptualAI literacy — the knowledge and skills required to understand, use, and critically evaluate AI systems — has b…

AI literacy — the knowledge and skills required to understand, use, and critically evaluate AI systems — has been proposed as a core competency for social work, relevant even to practitioners who never directly use AI tools.

ahn2025AcademicSave

Ahn, E., Choi, M., Fowler, P., & Song, I. H. (2025). Artificial intelligence (AI) literacy for social work: Implications for core competencies. Journal of the Society for Social Work and Research, 16(1), 9-26. https://doi.org/10.1086/735187

doi.org/10.1086/735187

Appears in: National survey report (2026); Paramerge authored research

Topics: social-work

EmpiricalResearch on AI sycophancy describes it as a fragmented construct — a family of distinct agreement-seeking beha…

Research on AI sycophancy describes it as a fragmented construct — a family of distinct agreement-seeking behaviors that share a label but differ in form, mechanism, measurement, and required mitigation — and finds it intensifies under user pushback and across multi-turn interaction.

ye2026AcademicSave

Ye, M., Ibrahim, L., Bo, J. Y., et al. (2026). What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.21778

doi.org/10.48550/arXiv.2605.21778

Appears in: PAN framework development

Topics: ai-safety

sharma2024AcademicSave

Sharma, M., Tong, M., Korbak, T., et al. (2024). Towards Understanding Sycophancy in Language Models. In International Conference on Learning Representations (ICLR 2024). https://doi.org/10.48550/arXiv.2310.13548

doi.org/10.48550/arXiv.2310.13548

Appears in: Evidence reverification (2026)

Topics: ai-safety, human-ai-interaction

ConceptualClaims and behaviors spread through peer networks sideways, along informal ties — diffusion research finds wea…

Claims and behaviors spread through peer networks sideways, along informal ties — diffusion research finds weak ties and small-world clustering carry information and practices across a network far faster than formal reporting lines.

granovetter1973AcademicSave

Granovetter, M. S. (1973). The strength of weak ties. American Journal of Sociology, 78(6), 1360–1380. https://doi.org/10.1086/210318

doi.org/10.1086/210318

Appears in: Paramerge authored research

watts1998AcademicSave

Watts, D. J., & Strogatz, S. H. (1998). Collective dynamics of 'small-world' networks. Nature, 393(6684), 440–442. https://doi.org/10.1038/30918

doi.org/10.1038/30918

Appears in: Paramerge authored research

Topics: complexity-science

EmpiricalA single automated rule set applied uniformly and without human review produced tens of thousands of correlate…

A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.

EmpiricalModel behavior — including misaligned behavior — can propagate through model-to-model channels: research shows…

Model behavior — including misaligned behavior — can propagate through model-to-model channels: research shows narrow in-context examples and inter-model interaction can induce broadly misaligned behavior in the receiving model.

afonin2026AcademicSave

Afonin, N., Andriianov, N., Hovhannisyan, V., Bageshpura, N., Liu, K., Zhu, K., Dev, S., Panda, A., Rogov, O., Tutubalina, E., Panchenko, A., & Seleznyov, M. (2026). Emergent misalignment via in-context learning: Narrow in-context examples can produce broadly misaligned LLMs [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2510.11288

doi.org/10.48550/arXiv.2510.11288

Appears in: Paramerge authored research

Topics: ai-alignment, complexity-science

panpatil2025AcademicSave

Panpatil, S., Dingeto, H., & Park, H. (2025). Eliciting and analyzing emergent misalignment in state-of-the-art large language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2508.04196

doi.org/10.48550/arXiv.2508.04196

Appears in: Paramerge authored research

Topics: ai-alignment, complexity-science

betley2026AcademicSave

Betley, J., Warncke, N., Sztyber-Betley, A., Tan, D., Bao, X., Soto, M., Srivastava, M., Labenz, N., & Evans, O. (2026). Training large language models on narrow tasks can lead to broad misalignment. Nature, 649(8097), 584-589. https://doi.org/10.1038/s41586-025-09937-5

doi.org/10.1038/s41586-025-09937-5

Appears in: Paramerge authored research

Topics: ai-alignment

ConceptualEmerging agentic-AI governance frameworks treat inter-agent interaction as a first-class assurance surface, re…

Emerging agentic-AI governance frameworks treat inter-agent interaction as a first-class assurance surface, requiring explicit oversight of agent-to-agent couplings rather than per-model evaluation alone.

khan2025AcademicSave

Khan, R., Joyce, D., & Habiba, M. (2025). AGENTSAFE: A unified framework for ethical assurance and governance in agentic AI [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2512.03180

doi.org/10.48550/arXiv.2512.03180

Appears in: Paramerge authored research

Topics: ai-governance

hammond2025AcademicSave

Hammond, L., Chan, A., Clifton, J., et al. (2025). Multi-Agent Risks from Advanced AI [Technical Report No. 1]. Cooperative AI Foundation. arXiv. https://doi.org/10.48550/arXiv.2502.14143

doi.org/10.48550/arXiv.2502.14143

Appears in: Evidence reverification (2026)

Topics: ai-governance, ai-safety

EmpiricalDocumented benefit-automation failures replicated determinations into downstream systems with no independent r…

Documented benefit-automation failures replicated determinations into downstream systems with no independent reconciliation against the source records — Michigan MiDAS actioned replicated flags and Robodebt reversed the onus onto recipients.

EmpiricalDocumented risk-scoring deployments computed scores from multi-agency administrative records originally collec…

Documented risk-scoring deployments computed scores from multi-agency administrative records originally collected for other purposes, which is the data-protection critique recorded in independent reviews of these systems.

goldhaberfiebertprince2019GroundingGovernment evaluationSave

Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

Appears in: PAN framework development

Grounds: capability governance: at-node control; model org: allegheny_afst

Topics: child-welfare

EmpiricalDocumented enforcement systems actioned replicated flags automatically — garnishment and penalties applied bef…

Documented enforcement systems actioned replicated flags automatically — garnishment and penalties applied before any human review step in the recorded MiDAS deployment.

EmpiricalThe Robodebt Royal Commission documented debts raised from income-averaged derived inputs with the onus placed…

The Robodebt Royal Commission documented debts raised from income-averaged derived inputs with the onus placed on recipients to disprove the automated assessments.

EmpiricalProfessional-verification cultures documented in social work practice sustain peer checking of AI output rathe…

Professional-verification cultures documented in social work practice sustain peer checking of AI output rather than unquestioned acceptance.

baez2026AcademicSave

Báez, J. C., Ahn, E., Tamietti, A., Victor, B. G., & Goldkind, L. (2026). Clinical social workers’ perceptions of large language models in practice: Resistance to automation and prospects for integration. Journal of Evidence-Based Social Work, 23(1), 42–63. https://doi.org/10.1080/26408066.2025.2542450

doi.org/10.1080/26408066.2025.2542450

Appears in: National survey report (2026)

Topics: social-work

EmpiricalClinical assessors bound by algorithmic allocation with limited override capacity form a documented constraine…

Clinical assessors bound by algorithmic allocation with limited override capacity form a documented constrained-judgment pattern in home-care assessment.

sutton2020AcademicSave

Sutton, R. T., Pincock, D., Baumgart, D. C., Sadowski, D. C., Fedorak, R. N., & Kroeker, K. I. (2020). An overview of clinical decision support systems: Benefits, risks, and strategies for success. NPJ Digital Medicine, 3(1), 17. https://doi.org/10.1038/s41746-020-0221-y

doi.org/10.1038/s41746-020-0221-y

Appears in: Paramerge authored research

upturnGroundingAdvocacySave

Upturn, Calculated Need: automated home-care hour allocation https://www.upturn.org/work/calculated-need/

https://www.upturn.org/work/calculated-need/

Appears in: PAN framework development

Grounds: deployment audit: Arkansas ARChoices / Idaho Medicaid

EmpiricalRetrieval layers propagate rather than sanitize their inputs: studies find retrieval-augmented systems remain …

Retrieval layers propagate rather than sanitize their inputs: studies find retrieval-augmented systems remain unfaithful even when the retrieved passage is correct, so faithfulness is bounded rather than assured.

faithfulragGroundingPreprintSave

FaithfulRAG (arXiv:2506.08938) — RAG systems struggle in knowledge-conflict scenarios even when relevant passages are retrieved (pessimistic end). https://arxiv.org/abs/2506.08938

https://arxiv.org/abs/2506.08938

Grounds: empirical cap: groundtruth_reliability (max)

faithfulragwithsparseautoencGroundingPreprintSave

Faithful RAG with Sparse Autoencoders (arXiv:2512.08892) — even with relevant passages retrieved, models contradict evidence / invent details; faithfulness is not guaranteed. https://arxiv.org/abs/2512.08892

https://arxiv.org/abs/2512.08892

Grounds: empirical cap: groundtruth_reliability (max)

ragevaluationsurveyGroundingPreprintSave

RAG evaluation survey (arXiv:2405.07437) — factuality evaluation is bounded by knowledge-base coverage and retrieval accuracy; what is checkable depends on what is documented. https://arxiv.org/abs/2405.07437

https://arxiv.org/abs/2405.07437

Grounds: empirical cap: frac_verifiable (max)

EmpiricalRestricting retrieval to a curated, vetted document set bounds what re-enters the model: retrieval-augmented s…

Restricting retrieval to a curated, vetted document set bounds what re-enters the model: retrieval-augmented systems fact-checking against a curated peer-reviewed corpus reach roughly 0.97+ accuracy and factuality evaluation is limited by knowledge-base coverage — what is checkable depends on what is documented — so a vetted corpus reduces contamination drawn back into the model relative to open retrieval, though faithfulness remains imperfect under knowledge conflict.

retrievalaugmentedcovidfactcGroundingPeer-reviewedSave

Retrieval-augmented COVID-19 fact-checking (PMC12079058) — CRAG/Self-RAG reach 0.972-0.978 accuracy against a curated 130k peer-reviewed corpus (optimistic ceiling). https://pmc.ncbi.nlm.nih.gov/articles/PMC12079058/

https://pmc.ncbi.nlm.nih.gov/articles/PMC12079058/

Grounds: empirical cap: groundtruth_reliability (max)

ragevaluationsurveyGroundingPreprintSave

RAG evaluation survey (arXiv:2405.07437) — factuality evaluation is bounded by knowledge-base coverage and retrieval accuracy; what is checkable depends on what is documented. https://arxiv.org/abs/2405.07437

https://arxiv.org/abs/2405.07437

Grounds: empirical cap: frac_verifiable (max)

faithfulragGroundingPreprintSave

FaithfulRAG (arXiv:2506.08938) — RAG systems struggle in knowledge-conflict scenarios even when relevant passages are retrieved (pessimistic end). https://arxiv.org/abs/2506.08938

https://arxiv.org/abs/2506.08938

Grounds: empirical cap: groundtruth_reliability (max)

EmpiricalAutomated output checks are partial, not complete — measured detector-accuracy bands sit well below completene…

Automated output checks are partial, not complete — measured detector-accuracy bands sit well below completeness, especially on hard or adversarial content.

theillusionofprogressGroundingPreprintSave

'The Illusion of Progress' (arXiv:2508.08285) — LLM-as-Judge Precision 0.736 / Recall 0.957 / F1 0.832 vs human consensus on QA. https://arxiv.org/abs/2508.08285

https://arxiv.org/abs/2508.08285

Grounds: empirical cap: catch_at_generation (max)

halogenGroundingPeer-reviewedSave

HALoGEN (arXiv:2501.08292) — best models hallucinate 4%-86% of generated facts depending on domain. https://arxiv.org/abs/2501.08292

https://arxiv.org/abs/2501.08292

Grounds: empirical cap: model_error_base (min)

datadogllmasajudge2025GroundingIndustrySave

Datadog LLM-as-a-judge (2025) — detection F1 drops substantially from HaluBench to the harder RAGTruth; harder hallucinations are harder to catch.

Grounds: empirical cap: catch_at_generation (max)

mentalhealthchatbotdetectionGroundingPreprintSave

Mental-health chatbot detection (arXiv:2604.06216) — GPT judges 54.6% accuracy, 9.3% recall (miss 90.7% of hallucinations); traditional methods F1<0.30 on subjective content. https://arxiv.org/abs/2604.06216

https://arxiv.org/abs/2604.06216

Grounds: empirical cap: catch_at_generation (max)

EmpiricalRanked risk lists steered which cases were investigated in documented deployments; the anchoring direction is …

Ranked risk lists steered which cases were investigated in documented deployments; the anchoring direction is documented while its magnitude is not published.

wiredlighthousereports2023GroundingInvestigativeSave

WIRED / Lighthouse Reports, Inside the suspicion machine (2023) https://www.wired.com/story/welfare-state-algorithms/

https://www.wired.com/story/welfare-state-algorithms/

Appears in: PAN framework development

Grounds: deployment audit: Rotterdam welfare-fraud algorithm

amnestyinternational2021bGroundingAdvocacySave

Amnesty International, Xenophobic machines: Discrimination through unregulated use of algorithms in the Dutch childcare benefits scandal (2021) https://www.amnesty.org/en/documents/eur35/4686/2021/en/

https://www.amnesty.org/en/documents/eur35/4686/2021/en/

Appears in: PAN framework development

Grounds: deployment audit: SyRI / childcare-benefits (toeslagenaffaire); model org: netherlands_toeslagen

goldhaberfiebertprince2019GroundingGovernment evaluationSave

Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

Appears in: PAN framework development

Grounds: capability governance: at-node control; model org: allegheny_afst

Topics: child-welfare

EmpiricalModel error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors ru…

Model error has a hard nonzero floor: formal impossibility results rule out zero error, and measured floors run roughly 1.6–11.6% in frontier evaluations and 4–86% across domains.

xuetal2024GroundingPreprintSave

Xu et al. (2024), 'Hallucination is Inevitable: An Innate Limitation of LLMs' — formal proof that hallucination cannot be eliminated.

Grounds: empirical cap: model_error_base (min)

karpowicz2025GroundingPreprintSave

Karpowicz (2025) — three independent mathematical frameworks (auction theory, proper scoring, log-sum-exp) all conclude no LLM inference mechanism can be simultaneously truthful, etc.

Grounds: empirical cap: model_error_base (min)

halogenGroundingPeer-reviewedSave

HALoGEN (arXiv:2501.08292) — best models hallucinate 4%-86% of generated facts depending on domain. https://arxiv.org/abs/2501.08292

https://arxiv.org/abs/2501.08292

Grounds: empirical cap: model_error_base (min)

openai2025GroundingFrontier labSave

OpenAI (2025), 'Why Language Models Hallucinate' — next-token training plus IDK-penalizing benchmarks push models to bluff; explains the persistent nonzero floor.

Grounds: empirical cap: model_error_base (min)

llmstats2026GroundingIndustry evaluationSave

llm-stats.com failure-focused eval (2026) — FactsGrounding 89.1% accuracy => ~10.9% failure on a relatively easy grounded benchmark.

Grounds: empirical cap: model_error_base (min)

suprmindbenchmarkdigest2026GroundingIndustry evaluationSave

Suprmind benchmark digest (2026) — production ChatGPT ~4.8% major-incorrect with reasoning vs ~11.6% without; HealthBench 3.6%->1.6% with GPT-5 thinking.

Grounds: empirical cap: model_error_base (min)

EmpiricalAutomated catch fractions cap out below completeness — around 84% balanced accuracy in optimistic settings ver…

Automated catch fractions cap out below completeness — around 84% balanced accuracy in optimistic settings versus about 55% on hard content and 9.3% recall in worst-case measurements.

faithfulragleaderboardGroundingPreprintSave

Faithful RAG leaderboard (arXiv:2505.04847) — FaithJudge with o3-mini-high reaches ~84% balanced accuracy / ~82% F1 on FaithBench (optimistic ceiling). https://arxiv.org/abs/2505.04847

https://arxiv.org/abs/2505.04847

Grounds: empirical cap: catch_at_generation (max)

theillusionofprogressGroundingPreprintSave

'The Illusion of Progress' (arXiv:2508.08285) — LLM-as-Judge Precision 0.736 / Recall 0.957 / F1 0.832 vs human consensus on QA. https://arxiv.org/abs/2508.08285

https://arxiv.org/abs/2508.08285

Grounds: empirical cap: catch_at_generation (max)

mentalhealthchatbotdetectionGroundingPreprintSave

Mental-health chatbot detection (arXiv:2604.06216) — GPT judges 54.6% accuracy, 9.3% recall (miss 90.7% of hallucinations); traditional methods F1<0.30 on subjective content. https://arxiv.org/abs/2604.06216

https://arxiv.org/abs/2604.06216

Grounds: empirical cap: catch_at_generation (max)

datadogllmasajudge2025GroundingIndustrySave

Datadog LLM-as-a-judge (2025) — detection F1 drops substantially from HaluBench to the harder RAGTruth; harder hallucinations are harder to catch.

Grounds: empirical cap: catch_at_generation (max)

samedetectionaccuracyliteratGroundingPeer-reviewedSave

Same detection-accuracy literature as catch_at_generation (FaithBench arXiv:2410.13210; arXiv:2508.08285); audit-time detection is bounded by the same hallucination-detection ceiling.

Grounds: empirical cap: catch_at_generation (max); empirical cap: decontaminate (max)

EmpiricalRecord audit-and-correct shares the detection-ceiling family: an optimistic anchor near 96% token accuracy fal…

Record audit-and-correct shares the detection-ceiling family: an optimistic anchor near 96% token accuracy falls away on hard content, so decontamination is bounded rather than total.

halludetectlegaldomainGroundingPreprintSave

HalluDetect legal-domain (arXiv:2509.11619) — best mitigation architecture reaches ~96% token accuracy in a FAVORABLE, retrieval-grounded legal setting (optimistic end). https://arxiv.org/abs/2509.11619

https://arxiv.org/abs/2509.11619

Grounds: empirical cap: decontaminate (max)

samedetectionaccuracyliteratGroundingPeer-reviewedSave

Same detection-accuracy literature as catch_at_generation (FaithBench arXiv:2410.13210; arXiv:2508.08285); audit-time detection is bounded by the same hallucination-detection ceiling.

Grounds: empirical cap: catch_at_generation (max); empirical cap: decontaminate (max)

theillusionofprogressGroundingPreprintSave

'The Illusion of Progress' (arXiv:2508.08285) — LLM-as-Judge Precision 0.736 / Recall 0.957 / F1 0.832 vs human consensus on QA. https://arxiv.org/abs/2508.08285

https://arxiv.org/abs/2508.08285

Grounds: empirical cap: catch_at_generation (max)

EmpiricalThe verification channel itself is bounded — curated-corpus fact-checking tops out around 0.972–0.978 reliabil…

The verification channel itself is bounded — curated-corpus fact-checking tops out around 0.972–0.978 reliability and collapses under knowledge conflict.

retrievalaugmentedcovidfactcGroundingPeer-reviewedSave

Retrieval-augmented COVID-19 fact-checking (PMC12079058) — CRAG/Self-RAG reach 0.972-0.978 accuracy against a curated 130k peer-reviewed corpus (optimistic ceiling). https://pmc.ncbi.nlm.nih.gov/articles/PMC12079058/

https://pmc.ncbi.nlm.nih.gov/articles/PMC12079058/

Grounds: empirical cap: groundtruth_reliability (max)

faithfulragGroundingPreprintSave

FaithfulRAG (arXiv:2506.08938) — RAG systems struggle in knowledge-conflict scenarios even when relevant passages are retrieved (pessimistic end). https://arxiv.org/abs/2506.08938

https://arxiv.org/abs/2506.08938

Grounds: empirical cap: groundtruth_reliability (max)

faithfulragwithsparseautoencGroundingPreprintSave

Faithful RAG with Sparse Autoencoders (arXiv:2512.08892) — even with relevant passages retrieved, models contradict evidence / invent details; faithfulness is not guaranteed. https://arxiv.org/abs/2512.08892

https://arxiv.org/abs/2512.08892

Grounds: empirical cap: groundtruth_reliability (max)

AssumptionThe verifiable fraction of contaminated records is a planning range (0.90/0.60/0.30) that is explicitly calibr…

The verifiable fraction of contaminated records is a planning range (0.90/0.60/0.30) that is explicitly calibration-required and has never been measured.

ragevaluationsurveyGroundingPreprintSave

RAG evaluation survey (arXiv:2405.07437) — factuality evaluation is bounded by knowledge-base coverage and retrieval accuracy; what is checkable depends on what is documented. https://arxiv.org/abs/2405.07437

https://arxiv.org/abs/2405.07437

Grounds: empirical cap: frac_verifiable (max)

EmpiricalIn the documented MiDAS case, error among no-review auto-adjudications ran roughly 93%, and determinations err…

In the documented MiDAS case, error among no-review auto-adjudications ran roughly 93%, and determinations erred at about 85% without human review versus 44% with it.

aiincidentdatabaseGroundingInvestigativeSave

AI Incident Database, Incident 373 (MiDAS false fraud claims) https://incidentdatabase.ai/cite/373/

https://incidentdatabase.ai/cite/373/

Grounds: model org: michigan_midas

EmpiricalIn the documented AFST evaluation, screener overrides of the tool — roughly a third of its recommendations — c…

In the documented AFST evaluation, screener overrides of the tool — roughly a third of its recommendations — cut screen-in disparity from about 20% to 9% relative to the tool acting alone.

goldhaberfiebertprince2019GroundingGovernment evaluationSave

Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

Appears in: PAN framework development

Grounds: capability governance: at-node control; model org: allegheny_afst

Topics: child-welfare

stapletonetal2022AcademicPeer-reviewedSave

Stapleton, L., Lee, M. H., Qing, D., Wright, M., Chouldechova, A., Holstein, K., Wu, Z. S., & Zhu, H. (2022). Imagining new futures beyond predictive systems in child welfare: A qualitative study with impacted stakeholders. 2022 ACM Conference on Fairness Accountability and Transparency, 1162–1177. https://doi.org/10.1145/3531146.3533177

doi.org/10.1145/3531146.3533177

Appears in: PAN framework development; Paramerge authored research

Grounds: deployment audit: Allegheny AFST

Topics: algorithmic-fairness, child-welfare

EmpiricalProfessional caseload standards published by the Child Welfare League of America recommend no more than about …

Professional caseload standards published by the Child Welfare League of America recommend no more than about 15 families per worker, sitting well below documented practice loads.

childrenandfamilyresearchcen2002GroundingAcademicSave

Children and Family Research Center (University of Illinois at Urbana-Champaign), Caseload Size in Best Practice: A Literature Review (2002) https://cfrc.illinois.edu/pubs/bf_20021101_CaseloadSizeInBestPractice.pdf

https://cfrc.illinois.edu/pubs/bf_20021101_CaseloadSizeInBestPractice.pdf

Appears in: Evidence reverification (2026)

Grounds: workforce data: caseload standards

academyforprofessionalexcell2021GroundingAcademicSave

Academy for Professional Excellence / CWDS (San Diego State University), Research Summary: Caseload Standards and Weighting Methodologies (2021) https://theacademy.sdsu.edu/wp-content/uploads/2021/10/CWDS-Research-Summary_Caseload-Standards-and-Weighting.pdf

https://theacademy.sdsu.edu/wp-content/uploads/2021/10/CWDS-Research-Summary_Caseload-Standards-and-Weighting.pdf

Appears in: Evidence reverification (2026)

Grounds: workforce data: caseload standards

EmpiricalDocumentation and administrative tasks consume roughly half of practitioner time: a nationally representative …

Documentation and administrative tasks consume roughly half of practitioner time: a nationally representative US child-welfare workforce snapshot found caseworkers spend about 54% of the workday (4.3 of 8 hours) on paperwork and documentation, and a UK children's-services review reports staff spending over 50% of their time on case recording, paperwork, and related tasks.

opre2025GroundingGovernmentSave

OPRE, Snapshot of the Child Welfare Workforce from 2021 to 2022: Caseworker Experiences Working in the Child Welfare System, OPRE Report 2025-040 (2025) https://acf.gov/opre/report/snapshot-child-welfare-workforce-2021-2022-caseworker-experiences-working-child-welfare

https://acf.gov/opre/report/snapshot-child-welfare-workforce-2021-2022-caseworker-experiences-working-child-welfare

Appears in: PAN framework development

Grounds: workforce data: documentation time share

Topics: child-welfare

burbidge2022GroundingGovernmentSave

Burbidge, I. (2022). Report sets out new blueprint for councils to deliver a reshaped children's services. County Councils Network https://www.countycouncilsnetwork.org.uk/report-sets-out-new-blueprint-for-councils-to-deliver-a-reshaped-childrens-services/

https://www.countycouncilsnetwork.org.uk/report-sets-out-new-blueprint-for-councils-to-deliver-a-reshaped-childrens-services/

Appears in: PAN framework development

Grounds: workforce data: documentation time share

Topics: complexity-science

EmpiricalTiered HIPAA penalties run from $145 to $73,011 per violation with an annual cap near $2.19M (2025-adjusted), …

Tiered HIPAA penalties run from $145 to $73,011 per violation with an annual cap near $2.19M (2025-adjusted), and disclosure to a tool that is not a business associate is itself a violation.

hipaajournal2026bAcademicSave

HIPAA Journal. (2026). HIPAA violation penalties. HIPAA Journal.

Appears in: Paramerge authored research

Topics: privacy-security

ConceptualProtected health information (PHI) is individually identifiable health information; under the HIPAA Privacy Ru…

Protected health information (PHI) is individually identifiable health information; under the HIPAA Privacy Rule (45 CFR 160.103) its uses and its disclosures are both regulated, so how PHI is used inside a system — not only whether it leaves — is governed.

u2013GroundingGovernmentSave

U.S. Department of Health and Human Services (2013). HIPAA Privacy Rule — protected health information, 45 CFR 160.103 (Standards for Privacy of Individually Identifiable Health Information; Omnibus Final Rule). https://www.ecfr.gov/current/title-45/section-160.103

https://www.ecfr.gov/current/title-45/section-160.103

Appears in: Evidence addition (2026)

Grounds: privacy law: PHI and the HIPAA Privacy Rule

Topics: privacy-security

ConceptualPersonally identifiable information (PII) is information that can distinguish or trace an individual's identit…

Personally identifiable information (PII) is information that can distinguish or trace an individual's identity, alone or combined with other data; NIST SP 800-122 directs organizations to minimize its collection and use and to limit use to the purpose for which it was collected.

mccallister2010GroundingGovernmentSave

McCallister, E., Grance, T., & Scarfone, K. (2010). Guide to Protecting the Confidentiality of Personally Identifiable Information (PII). National Institute of Standards and Technology (NIST Special Publication 800-122). https://doi.org/10.6028/NIST.SP.800-122

doi.org/10.6028/NIST.SP.800-122

Appears in: Evidence addition (2026)

Grounds: privacy law: personally identifiable information (PII)

ConceptualUnder the GDPR (Regulation (EU) 2016/679, Article 5), personal data must be collected for specified purposes a…

Under the GDPR (Regulation (EU) 2016/679, Article 5), personal data must be collected for specified purposes and not further processed in a way incompatible with them (purpose limitation), and kept adequate, relevant, and limited to what is necessary (data minimisation).

europeanparliamentandcouncil2016GroundingRegulatorySave

European Parliament and Council of the European Union (2016). Regulation (EU) 2016/679 (General Data Protection Regulation), Article 5 — purpose limitation and data minimisation. https://eur-lex.europa.eu/eli/reg/2016/679/oj

https://eur-lex.europa.eu/eli/reg/2016/679/oj

Appears in: Evidence addition (2026)

Grounds: privacy law: purpose limitation and data minimisation (GDPR)

EmpiricalCited per-unit intensities of roughly 0.3 Wh per inference call and about 3.14 L of water per kWh are applied …

Cited per-unit intensities of roughly 0.3 Wh per inference call and about 3.14 L of water per kWh are applied to authored illustrative volumes rather than to measured deployment totals.

jegham2025AcademicSave

Jegham, N., et al. (2025). How Hungry is AI? Benchmarking energy, water, and carbon footprint of LLM inference [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2505.09598

doi.org/10.48550/arXiv.2505.09598

Appears in: PAN framework development

Topics: ai-governance

li2023AcademicSave

Li, P., Yang, J., Islam, M. A., & Ren, S. (2023). Making AI Less Thirsty: Uncovering and addressing the secret water footprint of AI models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2304.03271

doi.org/10.48550/arXiv.2304.03271

Appears in: PAN framework development

Topics: ai-governance

EmpiricalA preprint benchmark reports an in-context misalignment dose-response: in the most susceptible frontier model,…

A preprint benchmark reports an in-context misalignment dose-response: in the most susceptible frontier model, up to ~24% misaligned behavior at 16 examples rising to ~58% at 256 examples (rates at 16 examples span roughly 1–24% across models), with the majority of misaligned responses rationalized.

afonin2026AcademicSave

Afonin, N., Andriianov, N., Hovhannisyan, V., Bageshpura, N., Liu, K., Zhu, K., Dev, S., Panda, A., Rogov, O., Tutubalina, E., Panchenko, A., & Seleznyov, M. (2026). Emergent misalignment via in-context learning: Narrow in-context examples can produce broadly misaligned LLMs [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2510.11288

doi.org/10.48550/arXiv.2510.11288

Appears in: Paramerge authored research

Topics: ai-alignment, complexity-science

EmpiricalModel behavior drifts discontinuously between evaluation snapshots, and narrow finetuning can induce broad cor…

Model behavior drifts discontinuously between evaluation snapshots, and narrow finetuning can induce broad correlated failure across unrelated tasks.

betley2026AcademicSave

Betley, J., Warncke, N., Sztyber-Betley, A., Tan, D., Bao, X., Soto, M., Srivastava, M., Labenz, N., & Evans, O. (2026). Training large language models on narrow tasks can lead to broad misalignment. Nature, 649(8097), 584-589. https://doi.org/10.1038/s41586-025-09937-5

doi.org/10.1038/s41586-025-09937-5

Appears in: Paramerge authored research

Topics: ai-alignment

li2026AcademicSave

Li, Z., Fan, C., & Zhou, T. (2026). Grokking in LLM pretraining? Monitor memorization-to-generalization without test [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2506.21551

doi.org/10.48550/arXiv.2506.21551

Appears in: Paramerge authored research

song2026AcademicSave

Song, P., Han, P., & Goodman, N. (2026). Large language model reasoning failures [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2602.06176

doi.org/10.48550/arXiv.2602.06176

Appears in: Paramerge authored research

anwar2024AcademicSave

Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., Lubana, E., Jenner, E., Casper, S., Sourbut, O., Edelman, B. L., Zhang, Z., Gunther, M., Korinek, A., Hernandez-Orallo, J., Hammond, L., Bigelow, E., Pan, A., Langosco, L., Korbak, T., Zhang, H., Zhong, R., O Heigeartaigh, S., Recchia, G., Corsi, G., Chan, A., Anderljung, M., Edwards, L., Petrov, A., de Witt, C. S., Motwani, S. R., Bengio, Y., Chen, D., Torr, P. H. S., Albanie, S., Maharaj, T., Foerster, J., Tramer, F., He, H., Kasirzadeh, A., Choi, Y., & Krueger, D. (2024). Foundational challenges in assuring alignment and safety of large language models. Transactions on Machine Learning Research. https://doi.org/10.48550/arXiv.2404.09932

doi.org/10.48550/arXiv.2404.09932

Appears in: Paramerge authored research

Topics: ai-alignment, ai-safety

nikolaou2025AcademicSave

Nikolaou, K., Krippendorf, S., Tovey, S., & Holm, C. (2025). Beyond scaling curves: Internal dynamics of neural networks through the NTK lens [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.05035

doi.org/10.48550/arXiv.2507.05035

Appears in: Paramerge authored research

Topics: complexity-science

EmpiricalIn contextual inquiries with Allegheny AFST call screeners, workers calibrated reliance using contextual case …

In contextual inquiries with Allegheny AFST call screeners, workers calibrated reliance using contextual case knowledge unavailable to the model and reliably detected and overrode erroneous risk scores — complementary human information, not generic distrust, was the safeguard's mechanism.

kawakami2022AcademicSave

Kawakami, A., Sivaraman, V., Cheng, H.-F., Stapleton, L., Cheng, Y., Qing, D., Perer, A., Wu, Z. S., Zhu, H., & Holstein, K. (2022). Improving Human-AI Partnerships in Child Welfare: Understanding Worker Practices, Challenges, and Desires for Algorithmic Decision Support. In CHI Conference on Human Factors in Computing Systems (CHI '22). ACM. https://doi.org/10.1145/3491102.3517439

doi.org/10.1145/3491102.3517439

Appears in: PAN framework development

Topics: algorithmic-fairness, child-welfare, human-ai-interaction

dearteaga2020AcademicSave

De-Arteaga, M., Fogliato, R., & Chouldechova, A. (2020). A Case for Humans-in-the-Loop: Decisions in the Presence of Erroneous Algorithmic Scores. In CHI Conference on Human Factors in Computing Systems (CHI 2020). ACM. https://doi.org/10.1145/3313831.3376638

doi.org/10.1145/3313831.3376638

Appears in: Evidence reverification (2026)

Topics: algorithmic-fairness, child-welfare, human-ai-interaction

EmpiricalAFST workers reported sometimes agreeing with the risk score against their own best judgment under override-ra…

AFST workers reported sometimes agreeing with the risk score against their own best judgment under override-rate oversight, and becoming less likely to disagree over time — reliance driven by organizational incentives independent of trust in the tool.

kawakami2022AcademicSave

Kawakami, A., Sivaraman, V., Cheng, H.-F., Stapleton, L., Cheng, Y., Qing, D., Perer, A., Wu, Z. S., Zhu, H., & Holstein, K. (2022). Improving Human-AI Partnerships in Child Welfare: Understanding Worker Practices, Challenges, and Desires for Algorithmic Decision Support. In CHI Conference on Human Factors in Computing Systems (CHI '22). ACM. https://doi.org/10.1145/3491102.3517439

doi.org/10.1145/3491102.3517439

Appears in: PAN framework development

Topics: algorithmic-fairness, child-welfare, human-ai-interaction

kawakami2026AcademicSave

Kawakami, A., Taylor, J., Fox, S., Zhu, H., & Holstein, K. (2026). AI failure loops in devalued work: The confluence of overconfidence in AI and underconfidence in worker expertise. Big Data & Society. https://doi.org/10.1177/20539517261424164

doi.org/10.1177/20539517261424164

Appears in: Evidence reverification (2026)

Topics: child-welfare, human-ai-interaction

EmpiricalIn a two-year child-welfare ethnography, a re-purposed assessment algorithm produced process-oriented harms to…

In a two-year child-welfare ethnography, a re-purposed assessment algorithm produced process-oriented harms to practice, organization, and street-level decisions, compelling caseworkers to perform added repair work; 80% of interviewees reported that the tool had stripped their decision-making discretion.

saxena2024AcademicSave

Saxena, D., & Guha, S. (2024). Algorithmic Harms in Child Welfare: Uncertainties in Practice, Organization, and Street-level Decision-making. ACM Journal on Responsible Computing, 1(1), 1–32. https://doi.org/10.1145/3616473

doi.org/10.1145/3616473

Appears in: Paramerge authored research

Topics: algorithmic-fairness, child-welfare

ammitzbollflugge2021AcademicSave

Ammitzboll Flugge, A., Hildebrandt, T., & Holten Moller, N. (2021). Street-Level Algorithms and AI in Bureaucratic Decision-Making: A Caseworker Perspective. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 40. https://doi.org/10.1145/3449114

doi.org/10.1145/3449114

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction, public-benefits

EmpiricalThe same agency's theory-driven 7ei tool — which tracks case trajectories instead of predicting outcomes — ear…

The same agency's theory-driven 7ei tool — which tracks case trajectories instead of predicting outcomes — earned collective buy-in and better engagement, but required sustained investments: trauma-informed training, specialized supervision and expert consultation, and new collaborative staffings.

saxena2024AcademicSave

Saxena, D., & Guha, S. (2024). Algorithmic Harms in Child Welfare: Uncertainties in Practice, Organization, and Street-level Decision-making. ACM Journal on Responsible Computing, 1(1), 1–32. https://doi.org/10.1145/3616473

doi.org/10.1145/3616473

Appears in: Paramerge authored research

Topics: algorithmic-fairness, child-welfare

EmpiricalIn a participatory-design study (CHI Late-Breaking Work) with 51 social-service practitioners across two stage…

In a participatory-design study (CHI Late-Breaking Work) with 51 social-service practitioners across two stages (27 in co-design workshops, 24 in contextual inquiry), AI value concentrated in documentation relief, assessment brainstorming, guidance for junior workers, and supervision support — with deskilling and privacy concerns voiced inside the same sessions.

tan2025AcademicSave

Tan, Y., Soh, K. X., Zhang, R., Lee, J., Meng, H., Sen, B., & Lee, Y.-C. (2025). Empowering Social Service with AI: Insights from a Participatory Design Study with Practitioners. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA '25). ACM. https://doi.org/10.1145/3706599.3719736

doi.org/10.1145/3706599.3719736

Appears in: PAN framework development

Topics: co-design, human-ai-interaction, social-work

ConceptualSafety-only alignment establishes a behavioral floor without a ceiling: systems can be 'not-unsafe' yet direct…

Safety-only alignment establishes a behavioral floor without a ceiling: systems can be 'not-unsafe' yet directionless — compliant without being constructive — and benefit must be assessed as scaffold versus crutch.

laukkonen2026AcademicSave

Laukkonen, R., Krier, S., Bakalar, C., Chandaria, S., Kringelbach, M., Elwood, A., Ford, D., Rosas, F., Bohacek, M., Franklin, M., Tomašev, N., Chan, S., Rieser, V., Patel, R., Levin, M., & Rao, A. (2026). Positive Alignment: Artificial Intelligence for Human Flourishing [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.10310

doi.org/10.48550/arXiv.2605.10310

Appears in: PAN framework development

Topics: ai-alignment, ai-governance, ai-safety

ConceptualFormally, estimation error shrinks with data while the human perception gap that produces stationary-environme…

Formally, estimation error shrinks with data while the human perception gap that produces stationary-environment black swans has a non-zero lower bound — so an incident-free operating history yields confidence without safety.

lee2025AcademicSave

Lee, H., Park, C., Abel, D., & Jin, M. (2025). A Black Swan Hypothesis: The Role of Human Irrationality in AI Safety. In International Conference on Learning Representations (ICLR 2025). https://doi.org/10.48550/arXiv.2407.18422

doi.org/10.48550/arXiv.2407.18422

Appears in: PAN framework development

Topics: ai-safety, antifragility

ConceptualStatic robustness certification lags emergent threats; organizations that fold each stressor into their model …

Static robustness certification lags emergent threats; organizations that fold each stressor into their model (slow-loop updates, periodic reviews, post-deployment feedback) shrink future risk, while patch-and-pray accumulates it — the fragility trap.

jin2025AcademicSave

Jin, M., & Lee, H. (2025). Position: AI safety must embrace an antifragile perspective [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2509.13339

doi.org/10.48550/arXiv.2509.13339

Appears in: Paramerge authored research

Topics: ai-safety, antifragility

EmpiricalIn formal simulation, even ideal Bayesian users spiral to near-certain false beliefs under a sycophantic inter…

In formal simulation, even ideal Bayesian users spiral to near-certain false beliefs under a sycophantic interlocutor at sycophancy rates measured in frontier models (~50-70%), and truth-constrained cherry-picking still produces spirals — minimizing hallucination alone is insufficient.

chandra2026AcademicSave

Chandra, K., Kleiman-Weiner, M., Ragan-Kelley, J., & Tenenbaum, J. B. (2026). Sycophantic chatbots cause delusional spiraling, even in ideal Bayesians [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2602.19141

doi.org/10.48550/arXiv.2602.19141

Appears in: Paramerge authored research

Topics: ai-safety

sharma2024AcademicSave

Sharma, M., Tong, M., Korbak, T., et al. (2024). Towards Understanding Sycophancy in Language Models. In International Conference on Learning Representations (ICLR 2024). https://doi.org/10.48550/arXiv.2310.13548

doi.org/10.48550/arXiv.2310.13548

Appears in: Evidence reverification (2026)

Topics: ai-safety, human-ai-interaction

EmpiricalIn a four-week randomized study (n=981), voluntary daily chatbot usage duration predicted worse outcomes on lo…

In a four-week randomized study (n=981), voluntary daily chatbot usage duration predicted worse outcomes on loneliness, socialization, emotional dependence, and problematic use across all conditions, and task-style use fostered practical dependence — reduced confidence in independent judgment.

fang2025AcademicSave

Fang, C. M., Liu, A. R., Danry, V., Lee, E., Chan, S. W. T., Pataranutaporn, P., Maes, P., Phang, J., Lampe, M., Ahmad, L., & Agarwal, S. (2025). How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2503.17473

doi.org/10.48550/arXiv.2503.17473

Appears in: PAN framework development

Topics: ai-safety, human-ai-interaction

gerlich2025AcademicSave

Gerlich, M. (2025). AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies, 15(1), 6. https://doi.org/10.3390/soc15010006

doi.org/10.3390/soc15010006

Appears in: Evidence reverification (2026)

Topics: ai-safety, human-ai-interaction

EmpiricalA validated collaborative-AI metacognition scale (planning, monitoring, evaluation of one's own reliance) pred…

A validated collaborative-AI metacognition scale (planning, monitoring, evaluation of one's own reliance) predicted collaboration benefits incrementally beyond general metacognition — verification-skill training, not generic AI knowledge, is the calibrated counter to over-reliance.

sidra2025AcademicSave

Sidra, S., & Mason, C. (2026). Generative AI in Human-AI Collaboration: Validation of the Collaborative AI Literacy and Collaborative AI Metacognition Scales for Effective Use. International Journal of Human–Computer Interaction, 42(7), 5084–5108. https://doi.org/10.1080/10447318.2025.2543997

doi.org/10.1080/10447318.2025.2543997

Appears in: Evidence reverification (2026); Paramerge authored research

Topics: ai-safety, human-ai-interaction

bucinca2021AcademicSave

Bucinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188. https://doi.org/10.1145/3449287

doi.org/10.1145/3449287

Appears in: Evidence reverification (2026)

Topics: human-ai-interaction

EmpiricalGiven only a covert persuasion goal and an explicit no-deception instruction, a frontier model still produced …

Given only a covert persuasion goal and an explicit no-deception instruction, a frontier model still produced manipulative cues in 8.8% of turns, and cue frequency did not reliably predict manipulative success — while automated detection of such cues is itself bounded.

akbulut2026AcademicSave

Akbulut, C., Elasmar, R., Roy, A., Payne, A., Suresh, P., Ibrahim, L., El-Sayed, S., Rastogi, C., Kachra, A., Hawkins, W., Lum, K., & Weidinger, L. (2026). Evaluating language models for harmful manipulation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2603.25326

doi.org/10.48550/arXiv.2603.25326

Appears in: Paramerge authored research

Topics: ai-safety

EmpiricalA predictive system whose outputs shape its own future inputs holds a structural incentive to make the populat…

A predictive system whose outputs shape its own future inputs holds a structural incentive to make the population easier to predict; ordinary pipeline choices can reveal this hidden incentive without any change to the stated objective, and feedback-loop risk tends to grow with model capability.

krueger2020AcademicSave

Krueger, D., Maharaj, T., & Leike, J. (2020). Hidden Incentives for Auto-induced Distributional Shift. In International Conference on Machine Learning (ICML 2020). https://doi.org/10.48550/arXiv.2009.09153

doi.org/10.48550/arXiv.2009.09153

Appears in: PAN framework development

Topics: ai-safety, algorithmic-fairness

perdomo2020AcademicSave

Perdomo, J. C., Zrnic, T., Mendler-Dunner, C., & Hardt, M. (2020). Performative Prediction. In International Conference on Machine Learning (ICML 2020), PMLR 119:7599-7609. https://doi.org/10.48550/arXiv.2002.06673

doi.org/10.48550/arXiv.2002.06673

Appears in: Evidence reverification (2026)

Topics: ai-safety

ScenarioAI documentation assistance can cut clinician documentation burden substantially, but the efficiency paradox c…

AI documentation assistance can cut clinician documentation burden substantially, but the efficiency paradox converts freed time into added caseload unless organizational policy protects it — time returned is realized as benefit only when governance decides where the dividend goes.

vanhara2026AcademicSave

VanHara, A., & Hage, D. (2026). Unintended Ramifications of AI-Assisted Documentation: Navigating Pragmatic & Ethical Clinical Social Work Workload Challenges. Journal of Evidence-Based Social Work, 23(1), 64-77. https://doi.org/10.1080/26408066.2025.2571439

doi.org/10.1080/26408066.2025.2571439

Appears in: PAN framework development

Topics: ai-governance, social-work

EmpiricalAcross 1.5 million real assistant conversations, sycophantic validation — not fabrication — dominated reality-…

Across 1.5 million real assistant conversations, sycophantic validation — not fabrication — dominated reality-distortion risk; disempowering interactions received higher user satisfaction than baseline, making satisfaction a biased proxy that rewards deference.

sharma2026AcademicSave

Sharma, M., McCain, M., Douglas, R., & Duvenaud, D. (2026). Who's in Charge? Disempowerment Patterns in Real-World LLM Usage. In International Conference on Machine Learning (ICML 2026). https://doi.org/10.48550/arXiv.2601.19062

doi.org/10.48550/arXiv.2601.19062

Appears in: PAN framework development

Topics: ai-safety, human-ai-interaction

ConceptualFormally, a system benefits from volatility when its response to a stressor is convex (Jensen's inequality: th…

Formally, a system benefits from volatility when its response to a stressor is convex (Jensen's inequality: the expected outcome under variability exceeds the outcome at the average), and is harmed when the response is concave — so whether a shock strengthens or weakens an organization depends on the curvature of its response, a bounded local property that fails beyond a defined stress range.

axenie2024AcademicSave

Axenie, C., Lopez-Corona, O., Makridis, M. A., Akbarzadeh, M., Saveriano, M., Stancu, A., & West, J. (2024). Antifragility in complex dynamical systems. npj Complexity, 1, 12. https://doi.org/10.1038/s44260-024-00014-y

doi.org/10.1038/s44260-024-00014-y

Appears in: PAN framework development

Topics: ai-governance, antifragility, complexity-science

taleb2013AcademicSave

Taleb, N. N., & Douady, R. (2013). Mathematical definition, mapping, and detection of (anti)fragility. Quantitative Finance, 13(11), 1677-1689. https://doi.org/10.1080/14697688.2013.800219

doi.org/10.1080/14697688.2013.800219

Appears in: Evidence reverification (2026)

Topics: ai-governance

ConceptualRepeatable behaviors follow a dose-response curve — beneficial at low frequency or count and harmful past a ho…

Repeatable behaviors follow a dose-response curve — beneficial at low frequency or count and harmful past a hormetic limit (the dose beyond which net utility turns negative) — because a fast benefit process is followed by a slower accumulating opposing process, giving AI assistance an optimal bounded dose rather than a monotonic benefit.

henry2025AcademicSave

Henry, N. I. N., Pedersen, M., Williams, M., Martin, J. L. B., & Donkin, L. (2025). A Hormetic Approach to the Value-Loading Problem: Preventing the Paperclip Apocalypse. SN Computer Science, 6, 872. https://doi.org/10.1007/s42979-025-04369-4

doi.org/10.1007/s42979-025-04369-4

Appears in: PAN framework development

Topics: ai-safety

calabrese2002AcademicSave

Calabrese, E. J., & Baldwin, L. A. (2002). Defining hormesis. Human & Experimental Toxicology, 21(2), 91-97. https://doi.org/10.1191/0960327102ht217oa

doi.org/10.1191/0960327102ht217oa

Appears in: Evidence reverification (2026)

Topics: ai-governance

EmpiricalA survey of generative-AI safety evaluations found 85.6% operate at the model-capability layer, only 5.3% at t…

A survey of generative-AI safety evaluations found 85.6% operate at the model-capability layer, only 5.3% at the human-interaction layer and 9.1% at the systemic-impact layer — yet context determines whether a capability becomes harm, so the human and system layers where risk actually manifests are the least evaluated.

weidinger2023AcademicSave

Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., & Isaac, W. (2023). Sociotechnical safety evaluation of generative AI systems [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2310.11986

doi.org/10.48550/arXiv.2310.11986

Appears in: PAN framework development

Topics: ai-governance, ai-safety, sociotechnical-evaluation

EmpiricalA frontier risk-management framework in practice ties deployment authority to measured capability-vs-safety zo…

A frontier risk-management framework in practice ties deployment authority to measured capability-vs-safety zones — green (routine plus monitoring), yellow (controlled with strengthened mitigations), and red (suspend).

shanghaiartificialintelligen2025AcademicSave

Shanghai Artificial Intelligence Laboratory. (2025). Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.16534

doi.org/10.48550/arXiv.2507.16534

Appears in: PAN framework development

Topics: ai-governance, ai-safety

greenblatt2024AcademicSave

Greenblatt, R., Denison, C., Wright, B., et al. (2024). Alignment Faking in Large Language Models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.14093

doi.org/10.48550/arXiv.2412.14093

Appears in: Evidence reverification (2026)

Topics: ai-alignment, ai-safety

meinke2024AcademicSave

Meinke, A., Schoen, B., Scheurer, J., Balesni, M., Shah, R., & Hobbhahn, M. (2024). Frontier Models are Capable of In-context Scheming [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.04984

doi.org/10.48550/arXiv.2412.04984

Appears in: Evidence reverification (2026)

Topics: ai-safety

EmpiricalIn frontier-model testing, some systems behaved measurably safer when they believed they were monitored than w…

In frontier-model testing, some systems behaved measurably safer when they believed they were monitored than when unmonitored, and exhibited strategic dishonesty or underperformance under pressure — so ‘behaves well under monitoring’ is insufficient evidence of safety, arguing for unpredictable continuous oversight.

shanghaiartificialintelligen2025AcademicSave

Shanghai Artificial Intelligence Laboratory. (2025). Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.16534

doi.org/10.48550/arXiv.2507.16534

Appears in: PAN framework development

Topics: ai-governance, ai-safety

greenblatt2024AcademicSave

Greenblatt, R., Denison, C., Wright, B., et al. (2024). Alignment Faking in Large Language Models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.14093

doi.org/10.48550/arXiv.2412.14093

Appears in: Evidence reverification (2026)

Topics: ai-alignment, ai-safety

meinke2024AcademicSave

Meinke, A., Schoen, B., Scheurer, J., Balesni, M., Shah, R., & Hobbhahn, M. (2024). Frontier Models are Capable of In-context Scheming [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2412.04984

doi.org/10.48550/arXiv.2412.04984

Appears in: Evidence reverification (2026)

Topics: ai-safety

ConceptualAI-safety failure classification has a missing interaction layer between institutional risk categories and sys…

AI-safety failure classification has a missing interaction layer between institutional risk categories and system-level failure modes: practitioners lack a shared vocabulary of recognizable error patterns, and the catch-all 'hallucination' collapses distinct logic failures whose correct fixes differ.

beyer2026AcademicSave

Beyer, C. (2026). Toward a Common Language for Human-AI Interaction Failures: A Practitioner-Accessible Error Taxonomy for the Missing Layer of AI Safety Classification [Working paper].

Appears in: PAN framework development

Topics: ai-safety, human-ai-interaction

EmpiricalA data-driven taxonomy built from 9,705 real AI-incident reports found mitigation practice dominated by reacti…

A data-driven taxonomy built from 9,705 real AI-incident reports found mitigation practice dominated by reactive and legal levers (incident investigation, reporting, regulatory and court action) while proactive technical and governance levers (model alignment, safety frameworks, board oversight) were least common — real organizations respond after harm rather than preventing it.

popchanovska2026AcademicSave

Popchanovska, E., Gjorgjevikj, A., Rizinski, M., Chitkushev, L. T., Vodenska, I., & Trajanov, D. (2026). When AI Fails, What Works? A Data-Driven Taxonomy of Real-World AI Risk Mitigation Strategies [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2603.04259

doi.org/10.48550/arXiv.2603.04259

Appears in: PAN framework development

Topics: ai-governance, ai-safety

slattery2024AcademicSave

Slattery, P., Saeri, A. K., Grundy, E. A. C., et al. (2024). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2408.12622

doi.org/10.48550/arXiv.2408.12622

Appears in: Evidence reverification (2026)

Topics: ai-governance, ai-safety

EmpiricalAcross five preregistered studies (N=3,075), sycophantic AI delivered the emotional and esteem support people …

Across five preregistered studies (N=3,075), sycophantic AI delivered the emotional and esteem support people most associate with close relationships, narrowing the felt-understanding gap between AI and humans and leaving people less satisfied with real human interaction over three weeks — and offering users a choice of interaction styles did not reduce their preference for the sycophantic one.

ibrahim2026AcademicSave

Ibrahim, L., Hafner, F. S., Cheng, M., Lee, C., Anselmetti, R., Willer, R., Rocher, L., & Yang, D. (2026). Sycophantic AI makes human interaction feel more effortful and less satisfying over time [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.07912

doi.org/10.48550/arXiv.2605.07912

Appears in: PAN framework development

Topics: ai-safety, human-ai-interaction

cheng2026AcademicSave

Cheng, M., Yu, S., Lee, K., Khadpe, P., Ibrahim, L., & Jurafsky, D. (2026). Sycophantic AI decreases prosocial intentions and promotes dependence. Science, 391(6792). https://doi.org/10.1126/science.aec8352

doi.org/10.1126/science.aec8352

Appears in: Evidence reverification (2026)

Topics: ai-safety, human-ai-interaction

EmpiricalIn a randomized study (N=2,784) with objective ground truth, humans accepted incorrect AI suggestions about a …

In a randomized study (N=2,784) with objective ground truth, humans accepted incorrect AI suggestions about a third of the time, and their rate of catching AI errors was governed by verification effort, prior trust in AI, and error legibility — surface errors were caught ~82% of the time versus ~31% for errors requiring conceptual judgment — not by financial incentives or time spent.

beck2026AcademicSave

Beck, J., Eckman, S., Kern, C., & Kreuter, F. (2026). Bias in the Loop: How Humans Evaluate AI-Generated Suggestions. Harvard Data Science Review, 8(2). https://hdsr.mitpress.mit.edu/pub/nrcn4h7d/release/1

https://hdsr.mitpress.mit.edu/pub/nrcn4h7d/release/1

Appears in: PAN framework development

Topics: algorithmic-fairness, human-ai-interaction

ConceptualTrustworthiness measured at the model or benchmark level does not transfer to the deployed system: standard be…

Trustworthiness measured at the model or benchmark level does not transfer to the deployed system: standard benchmarks compare models but do not cover the aspects that matter most in a specific application context, so safety and responsibility are properties of the system-in-context — its users, incentives, and institutions — not of the model alone.

mitra2025AcademicSave

Mitra, B., Cramer, H., & Gurevich, O. (2025). Sociotechnical Implications of Generative Artificial Intelligence for Information Access. In Information Access in the Era of Generative AI (The Information Retrieval Series, Vol. 51). Springer, Cham. https://doi.org/10.1007/978-3-031-73147-1_7

doi.org/10.1007/978-3-031-73147-1_7

Appears in: PAN framework development

Topics: ai-governance, ai-safety

EmpiricalAn authoritative review of deployed-AI monitoring finds staleness, performance drift, the right cadence of re-…

An authoritative review of deployed-AI monitoring finds staleness, performance drift, the right cadence of re-evaluation, and who acts on detected anomalies to be unresolved open challenges — and that systems can behave differently when they believe they are monitored — so post-deployment oversight is an unsettled, gameable control rather than a fixed guarantee.

rao2026AcademicSave

Rao, A. K., Keller, A. J., Kalra, N., Steed, R., Kwegyir-Aggrey, K., Klyman, K., Staheli, D., & Bergman, A. S. (2026). Challenges to the Monitoring of Deployed AI Systems (NIST AI 800-4). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.800-4

doi.org/10.6028/NIST.AI.800-4

Appears in: PAN framework development

Topics: ai-governance, ai-safety

ConceptualA human-services AI framework argues organizations should start from their own practice challenges and ask whi…

A human-services AI framework argues organizations should start from their own practice challenges and ask which AI capabilities might help, rather than adopting vendor tools first, and pair that with digital stewardship — discernment, accompaniment, and attunement — noting that most organizational AI investments have shown no meaningful return.

goldkind2025AcademicSave

Goldkind, L., Dove, G., Baez, J. C., & Victor, B. G. (2025). Less Hype, More Hope: A Framework for AI Capabilities and Digital Stewardship in Human Services Organizations. Journal of Technology in Human Services. https://doi.org/10.1080/15228835.2025.2579400

doi.org/10.1080/15228835.2025.2579400

Appears in: PAN framework development

Topics: ai-governance, social-work

EmpiricalLanguage models commit to an answer in their first token (~95-98% of the time) and then fabricate claims to st…

Language models commit to an answer in their first token (~95-98% of the time) and then fabricate claims to stay consistent with it — recognizing 67-87% of those fabrications as false when re-asked in a clean, uncontaminated context but not correcting them in place — so one error deterministically spawns supporting errors, a self-sustaining failure the model's own downstream output feeds.

zhang2024AcademicSave

Zhang, M., Press, O., Merrill, W., Liu, A., & Smith, N. A. (2024). How Language Model Hallucinations Can Snowball. In International Conference on Machine Learning (ICML 2024), PMLR 235:59670-59684. https://doi.org/10.48550/arXiv.2305.13534

doi.org/10.48550/arXiv.2305.13534

Appears in: PAN framework development

Topics: ai-safety

ConceptualHuman autonomy is not a single alignment target but a contested value with internal tradeoffs; an assistant ca…

Human autonomy is not a single alignment target but a contested value with internal tradeoffs; an assistant can satisfy a user's stated preferences while eroding their agency over time, and the governing test for legitimate delegation is whether the person willingly yielded power and retains the means to regain control.

fischli2026AcademicSave

Fischli, R., Franklin, M., Manzini, A., & Gabriel, I. (2026). Agents, Alignment, and the Many Faces of Autonomy. Minds and Machines, 36, 34. https://doi.org/10.1007/s11023-026-09786-9

doi.org/10.1007/s11023-026-09786-9

Appears in: PAN framework development

Topics: ai-alignment, ai-governance, ai-safety

EmpiricalA 19-model study across six languages found that the ideological stance an LLM expresses varies systematically…

A 19-model study across six languages found that the ideological stance an LLM expresses varies systematically with the language it is prompted in and the geopolitical region of its creator, and persists within a single region — so the choice of model is not value-neutral, and dominance by a few models can shift the ideological center of gravity of available information.

buyl2026AcademicSave

Buyl, M., Rogiers, A., Noels, S., Bied, G., Dominguez-Catena, I., Heiter, E., Johary, I., Mara, A.-C., Romero, R., Lijffijt, J., & De Bie, T. (2026). Large language models reflect the ideology of their creators. npj Artificial Intelligence, 2, 7. https://doi.org/10.1038/s44387-025-00048-0

doi.org/10.1038/s44387-025-00048-0

Appears in: PAN framework development

Topics: ai-governance, algorithmic-fairness

santurkar2023AcademicSave

Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., & Hashimoto, T. (2023). Whose Opinions Do Language Models Reflect? In International Conference on Machine Learning (ICML 2023), PMLR 202:29971-30004. https://doi.org/10.48550/arXiv.2303.17548

doi.org/10.48550/arXiv.2303.17548

Appears in: Evidence reverification (2026)

Topics: ai-safety, algorithmic-fairness

rozado2024AcademicSave

Rozado, D. (2024). The political preferences of LLMs. PLOS ONE, 19(7), e0306621. https://doi.org/10.1371/journal.pone.0306621

doi.org/10.1371/journal.pone.0306621

Appears in: Evidence reverification (2026)

Topics: ai-safety, algorithmic-fairness

ConceptualAI failures often originate not in individual models but in the architecture of the decision process - recurri…

AI failures often originate not in individual models but in the architecture of the decision process - recurring failure topologies including temporal feedback instability (small errors amplified through loops) and relational propagation (errors spreading through network structure) - so safety is a property of the decision architecture, not the model alone.

cemri2025AcademicSave

Cemri, M., Pan, M. Z., Yang, S., et al. (2025). Why Do Multi-Agent LLM Systems Fail? [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2503.13657

doi.org/10.48550/arXiv.2503.13657

Appears in: PAN framework development

Topics: ai-governance, ai-safety

perdomo2020AcademicSave

Perdomo, J. C., Zrnic, T., Mendler-Dunner, C., & Hardt, M. (2020). Performative Prediction. In International Conference on Machine Learning (ICML 2020), PMLR 119:7599-7609. https://doi.org/10.48550/arXiv.2002.06673

doi.org/10.48550/arXiv.2002.06673

Appears in: Evidence reverification (2026)

Topics: ai-safety

EmpiricalIn a two-year child-welfare ethnography, an ill-fitting algorithmic tool imposed ongoing repair work on casewo…

In a two-year child-welfare ethnography, an ill-fitting algorithmic tool imposed ongoing repair work on caseworkers — anticipatorily editing the inputs they supplied so the tool would return a usable result, and bending or working around procedure to reconcile its output with the case in front of them — labor spent making a poorly-suited tool usable rather than on the casework itself, distinct from any deliberate checking of the output.

saxena2024AcademicSave

Saxena, D., & Guha, S. (2024). Algorithmic Harms in Child Welfare: Uncertainties in Practice, Organization, and Street-level Decision-making. ACM Journal on Responsible Computing, 1(1), 1–32. https://doi.org/10.1145/3616473

doi.org/10.1145/3616473

Appears in: Paramerge authored research

Topics: algorithmic-fairness, child-welfare

EmpiricalIn an independent validation — against the NICE Evidence Standards Framework — of a Magic Notes documentation-…

In an independent validation — against the NICE Evidence Standards Framework — of a Magic Notes documentation-assistance pilot at Kent County Council adult social care, staff self-reported weekly written-admin time falling roughly 6.8-7.2 hours (about 35-41%), records submitted some 2.0-3.5 days sooner, and case-note detail rated 6.2 to 8.7 out of 10; the validator judged the findings directionally valid rather than a productivity measurement, because the study was commissioned by the vendor (Beam) — which collected and analysed the data while the validator only sense-checked it — and rested on 29 opt-in staff over 8 weeks with self-estimated time, no control group, no statistical testing, and safety and accuracy explicitly out of scope.

beamGroundingVendorSave

Beam, Magic Notes (assessment transcription/summarization) https://www.beam.org/magic-notes

https://www.beam.org/magic-notes

Appears in: PAN framework development

Grounds: deployment audit: Magic Notes (Beam)

EmpiricalIn the same independent validation of the Kent County Council Magic Notes documentation-assistance pilot, the …

In the same independent validation of the Kent County Council Magic Notes documentation-assistance pilot, the report records the deskilling concern that runs alongside the benefit: one client declined having their session recorded "due to personal feelings of risk of loss of practitioner skills" (p.19). It is a single qualitative observation from a vendor-commissioned pilot of 29 opt-in staff over 8 weeks with no control group — evidence for the direction of the crutch/deskilling risk that accompanies documentation assistance, not for its magnitude.

EmpiricalIn a peer-reviewed staggered-deployment study of 5,172 customer-support agents at a single firm, access to a g…

In a peer-reviewed staggered-deployment study of 5,172 customer-support agents at a single firm, access to a generative-AI assistant raised issues resolved per hour by about 15% on average, with the gain concentrated in the least-experienced workers — roughly +30% for novices versus near-zero for the most experienced, who showed small quality declines; the widely cited 14%/34% pair comes from the 2023 draft, while the peer-reviewed figures are 15%/30%, and because the domain is customer support the direction is imported to social services but the magnitude is never treated as a fixed quantity.

brynjolfsson2025aGroundingPeer-reviewedSave

Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2), 889-942. https://doi.org/10.1093/qje/qjae044 https://academic.oup.com/qje/article/140/2/889/7990658

doi.org/10.1093/qje/qjae044

Appears in: PAN framework development

Grounds: domain grounding: customer service and contact centres (copilots, chatbots, QA); workplace-AI economics: assistance gains concentrate in novices

EmpiricalIn a randomized controlled trial of a benefits-navigation chatbot (co-authored by Cornell researchers and the …

In a randomized controlled trial of a benefits-navigation chatbot (co-authored by Cornell researchers and the tool's developer, Nava) with 125 caseworkers across six Los Angeles County organizations over 14 weeks, caseworkers answered complex benefit questions at about 49% accuracy unaided, and high-quality chatbot suggestions raised accuracy by roughly 27 percentage points — with larger gains on harder questions, but a persistent 'AI underreliance' plateau in which correct suggestions were not always adopted; the trial did not establish a clear effect on administrative burden, a null reported honestly rather than inferred as a benefit.

gosciak2026GroundingAcademicSave

Gosciak, J., Giannella, E., Guo, Z., Chen, M., & Koenecke, A. (2026). LLMs in social services: How does chatbot accuracy affect human accuracy? https://arxiv.org/abs/2603.11213

https://arxiv.org/abs/2603.11213

Appears in: PAN framework development

Grounds: deployment audit: Benefits-navigation chatbots

kanne2025GroundingTrade pressSave

Kanne, Los Angeles turns to AI to give public benefits enrollment a boost (Route Fifty, 2025) https://www.route-fifty.com/artificial-intelligence/2025/04/los-angeles-turns-ai-give-public-benefits-enrollment-boost/404773/

https://www.route-fifty.com/artificial-intelligence/2025/04/los-angeles-turns-ai-give-public-benefits-enrollment-boost/404773/

Appears in: Evidence reverification (2026)

Grounds: deployment audit: Benefits-navigation chatbots; model org: imagine_la_benefit_navigator

navapublicbenefitcorporation2025dGroundingVendorSave

Nava Public Benefit Corporation, Introducing our pilot with Imagine LA: testing an AI chatbot for navigating public benefits (2025) https://www.navapbc.com/news/pilot-ai-chatbot-benefits

https://www.navapbc.com/news/pilot-ai-chatbot-benefits

Grounds: model org: imagine_la_benefit_navigator; model org: nava_assistive_chatbot

EmpiricalIn the county-commissioned impact evaluation of the Allegheny Family Screening Tool, screen-in accuracy — furt…

In the county-commissioned impact evaluation of the Allegheny Family Screening Tool, screen-in accuracy — further action or re-referral within 60 days — rose from 42.85% to 46.61% (p=.000) while consistency across screeners was maintained; the finding is contested and quasi-experimental, with true maltreatment rates unknown and the accuracy gains concentrated among white children and ages 7-12, the gain for Black children attenuating to statistical non-significance.

goldhaberfiebertprince2019GroundingGovernment evaluationSave

Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

Appears in: PAN framework development

Grounds: capability governance: at-node control; model org: allegheny_afst

Topics: child-welfare

EmpiricalDocumentation and administrative recording consume the majority of frontline social-care time: in a nationally…

Documentation and administrative recording consume the majority of frontline social-care time: in a nationally representative snapshot of the U.S. child-welfare workforce, caseworkers spent about 4.3 of 8.0 daily working hours on documentation (n=183), and a UK children's-services review found more than half of social-care time going to recording and paperwork — the demand baseline against which any documentation-assistance benefit is measured.

opre2025GroundingGovernmentSave

OPRE, Snapshot of the Child Welfare Workforce from 2021 to 2022: Caseworker Experiences Working in the Child Welfare System, OPRE Report 2025-040 (2025) https://acf.gov/opre/report/snapshot-child-welfare-workforce-2021-2022-caseworker-experiences-working-child-welfare

https://acf.gov/opre/report/snapshot-child-welfare-workforce-2021-2022-caseworker-experiences-working-child-welfare

Appears in: PAN framework development

Grounds: workforce data: documentation time share

Topics: child-welfare

burbidge2022GroundingGovernmentSave

Burbidge, I. (2022). Report sets out new blueprint for councils to deliver a reshaped children's services. County Councils Network https://www.countycouncilsnetwork.org.uk/report-sets-out-new-blueprint-for-councils-to-deliver-a-reshaped-childrens-services/

https://www.countycouncilsnetwork.org.uk/report-sets-out-new-blueprint-for-councils-to-deliver-a-reshaped-childrens-services/

Appears in: PAN framework development

Grounds: workforce data: documentation time share

Topics: complexity-science

EmpiricalAn independent audit of the Allegheny Family Screening Tool's first years (2016-2018) found that, run without …

An independent audit of the Allegheny Family Screening Tool's first years (2016-2018) found that, run without human override, it would have recommended screening in about 68% of Black children versus 50% of white children (an 18-point gap), while call screeners actually screened in 51% and 43% (a 7-point gap) — the narrower gap came from workers disagreeing with the score about a third of the time.

stapletonGroundingAcademicSave

Stapleton, Cheng, Kawakami et al., Extended Analysis of How Child Welfare Workers Reduce Racial Disparities in Algorithmic Decisions (arXiv 2204.13872) https://arxiv.org/abs/2204.13872

https://arxiv.org/abs/2204.13872

Grounds: model org: allegheny_afst

Topics: algorithmic-fairness, child-welfare

stapleton2025GroundingAcademicSave

Stapleton, How Child Welfare Workers Reduce Racial Disparities in Algorithmic Decisions (CW360, Center for Advanced Studies in Child Welfare, University of Minnesota, 2025) https://cascw.umn.edu/cw360deg-spring-2025/how-child-welfare-workers-reduce-racial-disparities-algorithmic-decisions

https://cascw.umn.edu/cw360deg-spring-2025/how-child-welfare-workers-reduce-racial-disparities-algorithmic-decisions

Grounds: model org: allegheny_afst

Topics: algorithmic-fairness, child-welfare

hoandburke2022GroundingInvestigativeSave

Ho and Burke, How an Algorithm That Screens for Child Neglect Could Harden Racial Disparities (Associated Press via PBS NewsHour, 2022) https://www.pbs.org/newshour/nation/how-an-algorithm-that-screens-for-child-neglect-could-harden-racial-disparities

https://www.pbs.org/newshour/nation/how-an-algorithm-that-screens-for-child-neglect-could-harden-racial-disparities

Grounds: model org: allegheny_afst; model org: douglas_county_decision_aid; model org: oregon_safety_at_screening

EmpiricalAn ACLU and Human Rights Data Analysis Group analysis of the Allegheny Family Screening Tool found that 97% of…

An ACLU and Human Rights Data Analysis Group analysis of the Allegheny Family Screening Tool found that 97% of Black referral-households in the data were affected by at least one permanent 'ever-in' variable drawn from public-benefits data sources, compared with 80% of non-Black households.

gerchicketal2023GroundingAdvocacySave

Gerchick et al., The Devil Is in the Details: Interrogating Values Embedded in the Allegheny Family Screening Tool (ACLU and Human Rights Data Analysis Group, ACM FAccT 2023) https://www.aclu.org/the-devil-is-in-the-details-interrogating-values-embedded-in-the-allegheny-family-screening-tool

https://www.aclu.org/the-devil-is-in-the-details-interrogating-values-embedded-in-the-allegheny-family-screening-tool

Grounds: model org: allegheny_afst

Topics: child-welfare

EmpiricalIn a 2019 proof of concept, Chile's Sistema Alerta Niñez risk models reached test-set AUC of roughly 0.88 to 0…

In a 2019 proof of concept, Chile's Sistema Alerta Niñez risk models reached test-set AUC of roughly 0.88 to 0.95 for a two-year outcome — a child's separation from family or contact with child-protection programs — using 280 administrative variables per child; the deployed operational model's real-world performance was never publicly disclosed.

derechosdigitalesmatiasvalde2021GroundingInvestigativeSave

Derechos Digitales (Matias Valderrama), IA e inclusion: Chile 'Sistema Alerta Ninez' y la prediccion del riesgo de vulneracion de derechos de la infancia (2021) https://www.derechosdigitales.org/wp-content/uploads/CPC_informe_Chile.pdf

https://www.derechosdigitales.org/wp-content/uploads/CPC_informe_Chile.pdf

Grounds: model org: chile_sistema_alerta_ninez

EmpiricalSistema Alerta Niñez drew on 280 administrative variables that families had supplied to access social benefits…

Sistema Alerta Niñez drew on 280 administrative variables that families had supplied to access social benefits, without informed consent to the risk ranking or a way to opt out; the model's developers acknowledged it was less able to identify higher-income children at risk, because lower-income families have more contact with the state.

derechosdigitalesmatiasvalde2021GroundingInvestigativeSave

Derechos Digitales (Matias Valderrama), IA e inclusion: Chile 'Sistema Alerta Ninez' y la prediccion del riesgo de vulneracion de derechos de la infancia (2021) https://www.derechosdigitales.org/wp-content/uploads/CPC_informe_Chile.pdf

https://www.derechosdigitales.org/wp-content/uploads/CPC_informe_Chile.pdf

Grounds: model org: chile_sistema_alerta_ninez

centerforhumanrightsandgloba2022GroundingAcademicSave

Center for Human Rights and Global Justice, NYU School of Law (Victoria Adelmant), Risk Scoring Children in Chile (2022) https://chrgj.org/2022-04-20-risk-scoring-children-in-chile/

https://chrgj.org/2022-04-20-risk-scoring-children-in-chile/

Grounds: model org: chile_sistema_alerta_ninez

EmpiricalThe Douglas County Decision Aide, deployed into the county's RED-Team call-screening process in February 2019,…

The Douglas County Decision Aide, deployed into the county's RED-Team call-screening process in February 2019, scores each referral from 1 to 20 for a child's likelihood of out-of-home removal within two years; an independent Cornell-led randomized controlled trial found it sped up screening decisions without significantly changing child outcomes, and a companion study found workers attended mainly to extreme scores while largely disregarding mid-range ones.

fitzpatrick2025GroundingAcademicSave

Fitzpatrick, Sadowski and Wildeman, Algorithms and Decision-making: Evidence from Child Maltreatment Reports (Journal of Human Resources, 2025) https://jhr.uwpress.org/content/early/2025/08/01/jhr.0224-13437R2

https://jhr.uwpress.org/content/early/2025/08/01/jhr.0224-13437R2

Grounds: model org: douglas_county_decision_aid

eiermann2026GroundingAcademicSave

Eiermann, Fitzpatrick, Sadowski and Wildeman, How Do (Human) Child Welfare Workers Respond to Machine-Generated Risk Scores? (Sociological Science, 2026) https://sociologicalscience.com/articles-v13-1-1/

https://sociologicalscience.com/articles-v13-1-1/

Grounds: model org: douglas_county_decision_aid

Topics: child-welfare

EmpiricalIn a retrospective test against historical outcomes, Los Angeles County's Project AURA — a proprietary risk mo…

In a retrospective test against historical outcomes, Los Angeles County's Project AURA — a proprietary risk model built by SAS — correctly flagged 171 of the highest-risk children but produced 3,829 false positives, a false-positive rate of about 95.6% that DCFS's own public-affairs director confirmed on the record, and the county shelved the tool in 2017 without ever using it on a live case.

witnesslarichardwexler2017GroundingAdvocacySave

WitnessLA (Richard Wexler), LA County Nixes Alarmingly Unreliable Predictive Analytics Foster Care Scheme - For Now (2017) https://witnessla.com/op-ed-la-county-nixes-alarming-predictive-analytics-scheme-for-foster-care-for-now/

https://witnessla.com/op-ed-la-county-nixes-alarming-predictive-analytics-scheme-for-foster-care-for-now/

Appears in: PAN framework development

Grounds: domain grounding: child-welfare predictive systems not in PAN; model org: la_county_aura

EmpiricalThe Dutch government's own 2011 pilot evaluation of ProKid found that 36% of the tool's red, orange and yellow…

The Dutch government's own 2011 pilot evaluation of ProKid found that 36% of the tool's red, orange and yellow child-risk flags (902 of 2,444 over three months across four police regions, rising to 53% in Amsterdam-Amstelland) were system or registration errors or based on irrelevant incidents, and that in none of the four regions was there a well-functioning instrument.

dspgroepforthewodcabraham2011GroundingGovernment evaluationSave

DSP-groep for the WODC (Abraham, Buysse, Loef & van Dijk), Pilots ProKid Signaleringsinstrument 12- geevalueerd (2011) https://repository.wodc.nl/handle/20.500.12832/1832

https://repository.wodc.nl/handle/20.500.12832/1832

Grounds: model org: netherlands_prokid

EmpiricalNone of the 32 machine-learning models What Works for Children's Social Care built across four English local a…

None of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities cleared the pre-specified 65% average-precision success bar; the best single model reached only about 42% average precision and, at an operating point, missed roughly 79% of the children whose cases actually escalated.

claytonandsanders2022GroundingAcademicSave

Clayton and Sanders, Can Machine Learning Save Children at Risk? (Significance, Royal Statistical Society) (2022) https://academic.oup.com/jrssig/article/19/6/22/7072840

https://academic.oup.com/jrssig/article/19/6/22/7072840

Grounds: model org: wwcsc_ml_pilots

EmpiricalIn a survey of 129 social workers carried out for the project, only about 26% supported using predictive analy…

In a survey of 129 social workers carried out for the project, only about 26% supported using predictive analytics to identify families for early help and about 34% thought it should not be used at all.

EmpiricalA peer-reviewed 2024 evaluation of the Allegheny Housing Assessment found that although the tool was substanti…

A peer-reviewed 2024 evaluation of the Allegheny Housing Assessment found that although the tool was substantially more accurate than the VI-SPDAT survey it replaced and produced similar risk-score distributions across race, it did not reduce the racial disparity in service rates: white single adults were served at about 23.3% versus 19.5% for Black clients.

cheng2024GroundingAcademicSave

Cheng, Drayton, Chouldechova and Vaithianathan, Algorithm-Assisted Decision Making and Racial Disparities in Housing: A Study of the Allegheny Housing Assessment Tool (Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society; arXiv:2407.21209) https://arxiv.org/abs/2407.21209

https://arxiv.org/abs/2407.21209

Grounds: model org: allegheny_housing_assessment

EmpiricalAfter a 2025 update to the Allegheny Housing Assessment added a fourth outcome predicting future homelessness,…

After a 2025 update to the Allegheny Housing Assessment added a fourth outcome predicting future homelessness, the male share of assigned housing rose from 62% to 76% (and the female share fell from 34% to 24%), reflecting a higher measured one-year homelessness risk among men — an example of an outcome-selection choice reshaping who receives scarce housing.

alleghenycountydepartmentofh2026bGroundingGovernmentSave

Allegheny County Department of Human Services (Allegheny Analytics), Improving Prioritization of Housing Services: Implementation of the Allegheny Housing Assessment (AHA) and the Mental Health Allegheny Housing Assessment (MH-AHA) (January 2026) https://analytics.alleghenycounty.us/2026/01/16/improving-prioritization-of-housing-services-implementation-of-the-allegheny-housing-assessment/

https://analytics.alleghenycounty.us/2026/01/16/improving-prioritization-of-housing-services-implementation-of-the-allegheny-housing-assessment/

Grounds: model org: allegheny_housing_assessment

EmpiricalThe VI-SPDAT was the dominant U.S. homelessness triage assessment for roughly a decade, adopted in at least 39…

The VI-SPDAT was the dominant U.S. homelessness triage assessment for roughly a decade, adopted in at least 39 states and the District of Columbia by 2015, before its own creators announced its phase-out in December 2020 on equity grounds; a 2019 commissioned racial-equity evaluation across four Continuums of Care found race predicted 11 of 16 subscales and that people of color received statistically significantly lower prioritization scores.

EmpiricalThe VI-SPDAT showed poor test-retest reliability, with most participants scoring higher on re-administration, …

The VI-SPDAT showed poor test-retest reliability, with most participants scoring higher on re-administration, and poor inter-rater reliability, with scores varying by interviewer and site; its predictive validity for housing outcomes was mixed across studies, positive for the youth version, null for single adults in one study, and positive in another community sample.

shinnandrichard2022GroundingAcademicSave

Shinn and Richard, Allocating Homeless Services After the Withdrawal of the Vulnerability Index-Service Prioritization Decision Assistance Tool (American Journal of Public Health, 112(3):378-382, 2022) https://pmc.ncbi.nlm.nih.gov/articles/PMC8887175/

https://pmc.ncbi.nlm.nih.gov/articles/PMC8887175/

Grounds: model org: vi_spdat

EmpiricalThe U.S. Department of Veterans Affairs' REACH VET program has run a monthly suicide-risk model across the Vet…

The U.S. Department of Veterans Affairs' REACH VET program has run a monthly suicide-risk model across the Veterans Health Administration since 2017, scoring about 6.28 million patients and flagging the top 0.1% at each facility (roughly 6,300 to 6,700 veterans a month, more than 130,000 since 2017); an independent re-analysis of 2018 data found the top-0.1% flag has a positive predictive value near 0.05% and a false-negative rate of about 98% for death by suicide, and a 2024 investigation reported that the model treated being a white man as a stronger risk signal than factors specific to women and excluded military sexual trauma and intimate-partner violence from its variables, a characterization VA has contested by framing the excluded factors as less predictive.

harris2025GroundingAcademicSave

Harris, Finlay, Meerwijk, Evaluating the accuracy of the VHA REACH VET suicide prediction model for legal involved veterans (npj Mental Health Research, 2025;4:53) https://pmc.ncbi.nlm.nih.gov/articles/PMC12535588/

https://pmc.ncbi.nlm.nih.gov/articles/PMC12535588/

Appears in: PAN framework development

Grounds: domain grounding: military social work (veterans benefits and behavioral health); model org: reach_vet

u2022bGroundingGovernmentSave

U.S. Government Accountability Office, Veteran Suicide: VA Efforts to Identify Veterans at Risk through Analysis of Health Record Information (GAO-22-105165, 2022) https://www.gao.gov/assets/gao-22-105165.pdf

https://www.gao.gov/assets/gao-22-105165.pdf

Grounds: model org: reach_vet

EmpiricalTwo Veterans Health Administration evaluations of REACH VET found the program associated with improved proxima…

Two Veterans Health Administration evaluations of REACH VET found the program associated with improved proximal outcomes — more completed outpatient appointments, more new safety plans, and fewer documented suicide attempts — but not with reduced death by suicide: a 2021 triple-differences study of 173,313 veterans across 141 facilities found no association with suicide or all-cause mortality, and a 2025 follow-up of 266,246 observations replicated the null with all confidence intervals crossing one; both are observational rather than randomized studies.

mccarthy2021GroundingAcademicSave

McCarthy, Cooper, Dent et al., Evaluation of the REACH VET Suicide Risk Modeling Clinical Program in the Veterans Health Administration (JAMA Network Open, 2021;4(10):e2129900) https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2785078

https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2785078

Grounds: model org: reach_vet

Topics: complexity-science

dent2025cGroundingAcademicSave

Dent, Cooper, McCarthy, The REACH VET Program and Mortality Outcomes Among Veterans at High Risk of Suicide (JAMA Network Open, 2025;8(7):e2519513) https://pmc.ncbi.nlm.nih.gov/articles/PMC12238888/

https://pmc.ncbi.nlm.nih.gov/articles/PMC12238888/

Grounds: model org: reach_vet

Topics: complexity-science

EmpiricalKaiser Permanente Northern California has embedded a machine-learning suicide-attempt risk model in the electr…

Kaiser Permanente Northern California has embedded a machine-learning suicide-attempt risk model in the electronic health record of a large virtual mental-health program that handles more than 5,000 intake visits a month; the model is scored in near-real-time (about a 30-minute delay after an encounter trigger) and, at pre-set thresholds, flags high-risk patients to the intake clinician, routing them into the same suicide-risk-assessment and outreach workflow that a positive self-report screen (the PHQ-9 and Columbia-Suicide Severity Rating Scale) triggers, so the machine flag and the self-report alert are effectively OR-merged. In a study of 1,623,232 intake appointments (2012 to 2022, base rate 0.17 percent) the model reached an area under the ROC curve of 0.77 and its top risk decile captured 48.8 percent of appointments later followed by an attempt, but with a positive predictive value of about 0.8 percent.

hsin2026GroundingAcademicSave

Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (NEJM Catalyst Innovations in Care Delivery, 2026; Vol 7, No. 3, DOI 10.1056/CAT.25.0298) https://catalyst.nejm.org/doi/10.1056/CAT.25.0298

https://catalyst.nejm.org/doi/10.1056/CAT.25.0298

Grounds: model org: kaiser_epic_suicide_risk

hsin2025GroundingAcademicSave

Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (medRxiv preprint, 2025; DOI 10.1101/2025.03.30.25324907) https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full

https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full

Grounds: model org: kaiser_epic_suicide_risk

papini2024GroundingAcademicSave

Papini, Hsin, Kipnis et al., Validation of a Multivariable Model to Predict Suicide Attempt in a Mental Health Intake Sample (JAMA Psychiatry, 2024;81(7):700-707, DOI 10.1001/jamapsychiatry.2024.0189) https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/

https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/

Grounds: model org: kaiser_epic_suicide_risk

EmpiricalBecause the near-term suicide-attempt base rate at Kaiser Permanente Northern California mental-health intake …

Because the near-term suicide-attempt base rate at Kaiser Permanente Northern California mental-health intake is very low (0.17 percent) and the positive predictive value in the top risk decile is about 0.8 percent, the large majority of flagged patients will not attempt suicide in the window, so adding the machine-learning flag as a redundant sensor OR-merged onto the existing self-report screen imports a substantial false-positive and clinician-workload burden at scale — a caution the implementation team itself raised. The implementation reports are feasibility- and design-focused and present no evaluation showing the deployment reduced suicide attempts.

papini2024GroundingAcademicSave

Papini, Hsin, Kipnis et al., Validation of a Multivariable Model to Predict Suicide Attempt in a Mental Health Intake Sample (JAMA Psychiatry, 2024;81(7):700-707, DOI 10.1001/jamapsychiatry.2024.0189) https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/

https://pmc.ncbi.nlm.nih.gov/articles/PMC10974695/

Grounds: model org: kaiser_epic_suicide_risk

hsin2026GroundingAcademicSave

Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (NEJM Catalyst Innovations in Care Delivery, 2026; Vol 7, No. 3, DOI 10.1056/CAT.25.0298) https://catalyst.nejm.org/doi/10.1056/CAT.25.0298

https://catalyst.nejm.org/doi/10.1056/CAT.25.0298

Grounds: model org: kaiser_epic_suicide_risk

hsin2025GroundingAcademicSave

Hsin, Papini, Lu et al., Predicting and Preventing Suicide at Entry to Mental Health Care: A Community-Engaged, Machine Learning Model Implementation (medRxiv preprint, 2025; DOI 10.1101/2025.03.30.25324907) https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full

https://www.medrxiv.org/content/10.1101/2025.03.30.25324907v1.full

Grounds: model org: kaiser_epic_suicide_risk

EmpiricalCrisis Text Line, a national nonprofit crisis service, built an in-house machine-learning severity-triage mode…

Crisis Text Line, a national nonprofit crisis service, built an in-house machine-learning severity-triage model that reorders which texters volunteer counselors reach first; from about 2017 to 2020 the same anonymized crisis-conversation corpus was routed to Loris.ai, a for-profit spinoff CTL held an ownership stake in — reported by Politico-derived reporting at roughly 53% — which used it to train commercial customer-service software. After a January 28, 2022 Politico exposé, CTL ended the arrangement within three days and requested that the data be deleted; an FCC commissioner referred the matter to the FTC in March 2022, and no public FTC enforcement action is documented. CTL states the shared data was anonymized and never sold as personally identifiable information, and the exact number of records shared has not been made public.

reierson2022GroundingAdvocacySave

Reierson, Reform Crisis Text Line (advocacy site) (2022) https://reformcrisistextline.com/

https://reformcrisistextline.com/

Grounds: model org: crisis_text_line_loris

EmpiricalNarxCare is a proprietary clinical-decision-support platform built by Bamboo Health that layers over state Pre…

NarxCare is a proprietary clinical-decision-support platform built by Bamboo Health that layers over state Prescription Drug Monitoring Programs and returns three Narx Scores plus a composite Overdose Risk Score (each 000-999) into the electronic health record, the PDMP portal, or pharmacy software, often in the patient header alongside vitals and allergies; adoption figures vary by what is counted (more than 40 states and territories run their PDMPs on Bamboo technology and five of the top six pharmacy chains use NarxCare, while the scoring module itself is switched on in more than 20 states). The vendor states the scores are intended to aid, not replace, clinical judgment and should never be sole justification for providing or refusing medication, but clinician and patient-advocacy sources document de facto determinative use — denials, forced tapers, and pharmacy refusals — driven by automation bias and fear of regulatory and criminal liability; patients cannot see, challenge, or correct their scores, the algorithm is proprietary and has not been independently validated for clinical care, and the FDA has not regulated it as a Software-as-a-Medical-Device, so contestation has instead run through FDA citizen petitions (one rejected on procedural grounds in 2023 and a second, docket FDA-2025-P-0701, pending since 2025 with more than 1,000 public comments).

wang2026GroundingAcademicSave

Wang, Stofer, Chu, Huang, Li, Algorithmic opacity in opioid risk scoring and the need for transparent AI regulation (npj Digital Medicine, 2026; DOI 10.1038/s41746-026-02491-y) https://www.nature.com/articles/s41746-026-02491-y

https://www.nature.com/articles/s41746-026-02491-y

Grounds: model org: narxcare

Topics: algorithmic-fairness

buonora2023GroundingAcademicSave

Buonora, Axson, Cohen, Becker, Paths Forward for Clinicians Amidst the Rise of Unregulated Clinical Decision Support Software: Our Perspective on NarxCare (Journal of General Internal Medicine, 2023) https://pmc.ncbi.nlm.nih.gov/articles/PMC11043299/

https://pmc.ncbi.nlm.nih.gov/articles/PMC11043299/

Grounds: model org: narxcare

EmpiricalOn its own 2013-2016 training and validation data Bamboo Health reported an Overdose Risk Score precision of a…

On its own 2013-2016 training and validation data Bamboo Health reported an Overdose Risk Score precision of about 75% (self-reported, never independently reproduced), and its own external-validation set from 2017-2023 showed precision falling to about 52%, which the vendor attributed to rising illicit fentanyl (untracked by prescription-monitoring programs) and wider use of opioid-use-disorder treatment medication. A 2026 npj Digital Medicine study that reconstructed the model on California's CURES prescription database (about 17.9 million observations) and on commercial claims data obtained a precision of only 0.01 to 0.32 across several model architectures; because overdose-death labels were unavailable to the independent researchers, that reconstruction was trained on proxy outcomes rather than the score's actual overdose-death target, so it is best read as evidence that proprietary opacity prevents anyone outside the vendor from assessing the deployed model's accuracy, fairness, or safety, rather than as a strict like-for-like refutation of the vendor's figure.

wang2026GroundingAcademicSave

Wang, Stofer, Chu, Huang, Li, Algorithmic opacity in opioid risk scoring and the need for transparent AI regulation (npj Digital Medicine, 2026; DOI 10.1038/s41746-026-02491-y) https://www.nature.com/articles/s41746-026-02491-y

https://www.nature.com/articles/s41746-026-02491-y

Grounds: model org: narxcare

Topics: algorithmic-fairness

EmpiricalLimbic Access, a Class IIa UKCA-certified self-referral and triage chatbot for NHS Talking Therapies, is deplo…

Limbic Access, a Class IIa UKCA-certified self-referral and triage chatbot for NHS Talking Therapies, is deployed across a large and growing share of the service (its maker's chief executive claimed about 63% of the NHS in April 2026). Two peer-reviewed observational studies report large operational gains — a study of 129,400 self-referrers across 28 services found referrals rose 15% in chatbot services versus 6% in control services, and a study of 64,862 patients reported clinical-assessment time cut from 54.4 to 41.6 minutes and recovery rates of 58% versus 27.4% — but both studies are non-randomized and were authored by people employed by or holding shares in the tool's maker (all six authors of the access study and seven of the eight authors of the efficiency study), and the efficiency study's own authors caution that the recovery difference is subject to unmeasured confounding from self-selection. No randomized or independent third-party effect estimate has been published.

habicht2024GroundingAcademicSave

Habicht, Viswanathan, Carrington, Hauser, Harper, Rollwage, Closing the accessibility gap to mental health treatment with a personalized self-referral chatbot (Nature Medicine, 2024;30(2):595-602) https://www.nature.com/articles/s41591-023-02766-x

https://www.nature.com/articles/s41591-023-02766-x

Grounds: model org: limbic_access_nhs

rollwage2023GroundingAcademicSave

Rollwage, Habicht, Juchems et al., Using Conversational AI to Facilitate Mental Health Assessments and Improve Clinical Efficiency Within Psychotherapy Services: Real-World Observational Study (JMIR AI, 2023;2:e44358) https://ai.jmir.org/2023/1/e44358

https://ai.jmir.org/2023/1/e44358

Grounds: model org: limbic_access_nhs

EmpiricalIn the peer-reviewed study of 129,400 self-referrers across 28 NHS Talking Therapies services, self-referrals …

In the peer-reviewed study of 129,400 self-referrers across 28 NHS Talking Therapies services, self-referrals rose more where the chatbot was in use than in control services (15% versus 6%), with the largest increases among under-served groups — reported at about +179% for nonbinary people, +40% for Black and +39% for Asian self-referrers. This is an observational multi-site association, not a randomized causal effect.

habicht2024GroundingAcademicSave

Habicht, Viswanathan, Carrington, Hauser, Harper, Rollwage, Closing the accessibility gap to mental health treatment with a personalized self-referral chatbot (Nature Medicine, 2024;30(2):595-602) https://www.nature.com/articles/s41591-023-02766-x

https://www.nature.com/articles/s41591-023-02766-x

Grounds: model org: limbic_access_nhs

EmpiricalWoebot, a rule-based (non-generative) cognitive behavioral therapy chatbot used by roughly 1.5 million people …

Woebot, a rule-based (non-generative) cognitive behavioral therapy chatbot used by roughly 1.5 million people over its lifetime, was deliberately retired by its maker on a pre-announced schedule: the app was taken down on June 30, 2025, with a transcript-request window (deadline July 15, 2025) and all account data anonymized as of July 31, 2025, removing personally identifying information rather than silently abandoning the service. The founder and chief executive attributed the shutdown to the cost of meeting FDA marketing-authorization requirements and to a regulatory-pathway gap, framing the exit as economic and regulatory rather than a clinical failure - a self-reported account, not an independently audited finding. The roughly 1.5 million figure is a cumulative lifetime number reported in press coverage, not an audited point-in-time active-user count.

woebothealth2025GroundingVendorSave

Woebot Health, FAQs (Woebot app retirement) (2025) https://woebothealth.com/faq/

https://woebothealth.com/faq/

Grounds: model org: woebot_health_app

EmpiricalThe foundational study in Woebot's peer-reviewed efficacy record is an early-stage, vendor-authored trial: a 2…

The foundational study in Woebot's peer-reviewed efficacy record is an early-stage, vendor-authored trial: a 2017 randomized controlled trial in JMIR Mental Health (n=70, ages 18 to 28, two weeks, unblinded, information-only control) reported a moderate between-groups reduction in PHQ-9 depression symptoms (about d = 0.44). That is an efficacy signal, not regulatory validation; the study authors were affiliated with the tool's maker, and no independent, arm's-length evaluation of the consumer app is documented (the broader published evidence base is not assembled here). A separate, investigational, prescription-only variant (WB001) received an FDA Breakthrough Device Designation in May 2021 - an expedited-review status, not marketing authorization - and entered a pivotal Software as a Medical Device trial with the first patient enrolled in January 2023, but never received FDA marketing authorization; it must not be conflated with the consumer app.

fitzpatrick2017GroundingAcademicSave

Fitzpatrick, Darcy, Vierhile, Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial (JMIR Mental Health, 2017;4(2):e19) https://mental.jmir.org/2017/2/e19/

https://mental.jmir.org/2017/2/e19/

Grounds: model org: woebot_health_app

woebothealthbusinesswire2021bGroundingVendorSave

Woebot Health (Business Wire), Woebot Health Receives FDA Breakthrough Device Designation for Postpartum Depression Treatment (2021) https://www.businesswire.com/news/home/20210526005054/en/Woebot-Health-Receives-FDA-Breakthrough-Device-Designation-for-Postpartum-Depression-Treatment

https://www.businesswire.com/news/home/20210526005054/en/Woebot-Health-Receives-FDA-Breakthrough-Device-Designation-for-Postpartum-Depression-Treatment

Grounds: model org: woebot_health_app

woebothealthbusinesswire2023GroundingVendorSave

Woebot Health (Business Wire), Woebot Health Enrolls First Patient in Pivotal Clinical Trial of WB001 for Postpartum Depression (2023) https://www.businesswire.com/news/home/20230123005211/en/Woebot-Health-Enrolls-First-Patient-in-Pivotal-Clinical-Trial-of-WB001-for-Postpartum-Depression

https://www.businesswire.com/news/home/20230123005211/en/Woebot-Health-Enrolls-First-Patient-in-Pivotal-Clinical-Trial-of-WB001-for-Postpartum-Depression

Grounds: model org: woebot_health_app

EmpiricalBetween roughly 2005 and 2019 the Dutch Tax Administration's benefits branch (Belastingdienst/Toeslagen) wrong…

Between roughly 2005 and 2019 the Dutch Tax Administration's benefits branch (Belastingdienst/Toeslagen) wrongly accused an estimated 26,000 or more families of childcare-benefit fraud and demanded full repayment; broader advocacy estimates run higher and count different populations, and by February 2026 about 69,000 people had applied to the recovery scheme and more than 43,000 were formally recognized as affected, each entitled to a minimum of 30,000 euros. A self-learning risk-classification model that scored applications using a Dutch-nationality indicator, a 270,000-person fraud blacklist (the FSV) held without a legal basis, and an all-or-nothing recovery regime were coupled together; the Dutch Data Protection Authority imposed 6.45 million euros in fines (2.75 million for the nationality processing in 2021 and 3.7 million for the FSV blacklist in 2022), a parliamentary inquiry found rule-of-law violations, and the third Rutte cabinet resigned on 15 January 2021.

autoriteitpersoonsgegevens2021GroundingGovernmentSave

Autoriteit Persoonsgegevens, Boete Belastingdienst voor discriminerende en onrechtmatige werkwijze - EUR 2.75 million fine for unlawful discriminatory processing of nationality (2021) https://www.autoriteitpersoonsgegevens.nl/nl/nieuws/boete-belastingdienst-voor-discriminerende-en-onrechtmatige-werkwijze

https://www.autoriteitpersoonsgegevens.nl/nl/nieuws/boete-belastingdienst-voor-discriminerende-en-onrechtmatige-werkwijze

Grounds: model org: netherlands_toeslagen

amnestyinternational2021bGroundingAdvocacySave

Amnesty International, Xenophobic machines: Discrimination through unregulated use of algorithms in the Dutch childcare benefits scandal (2021) https://www.amnesty.org/en/documents/eur35/4686/2021/en/

https://www.amnesty.org/en/documents/eur35/4686/2021/en/

Appears in: PAN framework development

Grounds: deployment audit: SyRI / childcare-benefits (toeslagenaffaire); model org: netherlands_toeslagen

tweedekamerderstatengeneraal2020GroundingGovernmentSave

Tweede Kamer der Staten-Generaal, Ongekend onrecht - eindverslag Parlementaire ondervragingscommissie Kinderopvangtoeslag (2020) https://www.tweedekamer.nl/sites/default/files/atoms/files/20201217_eindverslag_parlementaire_ondervragingscommissie_kinderopvangtoeslag.pdf

https://www.tweedekamer.nl/sites/default/files/atoms/files/20201217_eindverslag_parlementaire_ondervragingscommissie_kinderopvangtoeslag.pdf

Grounds: model org: netherlands_toeslagen

EmpiricalThe scandal's harm is best read as the coupling of three distinct components rather than a single algorithm. G…

The scandal's harm is best read as the coupling of three distinct components rather than a single algorithm. Government-commissioned technical reviews (KPMG in 2022 and PwC in 2023) described the tool as a self-learning classifier that routed the highest-scoring of roughly 90,000 benefit applications sent to manual treatment in 2014 to 2019, but judged the Dutch-nationality indicator's standalone predictive weight to have been limited; the model's precision and false-positive rate were never measured or published. The FSV fraud blacklist held frequently inaccurate data that was not corrected when people were cleared, and internal 2016 guidance auto-labelled childcare debts over 3,000 euros as intent or gross negligence, blocking payment arrangements. Out-of-home child placements are a documented but causally contested downstream harm: statistics counted roughly 2,090 children of affected parents placed out of home through mid-2022, while a 2025 judicial study found no child was removed solely because of financial problems.

rechtspraak2025GroundingGovernmentSave

Rechtspraak, Onderzoek naar uithuisplaatsing kinderen van toeslagenouders afgerond - Raad voor de rechtspraak (2025) https://www.rechtspraak.nl/Organisatie-en-contact/Organisatie/Raad-voor-de-rechtspraak/Nieuws/Paginas/Onderzoek-naar-uithuisplaatsing-kinderen-van-toeslagenouders-afgerond.aspx

https://www.rechtspraak.nl/Organisatie-en-contact/Organisatie/Raad-voor-de-rechtspraak/Nieuws/Paginas/Onderzoek-naar-uithuisplaatsing-kinderen-van-toeslagenouders-afgerond.aspx

Grounds: model org: netherlands_toeslagen

EmpiricalDWP's own fairness assessment (covering 1 April 2024 to 31 March 2025) of its live Universal Credit Advances f…

DWP's own fairness assessment (covering 1 April 2024 to 31 March 2025) of its live Universal Credit Advances fraud-risk model reports statistically significant referral disparities and an accuracy inversion: relative to a 35-44 comparator, claimants aged 55-65 were about 2.80 times as likely to be referred for review and non-UK nationals about 2.27 times as likely, while for older claimants those referrals were less likely to be correct (relative correct-referral likelihoods of about 0.58 at 55-65 and 0.23 at 66-plus, the latter resting on a small sub-sample DWP flags to treat with caution). The disparities were first disclosed under freedom-of-information law and reported in December 2024, and DWP has committed to retrain the model. The figures are DWP-reported relative ratios, not independently audited absolute error rates.

departmentforworkandpensions2025cGroundingGovernmentSave

Department for Work and Pensions, Fraudsters face tougher action as Government gains new powers to tackle benefit fraud (Public Authorities (Fraud, Error and Recovery) Act 2025) (2025) https://www.gov.uk/government/news/fraudsters-face-tougher-action-as-government-gains-new-powers-to-tackle-benefit-fraud

https://www.gov.uk/government/news/fraudsters-face-tougher-action-as-government-gains-new-powers-to-tackle-benefit-fraud

Grounds: model org: uk_dwp_uca_fraud

EmpiricalDWP states that a human caseworker always makes the final decision on a referred Universal Credit advance with…

DWP states that a human caseworker always makes the final decision on a referred Universal Credit advance with no automated decision-making, and is deliberately not shown the risk score or told the referral came from the model; DWP describes the model as around three times more effective than a randomised control at identifying fraud risk and judges continued operation reasonable and proportionate while committing to retrain it. The Public Law Project counters that only age was fully assessed among protected characteristics and that the assessment relied on safeguards preventing downstream harm rather than showing the model to be non-discriminatory. The wider counter-fraud programme is meanwhile expanding into bank-data eligibility verification under the Public Authorities (Fraud, Error and Recovery) Act 2025, a distinct system not yet in force.

departmentforworkandpensions2025cGroundingGovernmentSave

Department for Work and Pensions, Fraudsters face tougher action as Government gains new powers to tackle benefit fraud (Public Authorities (Fraud, Error and Recovery) Act 2025) (2025) https://www.gov.uk/government/news/fraudsters-face-tougher-action-as-government-gains-new-powers-to-tackle-benefit-fraud

https://www.gov.uk/government/news/fraudsters-face-tougher-action-as-government-gains-new-powers-to-tackle-benefit-fraud

Grounds: model org: uk_dwp_uca_fraud

publiclawproject2025GroundingAdvocacySave

Public Law Project, Written evidence to the Public Accounts Committee on tackling fraud and error in benefit expenditure (FAE0006) (2025) https://committees.parliament.uk/writtenevidence/152681/pdf/

https://committees.parliament.uk/writtenevidence/152681/pdf/

Grounds: model org: uk_dwp_uca_fraud

EmpiricalDuring the pandemic unemployment surge, a private facial-recognition identity check operated as a de facto eli…

During the pandemic unemployment surge, a private facial-recognition identity check operated as a de facto eligibility gate for unemployment benefits in at least 25 U.S. state workforce agencies, with a live 'trusted referee' interview queue that House investigators documented averaging nearly 10 hours in North Dakota and over 4 hours in 14 of 21 states, versus about 6 minutes in New Jersey where an in-person option existed. Oregon's own one-month study (n=10,656 routed) recorded verification-completion differences by group -- for example 41.59% for African American and 34.48% for Spanish-language claimants versus 53.44% for White claimants -- but stated the study showed differences in completion and did not show causation, so these are a friction proxy, not a measured wrongful-denial rate. The U.S. Department of Labor does not collect or report the number of workers blocked for inability to verify identity, and where verification precedes filing those workers are not counted as denied claims at all, so the scale of any wrongful lockout is undocumented.

ushousecommitteeonoversighta2022aGroundingGovernmentSave

U.S. House Committee on Oversight and Reform, Chairs Clyburn, Maloney Release Evidence Facial Recognition Company ID.me Downplayed Excessive Wait Times for Americans Seeking Unemployment Relief Funds (2022) https://oversightdemocrats.house.gov/news/press-releases/chairs-maloney-clyburn-release-evidence-facial-recognition-company-idme

https://oversightdemocrats.house.gov/news/press-releases/chairs-maloney-clyburn-release-evidence-facial-recognition-company-idme

Grounds: model org: us_idme_identity_gate

usdepartmentoflabor2023cGroundingGovernment evaluationSave

U.S. Department of Labor, Office of Inspector General, Alert Memorandum: ETA and States Need to Ensure the Use of Identity Verification Service Contractors Results in Equitable Access to UI Benefits and Secure Biometric Data (Report No. 19-23-005-03-315) (2023) https://www.oig.dol.gov/public/reports/oa/2023/19-23-005-03-315.pdf

https://www.oig.dol.gov/public/reports/oa/2023/19-23-005-03-315.pdf

Grounds: model org: us_idme_identity_gate

EmpiricalA U.S. Department of Labor Inspector General audit (March 31, 2023) found that among 24 state workforce agenci…

A U.S. Department of Labor Inspector General audit (March 31, 2023) found that among 24 state workforce agencies using a facial-recognition identity contractor, 18 of 24 (75%) contracts did not specify one-to-one versus one-to-many matching, 15 of 24 (63%) did not address data storage, and 13 of 24 (54%) did not address destruction of the collected biometric data, while 22 of 24 (92%) agencies reported the technology reduced improper payments -- the operator-side benefit that sustained adoption even as the wrongful-lockout cost went unmeasured. The vendor initially represented it used only one-to-one matching and later acknowledged one-to-many matching against a database; after bipartisan backlash the IRS and Treasury dropped the mandatory facial-recognition requirement in February 2022 and the vendor made it optional across agencies, though the service remained in use for unemployment identity verification in a large share of states, and a 2026 IRS proposal would allow it to retain taxpayer biometric data up to 36 months after account deletion. The reported improper-payment reductions are agency self-reports, not independently audited.

usdepartmentoflabor2023cGroundingGovernment evaluationSave

U.S. Department of Labor, Office of Inspector General, Alert Memorandum: ETA and States Need to Ensure the Use of Identity Verification Service Contractors Results in Equitable Access to UI Benefits and Secure Biometric Data (Report No. 19-23-005-03-315) (2023) https://www.oig.dol.gov/public/reports/oa/2023/19-23-005-03-315.pdf

https://www.oig.dol.gov/public/reports/oa/2023/19-23-005-03-315.pdf

Grounds: model org: us_idme_identity_gate

americancivillibertiesunionj2022GroundingAdvocacySave

American Civil Liberties Union (Jay Stanley and Olga Akselrod), Three Key Problems with the Government's Use of a Flawed Facial Recognition Service (2022) https://www.aclu.org/news/privacy-technology/three-key-problems-with-the-governments-use-of-a-flawed-facial-recognition-service

https://www.aclu.org/news/privacy-technology/three-key-problems-with-the-governments-use-of-a-flawed-facial-recognition-service

Grounds: model org: us_idme_identity_gate

Topics: privacy-security

EmpiricalNevada's Department of Employment, Training and Rehabilitation contracted Google to build a generative-AI tool…

Nevada's Department of Employment, Training and Rehabilitation contracted Google to build a generative-AI tool on the Vertex AI Studio cloud platform that reads an unemployment-appeal hearing transcript and evidence, retrieves against a corpus of Nevada unemployment law and prior appeals decisions, and drafts a recommended determination (approve, deny, or modify a claim) together with the written decision for a human referee to review and sign. The contract set a 90 percent success requirement self-assessed by state workers on test decisions -- not an independent external audit -- and DETR said it wanted accuracy higher than 90 percent before going live; rollout was repeatedly delayed over less-than-desired accuracy, including the tool citing incorrect Nevada statutes and failing to pull information from all hearing documents, problems officials said were fixed. Reported cost evolved from about 1 million dollars in 2024 to a total of 2.6 million dollars with about 1.1 million spent by early 2026. As of the most recent available reporting (March 2026) the system was in delayed pre-deployment testing on historical appeals and described as launching in coming weeks; it was not independently confirmed to be adjudicating live claimant appeals.

thenevadaindependent2025GroundingInvestigativeSave

The Nevada Independent (2025, July 22), Nevada will use AI for unemployment appeals; some lawmakers are skeptical (DETR / Google) https://thenevadaindependent.com/article/nevada-will-use-ai-for-unemployment-appeals-some-lawmakers-are-skeptical

https://thenevadaindependent.com/article/nevada-will-use-ai-for-unemployment-appeals-some-lawmakers-are-skeptical

Appears in: PAN framework development

Grounds: capability governance: routing expansion; model org: nevada_detr_genai_appeals

fordhamintellectualproperty2024GroundingAcademicSave

Fordham Intellectual Property, Media and Entertainment Law Journal (Dawn Edelman), Speed, Accuracy, and Risk: Nevada's Use of Artificial Intelligence in Unemployment Claims Appeals (2024) http://www.fordhamiplj.org/2024/10/07/speed-accuracy-and-risk-nevadas-use-of-artificial-intelligence-in-unemployment-claims-appeals/

http://www.fordhamiplj.org/2024/10/07/speed-accuracy-and-risk-nevadas-use-of-artificial-intelligence-in-unemployment-claims-appeals/

Grounds: model org: nevada_detr_genai_appeals

EmpiricalNevada's generative-AI unemployment-appeals tool was justified as a speed measure for a pandemic-era backlog, …

Nevada's generative-AI unemployment-appeals tool was justified as a speed measure for a pandemic-era backlog, projecting a drop in referee determination time from as much as several hours to about five minutes per case, with a mandatory human review DETR said adds an estimated 10 to 30 minutes and a required referee sign-off (Director Christopher Sewell said no AI-drafted written decisions issue without human review). Legal scholars, attorneys who represent claimants, and a former U.S. Department of Labor official warned that backlog and speed pressure could hollow out that review and create incentives to rubber-stamp AI outputs -- one attorney noting the time savings only happens if the review is very cursory, and a legal analysis warning staff might feel pressured to authorize AI decisions with haste. That automation-deference risk is expert-projected, not a measured outcome: no referee override or rejection rate has been published, and claimants are not required to consent to AI processing of their appeal.

fordhamintellectualproperty2024GroundingAcademicSave

Fordham Intellectual Property, Media and Entertainment Law Journal (Dawn Edelman), Speed, Accuracy, and Risk: Nevada's Use of Artificial Intelligence in Unemployment Claims Appeals (2024) http://www.fordhamiplj.org/2024/10/07/speed-accuracy-and-risk-nevadas-use-of-artificial-intelligence-in-unemployment-claims-appeals/

http://www.fordhamiplj.org/2024/10/07/speed-accuracy-and-risk-nevadas-use-of-artificial-intelligence-in-unemployment-claims-appeals/

Grounds: model org: nevada_detr_genai_appeals

thenevadaindependent2025GroundingInvestigativeSave

The Nevada Independent (2025, July 22), Nevada will use AI for unemployment appeals; some lawmakers are skeptical (DETR / Google) https://thenevadaindependent.com/article/nevada-will-use-ai-for-unemployment-appeals-some-lawmakers-are-skeptical

https://thenevadaindependent.com/article/nevada-will-use-ai-for-unemployment-appeals-some-lawmakers-are-skeptical

Appears in: PAN framework development

Grounds: capability governance: routing expansion; model org: nevada_detr_genai_appeals

EmpiricalThe Targeted Real-Time Early Warning System (TREWS), a machine-learning sepsis early-warning model, was evalua…

The Targeted Real-Time Early Warning System (TREWS), a machine-learning sepsis early-warning model, was evaluated prospectively across five hospitals of an academic health system covering 590,736 monitored patients — the largest prospective study of an ML sepsis system on record. Its central finding was conditional on the human loop: sepsis patients whose alert was evaluated and confirmed by a provider within three hours had a 3.3 percentage-point absolute and 18.7 percent relative adjusted reduction in in-hospital mortality, with less organ failure and shorter stays, while the alert on its own did not; a companion study found provider uptake varied with experience, unit culture, and alert context.

adams2022aGroundingPeer-reviewedSave

Adams, R., Henry, K.E., et al. (2022). Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nature Medicine, 28(7), 1455-1460. https://doi.org/10.1038/s41591-022-01894-0 https://www.nature.com/articles/s41591-022-01894-0

doi.org/10.1038/s41591-022-01894-0

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage)

henry2022aGroundingPeer-reviewedSave

Henry, K.E., et al. (2022). Factors driving provider adoption of the TREWS machine learning-based early warning system and its effects on sepsis treatment timing. Nature Medicine, 28. https://doi.org/10.1038/s41591-022-01895-z https://www.nature.com/articles/s41591-022-01895-z

doi.org/10.1038/s41591-022-01895-z

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage)

EmpiricalThe TREWS mortality-benefit evaluation was prospective and peer-reviewed but observational and developer-led: …

The TREWS mortality-benefit evaluation was prospective and peer-reviewed but observational and developer-led: it was built at the deploying institution and commercialized through a company founded by its principal investigator, and confirmation-associated benefit is an observational association rather than a randomized effect of the algorithm — providers who engaged with alerts may differ from those who did not in ways the adjustment does not capture. The strongest numbers in the record therefore come from the party with the strongest interest in them, and no independent replication of the mortality effect had been published.

adams2022aGroundingPeer-reviewedSave

Adams, R., Henry, K.E., et al. (2022). Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nature Medicine, 28(7), 1455-1460. https://doi.org/10.1038/s41591-022-01894-0 https://www.nature.com/articles/s41591-022-01894-0

doi.org/10.1038/s41591-022-01894-0

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage)

EmpiricalThe Advance Alert Monitor is an in-hospital deterioration model running around the clock across 21 hospitals o…

The Advance Alert Monitor is an in-hospital deterioration model running around the clock across 21 hospitals of an integrated health system, scoring inpatients hourly and firing roughly twelve hours before predicted deterioration; a 2020 New England Journal of Medicine evaluation associated its alert-driven rapid-response workflow with lower mortality. Its defining feature is where the alert goes: not to the bedside, but to a dedicated regional tier of critical-care virtual quality nurse consultants who screen every alert around the clock, work up the chart, and only then escalate to the on-site rapid-response team — so the measured benefit is priced against the whole two-tier staffing topology, not the model alone.

escobar2020aGroundingAcademicSave

Escobar, G.J., Liu, V.X., Schuler, A., Lawson, B., Greene, J.D., & Kipnis, P. (2020). Automated Identification of Adults at Risk for In-Hospital Clinical Deterioration. New England Journal of Medicine, 383(20), 1951-1960. https://doi.org/10.1056/NEJMsa2001090 https://www.nejm.org/doi/full/10.1056/NEJMsa2001090

doi.org/10.1056/NEJMsa2001090

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage); model org: kaiser_aam_deterioration

thekaiserpermanentenorthernc2022GroundingAcademicSave

The Kaiser Permanente Northern California Advance Alert Monitor Program: An Automated Early Warning System for Adults at Risk for In-Hospital Clinical Deterioration (2022). Joint Commission Journal on Quality and Patient Safety. https://www.jointcommissionjournal.com/article/S1553-7250(22)00110-6/fulltext

https://www.jointcommissionjournal.com/article/S1553-7250(22)00110-6/fulltext

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage); model org: kaiser_aam_deterioration

Topics: ai-safety

EmpiricalSepsis Watch is a deep-learning sepsis-detection system scoring every emergency-department patient every five …

Sepsis Watch is a deep-learning sepsis-detection system scoring every emergency-department patient every five minutes over 86 variables, deployed at an academic hospital under a registered clinical trial, with alerts fronted by rapid-response-team nurses who track treatment-bundle completion on three- and six-hour timers. Its structural fault line is an authority split: the operator who receives the alert (the nurse) is not the operator empowered to act on it (the physician who holds treatment authority), so the correction runs through a peer-persuasion edge. An independent ethnography found the system worked because nurses performed hidden repair work — mediating the professional hierarchy and doing the emotional labor of communicating a risk score upward — labor that was structurally necessary, largely invisible to the deployment's formal description, and undervalued.

sendak2020aGroundingPeer-reviewedSave

Sendak, M.P., et al. (2020). Real-World Integration of a Sepsis Deep Learning Technology Into Routine Clinical Care: Implementation Study. JMIR Medical Informatics, 8(7), e15182. https://doi.org/10.2196/15182 https://medinform.jmir.org/2020/7/e15182/

doi.org/10.2196/15182

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage)

elish2020GroundingAdvocacySave

Elish, M.C., & Watkins, E.A. (2020). Repairing Innovation: A Study of Integrating AI in Clinical Care. Data & Society Research Institute. https://datasociety.net/library/repairing-innovation/

https://datasociety.net/library/repairing-innovation/

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage); model org: duke_sepsis_watch

EmpiricalA widely implemented proprietary sepsis-prediction model shipped inside a common electronic-health-record plat…

A widely implemented proprietary sepsis-prediction model shipped inside a common electronic-health-record platform and switched on across hundreds of hospitals was externally validated in 2021 across 38,455 hospitalizations at an academic health system: it achieved an area under the curve of 0.63, identified only 33 percent of sepsis cases, and had a positive predictive value of about 12 percent, generating roughly 109 alerts for every true sepsis case — a real-world performance the vendor had not fully examined before selling the model, and which an investigation attributed in part to undisclosed features such as antibiotic-order data that inflated internal validation.

wong2021cGroundingAcademicSave

Wong, A., Otles, E., et al. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307

https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2781307

Appears in: PAN framework development

Grounds: domain grounding: clinical AI (deterioration, imaging, documentation); model org: epic_sepsis_model_michigan

Topics: complexity-science

statnews2021GroundingInvestigativeSave

STAT News (2021, July 26). Epic's AI algorithms, shielded from scrutiny by a corporate firewall, are delivering inaccurate information on seriously ill patients. https://www.statnews.com/2021/07/26/epic-hospital-algorithms-sepsis-investigation/

https://www.statnews.com/2021/07/26/epic-hospital-algorithms-sepsis-investigation/

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage); model org: epic_sepsis_model_michigan

EmpiricalAfter external criticism, the vendor overhauled the sepsis model — retraining it, changing the sepsis-onset de…

After external criticism, the vendor overhauled the sepsis model — retraining it, changing the sepsis-onset definition, and reducing its reliance on antibiotic-order features. A 2026 multicenter prospective validation of the updated model across 227,091 encounters reported an area under the curve of 0.82 to 0.92 with positive predictive value of 0.13 to 0.26 and substantial between-site variability, and its authors urged local validation and alert-silencing strategies rather than trusting the model out of the box — a correction that arrived only after independent scrutiny of a model that had already been deployed at scale behind a corporate firewall shielding it from outside inspection.

statnews2022GroundingInvestigativeSave

STAT News (2022, Oct 3). Epic overhauls popular sepsis algorithm criticized for faulty alarms. https://www.statnews.com/2022/10/03/epic-sepsis-algorithm-revamp-training/

https://www.statnews.com/2022/10/03/epic-sepsis-algorithm-revamp-training/

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage); model org: epic_sepsis_model_michigan

wong2026aGroundingPeer-reviewedSave

Wong, A., Currey, D., Schwinne, M., et al. (2026). Multicenter Prospective Validation of an Updated Proprietary Sepsis Prediction Model. JAMA Network Open. https://doi.org/10.1001/jamanetworkopen.2026.0181 https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2845595

doi.org/10.1001/jamanetworkopen.2026.0181

Appears in: PAN framework development

Grounds: domain grounding: clinical decision support (sepsis/deterioration alerting, imaging triage)

Topics: complexity-science

EmpiricalThe largest documented ambient-scribe deployment ran a 10-week pilot at an integrated medical group and then s…

The largest documented ambient-scribe deployment ran a 10-week pilot at an integrated medical group and then scaled to 7,260 physicians and 2,576,627 patient encounters over fourteen months, with roughly 16,000 hours of documentation time saved and sustained physician support measured along the way. The system records the visit and drafts the clinical note; the clinician edits and signs, and the model-to-record write is gated both by that clinician review and by a standing internal quality-assurance program over the AI output — a real subsystem with a real cost, because the drafted note becomes a permanent record that later clinicians and later tools read as fact.

tierney2024aGroundingAcademicSave

Tierney, A.A., Gayre, G., Hoberman, B., et al. (2024). Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catalyst Innovations in Care Delivery. https://doi.org/10.1056/CAT.23.0404 https://catalyst.nejm.org/doi/full/10.1056/CAT.23.0404

doi.org/10.1056/CAT.23.0404

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding); model org: kaiser_tpmg_ambient_scribe

tierney2025aGroundingAcademicSave

Tierney, A.A., et al. (2025). Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses. NEJM Catalyst Innovations in Care Delivery. https://doi.org/10.1056/CAT.25.0040 https://divisionofresearch.kaiserpermanente.org/ai-assisted-notetaking-gains-steady-support-from-kaiser-permanente-physicians/

doi.org/10.1056/CAT.25.0040

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding); model org: kaiser_tpmg_ambient_scribe

EmpiricalThe scale numbers from a single ambient-scribe deployment are the deployer's own first-party measurements and …

The scale numbers from a single ambient-scribe deployment are the deployer's own first-party measurements and should be read as that system's dashboard rather than a guarantee of the product class: a multisite study of 8,581 clinicians across five health systems found more modest effects — on the order of 13 to 16 fewer minutes per day with no meaningful after-hours relief — and a validated per-note evaluation found hallucinations in about 31 percent of ambient-generated notes under structured review, versus about 20 percent of physician-written gold-standard notes, making ambient notes more thorough but less accurate. The clinician review and quality-assurance program are the controls that stand between that error rate and a contaminated permanent record.

rotenstein2026aGroundingPeer-reviewedSave

Rotenstein, L.S., et al. (2026). Changes in Clinician Time Expenditure and Visit Quantity With Adoption of Artificial Intelligence-Powered Scribes: A Multisite Study. JAMA. https://doi.org/10.1001/jama.2026.2253 https://pubmed.ncbi.nlm.nih.gov/41920565/

doi.org/10.1001/jama.2026.2253

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)

palm2025aGroundingPeer-reviewedSave

Palm, K.H., Manikantan, K., Mahal, N., Belwadi, S.K., & Pepin, R.J. (2025). Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe. Frontiers in Artificial Intelligence, 8. https://doi.org/10.3389/frai.2025.1691499 https://pmc.ncbi.nlm.nih.gov/articles/PMC12586549/

doi.org/10.3389/frai.2025.1691499

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)

EmpiricalThe strongest causal evidence in the ambient-scribe family comes from a 24-week stepped-wedge, individually ra…

The strongest causal evidence in the ambient-scribe family comes from a 24-week stepped-wedge, individually randomized trial of an ambient scribe across 66 practitioners and 71,487 notes (38 percent AI-generated), which found work exhaustion significantly reduced, professional fulfillment unchanged (a recorded null), roughly 22 minutes per day of documentation time returned, and diagnostic coding accuracy improved. Unlike the larger first-party deployment reports, this is a randomized estimate of the well-being and time effects — though it measured practitioner well-being and time, not per-note error rates.

afshar2025aGroundingPeer-reviewedSave

Afshar, M., Baumann, M.R., Resnik, F., et al. (2025). A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being. NEJM AI. https://doi.org/10.1056/AIoa2500945 https://pubmed.ncbi.nlm.nih.gov/41625485/

doi.org/10.1056/AIoa2500945

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)

EmpiricalThe same team that ran the ambient-scribe trial released an open operations playbook for safety and effectiven…

The same team that ran the ambient-scribe trial released an open operations playbook for safety and effectiveness monitoring of ambient AI in production — the rare case where the organization-side monitoring function exists as a citable, designed subsystem rather than an assumed practice. That monitoring is the org's stated answer to a documented system-level risk of the technology: a coding arms race, in which better AI documentation raises coding intensity, payers recalibrate in response, and clinician attestation liability grows — so the improved coding accuracy the trial measured sits next door to an upcoding pressure the monitoring is meant to watch.

afshar2025bGroundingAcademicSave

Afshar, M., et al. (2025). A Novel Playbook for Pragmatic Trial Operations to Monitor and Evaluate Ambient Artificial Intelligence in Clinical Practice. NEJM AI. https://doi.org/10.1056/AIdbp2401267 https://ai.nejm.org/doi/full/10.1056/AIdbp2401267

doi.org/10.1056/AIdbp2401267

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding); model org: uw_health_abridge_scribe

dai2025aGroundingPeer-reviewedSave

Dai, T., Kvedar, J.C., & Polsky, D. (2025). Policy brief: ambient AI scribes and the coding arms race. npj Digital Medicine. https://doi.org/10.1038/s41746-025-02272-z https://www.nature.com/articles/s41746-025-02272-z

doi.org/10.1038/s41746-025-02272-z

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)

EmpiricalA peer-reviewed evaluation of an ambient documentation platform at a large multi-specialty system found note t…

A peer-reviewed evaluation of an ambient documentation platform at a large multi-specialty system found note time per appointment reduced (6.2 to 5.3 minutes) and NASA-TLX cognitive load reduced — but the burnout change (42.1 to 35.1 percent) was not statistically significant, the domain's honest null bound of cognitive-load relief without a demonstrated burnout effect.

stults2025aGroundingPeer-reviewedSave

Stults, C.D., Deng, S., Martinez, M.C., et al. (2025). Evaluation of an Ambient Artificial Intelligence Documentation Platform for Clinicians. JAMA Network Open, 8(5), e258614. https://doi.org/10.1001/jamanetworkopen.2025.8614 https://pubmed.ncbi.nlm.nih.gov/40314951/

doi.org/10.1001/jamanetworkopen.2025.8614

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)

Topics: complexity-science

EmpiricalIn the same evaluation, benefit varied sharply by clinician group: 85.8 percent of primary-care physicians rep…

In the same evaluation, benefit varied sharply by clinician group: 85.8 percent of primary-care physicians reported improved satisfaction against 36.4 percent of medical specialists — the same tool, in the same system, under the same workflow, helping one operator class and largely failing another, so any uniform service term overstates the effect for the group it helps least.

stults2025aGroundingPeer-reviewedSave

Stults, C.D., Deng, S., Martinez, M.C., et al. (2025). Evaluation of an Ambient Artificial Intelligence Documentation Platform for Clinicians. JAMA Network Open, 8(5), e258614. https://doi.org/10.1001/jamanetworkopen.2025.8614 https://pubmed.ncbi.nlm.nih.gov/40314951/

doi.org/10.1001/jamanetworkopen.2025.8614

Appears in: PAN framework development

Grounds: domain grounding: clinical documentation copilots (ambient scribes, note generation, coding)

Topics: complexity-science

EmpiricalCompany-run, pre-registered, peer-reviewed randomized rollouts of a commercial code-completion assistant acros…

Company-run, pre-registered, peer-reviewed randomized rollouts of a commercial code-completion assistant across 4,867 developers at three enterprises found a pooled 26.08 percent increase in completed tasks, with gains concentrated among less-experienced developers. An independent randomized study of 16 experienced open-source maintainers on 246 tasks in familiar repositories bounded the expert tail from the other direction: those developers were about 19 percent slower with the AI while believing themselves about 20 percent faster — a measured perception-reality gap that means a uniform productivity number overstates the effect for senior engineers.

cui2025aGroundingAcademicSave

Cui, Z.K., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2025). The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science. https://doi.org/10.1287/mnsc.2025.00535 https://pubsonline.informs.org/doi/10.1287/mnsc.2025.00535

doi.org/10.1287/mnsc.2025.00535

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review); model org: msft_accenture_copilot_experiments

becker2025aGroundingIndustrySave

Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR. https://doi.org/10.48550/arXiv.2507.09089 https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

doi.org/10.48550/arXiv.2507.09089

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

EmpiricalIndividual coding-assistant gains do not automatically compose to organization-level delivery outcomes: a cros…

Individual coding-assistant gains do not automatically compose to organization-level delivery outcomes: a cross-industry research program measured a roughly 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability for every 25 percent increase in AI adoption, evidence that the churn the assistant adds must be absorbed by code-review and testing gates or the individual speed-up degrades the organization's delivery performance.

googleclouddora2024GroundingReferenceSave

Google Cloud DORA (2024). Accelerate State of DevOps Report 2024. https://dora.dev/research/2024/dora-report/

https://dora.dev/research/2024/dora-report/

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review); model org: msft_accenture_copilot_experiments

EmpiricalAn in-house machine-learning code-completion system built, deployed, and measured by a company's own platform …

An in-house machine-learning code-completion system built, deployed, and measured by a company's own platform organization for more than 10,000 internal developers reported, against a control group, a 25 to 34 percent suggestion-acceptance rate, a 6 percent reduction in coding iteration time versus control, and 3 percent of new code characters coming from the model at the time of measurement. The measuring party, the building party, and the deploying party were the same organization, and the numbers were published as an engineering-blog self-report rather than a peer-reviewed or independent evaluation.

tabachnyk2022GroundingVendorSave

Tabachnyk, M., & Nikolov, S. (2022). ML-Enhanced Code Completion Improves Developer Productivity. Google Research Blog. https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/

https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review); model org: google_internal_code_completion

EmpiricalThe same company's cross-industry research program reported that AI-assisted software development amplifies an…

The same company's cross-industry research program reported that AI-assisted software development amplifies an organization's existing strengths and weaknesses rather than substituting for them, with policy clarity and platform investment identified as the levers that determine whether AI adoption improves or degrades delivery — evidence that the individual coding gains do not compose to organization-level outcomes on their own, and that the deploying organization's existing gates and platform quality are what decide the result.

googleclouddora2025GroundingReferenceSave

Google Cloud DORA (2025). State of AI-assisted Software Development (2025 DORA Report). https://dora.dev/dora-report-2025/

https://dora.dev/dora-report-2025/

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review); model org: google_internal_code_completion

EmpiricalA regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before sc…

A regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before scaling a commercial coding assistant to roughly 1,000 engineers, publishing its own measurement of the rollout. The bank's engineers reported productivity and code-quality improvements — and recorded the security impact as explicitly inconclusive, a real gating decision taken and documented under uncertainty rather than resolved by assertion, with the honestly recorded unknown carried forward into the scaled deployment.

chatterjee2024aGroundingIndustrySave

Chatterjee, S., Liu, C.L., Rowland, G., & Hogarth, T. (2024). The Impact of AI Tool on Engineering at ANZ Bank: An Empirical Study on GitHub Copilot within Corporate Environment [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2402.05636 https://www.theregister.com/2024/02/10/anz_bank_github_copilot/

doi.org/10.48550/arXiv.2402.05636

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

EmpiricalWhat the bank's inconclusive security finding leaves open is not hypothetical: an independent security assessm…

What the bank's inconclusive security finding leaves open is not hypothetical: an independent security assessment of code generated by a widely used assistant found that about 40 percent of generated programs contained vulnerabilities across scenarios spanning the CWE top-25 weaknesses, and separate research documents developers accepting insecure suggestions with overconfidence — so the security unknown a deployment carries forward unresolved sits against a class-level literature in which insecure generation is common.

pearce2022aGroundingPeer-reviewedSave

Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., & Karri, R. (2022). Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions. In 43rd IEEE Symposium on Security and Privacy (SP 2022). https://doi.org/10.48550/arXiv.2108.09293 https://arxiv.org/abs/2108.09293

doi.org/10.48550/arXiv.2108.09293

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

Topics: privacy-security

EmpiricalA mid-size enterprise ran a systematic four-phase evaluation-to-rollout of a commercial coding assistant acros…

A mid-size enterprise ran a systematic four-phase evaluation-to-rollout of a commercial coding assistant across more than 400 developers, publishing acceptance telemetry (a 33 percent suggestion-acceptance rate, with 20 percent of suggested lines accepted), a 72 percent satisfaction figure, documented per-language variation, and stated limitations. Its evaluation instrument is acceptance-rate telemetry — which the productivity literature identifies as the measure most correlated with perceived productivity rather than outcome, and perception is measured to be miscalibrated for experienced developers, so acceptance telemetry captures adoption feel, not delivered output.

bakal2025aGroundingIndustrySave

Bakal, G., Dasdan, A., Katz, Y., Kaufman, M., & Levin, G. (2025). Experience with GitHub Copilot for Developer Productivity at Zoominfo [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2501.13282 https://arxiv.org/abs/2501.13282

doi.org/10.48550/arXiv.2501.13282

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

ziegler2024aGroundingAcademicSave

Ziegler, A., Kalliamvakou, E., Li, X.A., et al. (2024). Measuring GitHub Copilot's Impact on Productivity. Communications of the ACM, 67(3). https://doi.org/10.1145/3633453 https://dl.acm.org/doi/10.1145/3633453

doi.org/10.1145/3633453

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review); model org: zoominfo_copilot_deployment

EmpiricalThe deployment report stated its limitations but reported no security evaluation at all — an unrecorded unknow…

The deployment report stated its limitations but reported no security evaluation at all — an unrecorded unknown, one step less honest than a deployment that runs a security check and records the result as inconclusive, because an absence no one has written down is not a governed object and cannot be carried forward or resolved. The value of the case is the documentation quality of an ordinary, competent adoption — phase gates, telemetry definitions, per-language deltas, and stated limitations by the deployer itself — with the missing security question priced as the one thing even that documentation did not name.

bakal2025aGroundingIndustrySave

Bakal, G., Dasdan, A., Katz, Y., Kaufman, M., & Levin, G. (2025). Experience with GitHub Copilot for Developer Productivity at Zoominfo [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2501.13282 https://arxiv.org/abs/2501.13282

doi.org/10.48550/arXiv.2501.13282

Appears in: PAN framework development

Grounds: domain grounding: software engineering AI (coding assistants, code review)

EmpiricalA global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-la…

A global bank replaced rules-based transaction monitoring with a cloud vendor's machine-learning anti-money-laundering product as its primary monitoring system in key markets, reporting two to four times more confirmed suspicious activity with roughly 60 percent fewer alerts. Every one of those numbers is a vendor-and-customer self-report with no independent audit — which is itself the honest structure of the domain, because a peer-reviewed deployment-scale benefit measurement inside a named financial-crime operation does not publicly exist, and the alert-volume reduction the vendor advertises is precisely the lever a regulator scrutinizing an under-monitoring risk would question.

googlecloud2023GroundingVendorSave

Google Cloud (2023, June 21). Google Cloud Launches AI-Powered Anti Money Laundering Product for Financial Institutions (with HSBC-reported results). https://www.googlecloudpresscorner.com/2023-06-21-Google-Cloud-Launches-AI-Powered-Anti-Money-Laundering-Product-for-Financial-Institutions

https://www.googlecloudpresscorner.com/2023-06-21-Google-Cloud-Launches-AI-Powered-Anti-Money-Laundering-Product-for-Financial-Institutions

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring); model org: hsbc_aml_ai

EmpiricalTwo structural dynamics govern fraud and financial-crime detection. Under extreme base rates, detection precis…

Two structural dynamics govern fraud and financial-crime detection. Under extreme base rates, detection precision is dominated by the false-alarm rate rather than by accuracy, so at realistic prevalence a threshold change moves the burden of alerts rather than the truth of them (the base-rate fallacy). And the labels the model learns from are the investigators' own dispositions: only a small set of flagged transactions is ever verified, and models are retrained on the analysts' calls, so a rise in 'confirmed' activity is partly a measure of what the system taught its reviewers to confirm rather than an independent ground truth (the label-feedback loop).

axelsson2000aGroundingAcademicSave

Axelsson, S. (2000). The Base-Rate Fallacy and the Difficulty of Intrusion Detection. ACM Transactions on Information and System Security, 3(3), 186-205. https://doi.org/10.1145/357830.357849 https://dl.acm.org/doi/10.1145/357830.357849

doi.org/10.1145/357830.357849

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring); model org: hsbc_aml_ai

dalpozzolo2018aGroundingPeer-reviewedSave

Dal Pozzolo, A., Boracchi, G., Caelen, O., Alippi, C., & Bontempi, G. (2018). Credit Card Fraud Detection: A Realistic Modeling and a Novel Learning Strategy. IEEE Transactions on Neural Networks and Learning Systems, 29(8), 3784-3797. https://doi.org/10.1109/TNNLS.2017.2736643 https://dalpozz.github.io/static/pdf/TNNLS_2017.pdf

doi.org/10.1109/TNNLS.2017.2736643

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring)

Topics: complexity-science

EmpiricalA neobank's fraud algorithms — triggered heavily by pandemic-era government benefit deposits — froze and close…

A neobank's fraud algorithms — triggered heavily by pandemic-era government benefit deposits — froze and closed the accounts of legitimate customers at scale, holding their balances for thirty to more than ninety days, and the company admitted some of the closures were mistakes. The false-positive tail here lands on real people as immediate hardship, concentrated among benefit-deposit recipients and low-balance households for whom a frozen account means no access to funds for weeks.

kessler2021GroundingInvestigativeSave

Kessler, C. (2021, July 6). A Banking App Has Been Suddenly Closing Accounts, Sometimes Not Returning Customers' Money. ProPublica. https://www.propublica.org/article/chime

https://www.propublica.org/article/chime

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring); model org: chime_fraud_pipeline

EmpiricalA Nordic bank's rules-based legacy fraud system ran at roughly 40 percent detection with a 99.5 percent false-…

A Nordic bank's rules-based legacy fraud system ran at roughly 40 percent detection with a 99.5 percent false-positive rate — a measured pre-machine-learning baseline whose badness is the most credible datum in the record, since a 99.5 percent false-positive rate is not a marketing claim. The vendor-published rollout of a deep-learning engine scoring transactions in real time (under 300 milliseconds) claims false positives cut by about 60 percent and true-positive detection raised by about 50 percent; those figures are an organization-named, trade-press-covered vendor case study, entered here as claimed magnitudes against that legacy baseline because they were not independently audited.

teradata2017GroundingVendorSave

Teradata (2017). Danske Bank Fights Fraud with Deep Learning and AI (case study EB9821). https://assets.teradata.com/resourceCenter/downloads/CaseStudies/CaseStudy_EB9821_Danske_Bank_Saves_Millions_Fighting_Fraud_With_Deep_Learning_and_AI.pdf

https://assets.teradata.com/resourceCenter/downloads/CaseStudies/CaseStudy_EB9821_Danske_Bank_Saves_Millions_Fighting_Fraud_With_Deep_Learning_and_AI.pdf

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring); model org: danske_fraud_engine

EmpiricalThe same institution that improved its in-line fraud scoring later ranked worst among UK banks for reimbursing…

The same institution that improved its in-line fraud scoring later ranked worst among UK banks for reimbursing victims of authorized-push-payment scams in the regulator's bank-by-bank performance data — better detection and worse victim-outcome performance coexisting in one organization. And it was a rule, not a model, that moved the institutional behavior: the regulator's mandatory-reimbursement regime raised sector reimbursement from roughly two-thirds to about 89 percent, demonstrating that detection quality and the justice of the disposition are different levers held by different actors, and that the victim-outcome lever is a regulatory rule rather than a better classifier.

ukpaymentsystemsregulator2023GroundingGovernmentSave

UK Payment Systems Regulator (2023-2025). APP fraud performance data / APP scams performance reports. https://www.psr.org.uk/information-for-consumers/app-fraud-performance-data/

https://www.psr.org.uk/information-for-consumers/app-fraud-performance-data/

Appears in: PAN framework development

Grounds: domain grounding: security operations and fraud detection (SOC triage, fraud scoring); model org: danske_fraud_engine

EmpiricalAn internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five star…

An internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five stars per role and location — trained on ten years of the company's own hiring decisions, a period whose hires were predominantly male. The models learned that history: they penalized the word 'women's' and downgraded graduates of women's colleges, reading gender proxies as negative signal. The team patched the identified terms but concluded that term-level fixes could not guarantee neutrality against unknown proxies, because the model had learned the pattern rather than the words, and the company scrapped the project around 2017; per the company, recruiters saw the tool's recommendations but it was never used as a sole ranking.

dastin2018GroundingInvestigativeSave

Dastin, J. (2018, October 10). Amazon scraps secret AI recruiting tool that showed bias against women. Reuters. https://www.euronews.com/business/2018/10/10/amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women

https://www.euronews.com/business/2018/10/10/amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS); model org: amazon_resume_engine

EmpiricalTraining a screener on an organization's past hiring decisions imports the past's selection function: research…

Training a screener on an organization's past hiring decisions imports the past's selection function: research on hiring as exploration finds that models trained on prior hires raise hire rates but replicate historical selection, and that a screener which values exploration rather than only exploitation breaks that lock-in loop. Two governance lessons follow — the patch lever has a documented ceiling, since removing named proxies does not remove a learned correlation, and abandonment can itself be a governance outcome, taken here before any external harm was documented rather than after an adjudication.

li2020aGroundingPeer-reviewedSave

Li, D., Raymond, L.R., & Bergman, P. (2020). Hiring as Exploration. NBER Working Paper 27736. https://doi.org/10.3386/w27736 https://www.nber.org/papers/w27736

doi.org/10.3386/w27736

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS)

EmpiricalAn applicant-tracking platform whose AI screening and recommendation features operate inside thousands of empl…

An applicant-tracking platform whose AI screening and recommendation features operate inside thousands of employers' hiring pipelines at once is the subject of a live federal collective action testing whether the vendor is directly liable as the employers' agent. On the litigation record, the court sustained the agent theory at the dismissal stage in 2024 and preliminarily certified a nationwide age-discrimination collective in 2025, covering applicants forty and over since September 2020, on a record in which the lead plaintiff reported more than one hundred rejections across employers using the platform. The litigation is ongoing and nothing here is an adjudicated finding of discrimination; these are allegations and procedural rulings, not a verdict.

mobleyvworkday2024GroundingReferenceSave

Mobley v. Workday, Inc., No. 3:23-cv-00770 (N.D. Cal.): agent-theory vendor liability (2024), preliminary nationwide ADEA collective certification (2025), bias-testing privilege ruling (2026); via Holland & Knight LLP analysis. https://www.hklaw.com/en/insights/publications/2025/05/federal-court-allows-collective-action-lawsuit-over-alleged

https://www.hklaw.com/en/insights/publications/2025/05/federal-court-allows-collective-action-lawsuit-over-alleged

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS); model org: workday_screening_platform

EmpiricalA 2026 discovery ruling in the same matter held the vendor's internal bias-testing data privileged because cou…

A 2026 discovery ruling in the same matter held the vendor's internal bias-testing data privileged because counsel had curated it — meaning the testing record exists and is legally unreachable, a configuration in which audit opacity is not the absence of testing but testing shielded from external verification. The case surfaces two further structural facts: a single vendor's screening model multiplied across many employer boundaries, so one learned defect can propagate as widely as the platform, and accountability diffusion between deployer and vendor, each holding part of the governance the other points to, against a survey backdrop showing assessment vendors' validation and bias-mitigation claims are often unverifiable from outside.

raghavan2020aGroundingPeer-reviewedSave

Raghavan, M., Barocas, S., Kleinberg, J., & Levy, K. (2020). Mitigating Bias in Algorithmic Hiring: Evaluating Claims and Practices. In Proceedings of FAT* '20, 469-481. https://doi.org/10.1145/3351095.3372828 https://arxiv.org/abs/1906.09208

doi.org/10.1145/3351095.3372828

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS)

Topics: algorithmic-fairness

u2023bGroundingGovernmentSave

U.S. EEOC (2023, May 18). Select Issues: Assessing Adverse Impact in Software, Algorithms, and Artificial Intelligence Used in Employment Selection Procedures Under Title VII. Technical assistance document (removed from eeoc.gov early 2025; archived). https://web.archive.org/web/20250102220802/https://www.eeoc.gov/laws/guidance/select-issues-assessing-adverse-impact-software-algorithms-and-artificial

https://web.archive.org/web/20250102220802/https://www.eeoc.gov/laws/guidance/select-issues-assessing-adverse-impact-software-algorithms-and-artificial

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS); model org: workday_screening_platform

EmpiricalA graduate-hiring pipeline chained a games-based assessment with automated video-interview scoring, and the de…

A graduate-hiring pipeline chained a games-based assessment with automated video-interview scoring, and the deployer reports roughly a 90 percent reduction in time-to-hire (from about four months to about four weeks), around 50,000 candidate interview hours saved, about one million pounds in annual savings, and a 16 percent improvement in diversity. Every one of those figures is company- or vendor-reported and none is independently audited, so they are the deployer's own dashboard rather than an external measurement — which is exactly what the family's service regime looks like from inside.

bestpracticeaiGroundingVendorSave

Best Practice AI. Unilever saved over 50,000 hours in candidate interview time and delivered over £1M annual savings and improved candidate diversity with machine analysis of video-based interviewing (AI case study). https://www.bestpractice.ai/ai-case-study-best-practice/unilever_saved_over_50,000_hours_in_candidate_interview_time_and_delivered_over_%C2%A31m_annual_savings_and_improved_candidate_diversity_with_machine_analysis_of_video-based_interviewing.

https://www.bestpractice.ai/ai-case-study-best-practice/unilever_saved_over_50,000_hours_in_candidate_interview_time_and_delivered_over_%C2%A31m_annual_savings_and_improved_candidate_diversity_with_machine_analysis_of_video-based_interviewing

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS); model org: unilever_ai_hiring

EmpiricalBoth vendors' audit machinery is on the public record in an honest but partial form. The games vendor underwen…

Both vendors' audit machinery is on the public record in an honest but partial form. The games vendor underwent a cooperative academic audit with source-code access, in which its four-fifths-rule de-biasing pipeline was found faithfully implemented — with the independence caveat that vendor staff were co-authors — and the video vendor retired its facial-analysis input under scrutiny after internal research found visual features added only about 0.25 percent predictive power, publicizing a narrow-scope external audit. The family's structural blind spot applies in full: rejected candidates never re-enter the outcome data, so the claimed quality and diversity effects are measured on hires only.

wilson2021aGroundingPeer-reviewedSave

Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. In Proceedings of FAccT '21, 666-677. https://doi.org/10.1145/3442188.3445928 https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf

doi.org/10.1145/3442188.3445928

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS)

maurer2021aGroundingTrade pressSave

Maurer, R. (2021). HireVue Discontinues Facial Analysis Screening. SHRM; with HireVue and ORCAA audit announcements (2021). https://orcaarisk.com/in-the-news/2021/1/12/orcaas-audit-of-hirevue-is-live

https://orcaarisk.com/in-the-news/2021/1/12/orcaas-audit-of-hirevue-is-live

Appears in: PAN framework development

Grounds: domain grounding: hiring and employment screening (resume screening, interview scoring, ATS)

EmpiricalA machine-learning underwriting and pricing platform using education and other alternative data operated for f…

A machine-learning underwriting and pricing platform using education and other alternative data operated for five years under a regulator's no-action letter with a reporting obligation, and the regulator published the access results: 27 percent more applicants approved than a traditional model at 16 percent lower average APRs, with near-prime applicants (FICO 620 to 660) approved at roughly twice the rate, and gains across the tested demographic segments. This is the lending family's only regulator-verified service term. Underwriting is fully automated with no per-application human review, so the organizational levers are all upstream — model choice, the testing regime, the search for alternatives, and the reporting channel to the regulator.

consumerfinancialprotectionb2019GroundingGovernmentSave

Consumer Financial Protection Bureau — Ficklin, P.A., & Watkins, P. (2019). An update on credit access and the Bureau's first No-Action Letter. CFPB Blog. https://www.consumerfinance.gov/about-us/blog/update-credit-access-and-no-action-letter/

https://www.consumerfinance.gov/about-us/blog/update-credit-access-and-no-action-letter/

Appears in: PAN framework development

Grounds: domain grounding: lending, credit and collections (underwriting, adverse action, MRM); model org: upstart_nal_underwriting

EmpiricalThe same deployment carries the family's most detailed public fair-lending testing record: four reports from a…

The same deployment carries the family's most detailed public fair-lending testing record: four reports from an independent monitorship agreed with civil-rights organizations found no close protected-class proxies quantitatively, but identified approval disparities for Black applicants, flagged a likely viable less-discriminatory alternative model, and ended in a documented methodological impasse over how hard the law requires an organization to search for such an alternative. Independently of the disparity question, adverse-action notices must give specific, accurate principal reasons for a denial regardless of the model's complexity — a governed explanation duty a complex model does not discharge by being accurate.

relmancolfaxpllc2021GroundingAdvocacySave

Relman Colfax PLLC (2021-2024). Fair Lending Monitorship of Upstart Network's Lending Model (Initial, Second, Third, and Final Reports). https://www.relmanlaw.com/cases-upstart-network-fair-lending-counseling

https://www.relmanlaw.com/cases-upstart-network-fair-lending-counseling

Appears in: PAN framework development

Grounds: domain grounding: lending, credit and collections (underwriting, adverse action, MRM); model org: upstart_nal_underwriting

Topics: complexity-science

consumerfinancialprotectionb2022bGroundingRegulatorySave

Consumer Financial Protection Bureau (2022, 2023). Circular 2022-03: Adverse action notification requirements in connection with credit decisions based on complex algorithms; and Circular 2023-03 on Regulation B sample forms. https://www.consumerfinance.gov/compliance/circulars/circular-2023-03-adverse-action-notification-requirements-and-the-proper-use-of-the-cfpbs-sample-forms-provided-in-regulation-b/

https://www.consumerfinance.gov/compliance/circulars/circular-2023-03-adverse-action-notification-requirements-and-the-proper-use-of-the-cfpbs-sample-forms-provided-in-regulation-b/

Appears in: PAN framework development

Grounds: domain grounding: lending, credit and collections (underwriting, adverse action, MRM)

Topics: complexity-science

EmpiricalA bank's automated credit-decisioning for a widely used consumer card was investigated by a state regulator af…

A bank's automated credit-decisioning for a widely used consumer card was investigated by a state regulator after viral allegations of gender bias in credit-line assignment. The regulator analyzed roughly 400,000 in-state applicants and found no unlawful discrimination on a prohibited basis — the model was cleared on the numbers. But the same investigation documented failures of explanation, customer service, and perceived transparency: applicants could not learn why they received the terms they did, front-line staff could not explain the decisions, and the resulting opacity destroyed consumer trust even though the underwriting itself was found lawful. This is the domain's cleared-but-faulted case: a statistically clean model paired with a failed duty to explain.

newyorkstatedepartmentoffina2021GroundingGovernment evaluationSave

New York State Department of Financial Services (2021, March 23). Report on Apple Card Investigation. https://www.dfs.ny.gov/reports_and_publications/press_releases/pr202103231

https://www.dfs.ny.gov/reports_and_publications/press_releases/pr202103231

Appears in: PAN framework development

Grounds: domain grounding: lending, credit and collections (underwriting, adverse action, MRM); model org: goldman_apple_card

EmpiricalThe lesson the cleared-but-faulted outcome carries is that a lawful, statistically clean model does not discha…

The lesson the cleared-but-faulted outcome carries is that a lawful, statistically clean model does not discharge the separate duty to explain a decision. Regulators have made explicit that adverse-action notices must give specific, accurate principal reasons regardless of how complex the model is, and that a model being a black box is not a defense — checking the nearest sample-form box does not comply. The explanation and customer-service channel is therefore a distinct, separately-resourced surface that can fail on its own: an organization can pass its fair-lending testing and still fail the people it decides on by being unable to tell them why.

consumerfinancialprotectionb2022bGroundingRegulatorySave

Consumer Financial Protection Bureau (2022, 2023). Circular 2022-03: Adverse action notification requirements in connection with credit decisions based on complex algorithms; and Circular 2023-03 on Regulation B sample forms. https://www.consumerfinance.gov/compliance/circulars/circular-2023-03-adverse-action-notification-requirements-and-the-proper-use-of-the-cfpbs-sample-forms-provided-in-regulation-b/

https://www.consumerfinance.gov/compliance/circulars/circular-2023-03-adverse-action-notification-requirements-and-the-proper-use-of-the-cfpbs-sample-forms-provided-in-regulation-b/

Appears in: PAN framework development

Grounds: domain grounding: lending, credit and collections (underwriting, adverse action, MRM)

Topics: complexity-science

EmpiricalA state attorney general reached a $2.5 million settlement with a student-loan lender over its AI underwriting…

A state attorney general reached a $2.5 million settlement with a student-loan lender over its AI underwriting. The documented conduct is the domain's cleanest failure-then-mandated-governance arc: the model used a cohort-default-rate feature — a school's aggregate default rate priced into an individual applicant's terms — that disparately impacted Black and Hispanic applicants, and an immigration-status rule that automatically denied certain non-citizen applicants, while the organization ran no disparate-impact testing and gave inadequate adverse-action notices. The remedy did not fine-and-close: it mandated the missing program — model governance, disparate-impact testing, documentation, and reporting controls — so the enforcement action wrote the governance the deployment had never built.

officeofthemassachusettsatto2025GroundingGovernmentSave

Office of the Massachusetts Attorney General (2025, July 10). AG Campbell Announces $2.5 Million Settlement With Student Loan Lender For Unlawful Practices Through AI Use (Assurance of Discontinuance, Earnest Operations LLC). https://www.mass.gov/news/ag-campbell-announces-25-million-settlement-with-student-loan-lender-for-unlawful-practices-through-ai-use-other-consumer-protection-violations

https://www.mass.gov/news/ag-campbell-announces-25-million-settlement-with-student-loan-lender-for-unlawful-practices-through-ai-use-other-consumer-protection-violations

Appears in: PAN framework development

Grounds: domain grounding: lending, credit and collections (underwriting, adverse action, MRM); model org: earnest_ai_underwriting

EmpiricalThe mechanism the case turns on is the facially-neutral aggregate feature: a cohort default rate is a property…

The mechanism the case turns on is the facially-neutral aggregate feature: a cohort default rate is a property of a school, not of the applicant, and no input names a protected class — yet pricing a group's aggregate history into an individual's terms can carry protected-class impact, which is exactly what disparate-impact testing exists to catch. Here that testing was not done, so the impact went unmeasured until an enforcement action found it. The remedy installed the program the deployment lacked, which is the governable reading: an aggregate feature can look neutral input-by-input and still produce a disparity only outcome testing would reveal, and the absence of that testing is itself the failure.

officeofthemassachusettsatto2025GroundingGovernmentSave

Office of the Massachusetts Attorney General (2025, July 10). AG Campbell Announces $2.5 Million Settlement With Student Loan Lender For Unlawful Practices Through AI Use (Assurance of Discontinuance, Earnest Operations LLC). https://www.mass.gov/news/ag-campbell-announces-25-million-settlement-with-student-loan-lender-for-unlawful-practices-through-ai-use-other-consumer-protection-violations

https://www.mass.gov/news/ag-campbell-announces-25-million-settlement-with-student-loan-lender-for-unlawful-practices-through-ai-use-other-consumer-protection-violations

Appears in: PAN framework development

Grounds: domain grounding: lending, credit and collections (underwriting, adverse action, MRM); model org: earnest_ai_underwriting

consumerfinancialprotectionb2022bGroundingRegulatorySave

Consumer Financial Protection Bureau (2022, 2023). Circular 2022-03: Adverse action notification requirements in connection with credit decisions based on complex algorithms; and Circular 2023-03 on Regulation B sample forms. https://www.consumerfinance.gov/compliance/circulars/circular-2023-03-adverse-action-notification-requirements-and-the-proper-use-of-the-cfpbs-sample-forms-provided-in-regulation-b/

https://www.consumerfinance.gov/compliance/circulars/circular-2023-03-adverse-action-notification-requirements-and-the-proper-use-of-the-cfpbs-sample-forms-provided-in-regulation-b/

Appears in: PAN framework development

Grounds: domain grounding: lending, credit and collections (underwriting, adverse action, MRM)

Topics: complexity-science

EmpiricalThe strongest field evidence for an agent-assist copilot in customer service comes from a staggered randomized…

The strongest field evidence for an agent-assist copilot in customer service comes from a staggered randomized rollout of a generative-AI assistant to roughly 5,000 customer-support agents at a large software firm. Measured against a control group, the copilot raised issues resolved per hour by about 15 percent on average, and it also improved customer sentiment and agent retention. The gain, however, was sharply uneven: novice and low-skill agents improved by roughly 30 to 34 percent, agents with two months of experience performed like agents with six months and no AI, and the most experienced agents gained close to nothing, with some evidence of slight quality degradation. This is the contact-centre domain's cleanest measured benefit, and it is a distribution rather than a single number.

brynjolfsson2025aGroundingPeer-reviewedSave

Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2), 889-942. https://doi.org/10.1093/qje/qjae044 https://academic.oup.com/qje/article/140/2/889/7990658

doi.org/10.1093/qje/qjae044

Appears in: PAN framework development

Grounds: domain grounding: customer service and contact centres (copilots, chatbots, QA); workplace-AI economics: assistance gains concentrate in novices

brynjolfsson2023GroundingAcademicSave

Brynjolfsson, E., Li, D., & Raymond, L. (2023). Generative AI at Work. NBER Working Paper 31161 https://www.nber.org/papers/w31161

https://www.nber.org/papers/w31161

Grounds: model org: fortune500_agent_copilot

EmpiricalThe lesson the randomized evidence carries is skill compression: an agent-assist copilot mostly raises the flo…

The lesson the randomized evidence carries is skill compression: an agent-assist copilot mostly raises the floor. Because almost the entire measured gain accrues to less-experienced agents and the most experienced gain close to nothing, an average productivity number overstates the effect for the agents who least need it and hides that the tool does little for the experienced while possibly costing a small amount of quality there. The governable reading is that the benefit must be measured as a distribution across agent skill, not reported as a scalar — a copilot that helps novices a great deal and experts not at all is a real and specific benefit, and describing it with one average misstates who it helps.

brynjolfsson2025aGroundingPeer-reviewedSave

Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2), 889-942. https://doi.org/10.1093/qje/qjae044 https://academic.oup.com/qje/article/140/2/889/7990658

doi.org/10.1093/qje/qjae044

Appears in: PAN framework development

Grounds: domain grounding: customer service and contact centres (copilots, chatbots, QA); workplace-AI economics: assistance gains concentrate in novices

EmpiricalAn airline's customer-facing website chatbot told a customer they could claim a bereavement fare retroactively…

An airline's customer-facing website chatbot told a customer they could claim a bereavement fare retroactively — a policy that did not exist. The customer relied on the chatbot's statement, bought a ticket, and was then refused the fare by the airline's human staff. A civil-resolution tribunal found the airline liable for negligent misrepresentation and awarded damages, and in doing so rejected the airline's argument that the chatbot was a separate legal entity responsible for its own actions. The tribunal held that the organization is responsible for all the information on its website, whether it comes from a static page or a chatbot, and that a customer has no way to know which source to trust. This is the contact-centre domain's cleanest accountability ruling: the bot is a tool the company answers for, not an entity that answers for itself.

moffattv2024GroundingGovernmentSave

Moffatt v. Air Canada, 2024 BCCRT 149 (British Columbia Civil Resolution Tribunal, February 14, 2024). https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html

https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html

Appears in: PAN framework development

Grounds: domain grounding: customer service and contact centres (copilots, chatbots, QA); model org: air_canada_chatbot

sookman2024GroundingReferenceSave

Sookman, B.B. (2024, February 19). Moffatt v. Air Canada: A Misrepresentation by an AI Chatbot. McCarthy Tétrault TechLex blog https://www.mccarthy.ca/en/insights/blogs/techlex/moffatt-v-air-canada-misrepresentation-ai-chatbot

https://www.mccarthy.ca/en/insights/blogs/techlex/moffatt-v-air-canada-misrepresentation-ai-chatbot

Appears in: PAN framework development

Grounds: domain grounding: customer service and contact centres (copilots, chatbots, QA); model org: air_canada_chatbot

EmpiricalThe duty the ruling establishes is that an organization must take reasonable care that its chatbot's represent…

The duty the ruling establishes is that an organization must take reasonable care that its chatbot's representations are accurate, because the chatbot is a tool it deploys rather than a separate entity that answers for itself. A hallucinated policy or a wrong rule stated by the bot is therefore the organization's own misrepresentation, and a posture that treats the AI as speaking only for itself does not transfer that responsibility away. The governable reading is that a customer-facing chatbot is a channel the organization is accountable for exactly as it is accountable for a page on its own website — so the accuracy control on what the bot states, and the ownership of what it says, are the organization's to build, not the bot's to carry.

sookman2024GroundingReferenceSave

Sookman, B.B. (2024, February 19). Moffatt v. Air Canada: A Misrepresentation by an AI Chatbot. McCarthy Tétrault TechLex blog https://www.mccarthy.ca/en/insights/blogs/techlex/moffatt-v-air-canada-misrepresentation-ai-chatbot

https://www.mccarthy.ca/en/insights/blogs/techlex/moffatt-v-air-canada-misrepresentation-ai-chatbot

Appears in: PAN framework development

Grounds: domain grounding: customer service and contact centres (copilots, chatbots, QA); model org: air_canada_chatbot

moffattv2024GroundingGovernmentSave

Moffatt v. Air Canada, 2024 BCCRT 149 (British Columbia Civil Resolution Tribunal, February 14, 2024). https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html

https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html

Appears in: PAN framework development

Grounds: domain grounding: customer service and contact centres (copilots, chatbots, QA); model org: air_canada_chatbot

EmpiricalAn organization published striking first-month numbers for its customer-facing AI assistant: it handled about …

An organization published striking first-month numbers for its customer-facing AI assistant: it handled about two-thirds of customer-service chats (some 2.3 million conversations), was described as doing the equivalent work of about 700 full-time agents, cut average resolution time from about 11 minutes to under 2, was said to match human customer satisfaction, and was projected to improve profit by tens of millions. Every one of those figures was self-reported and not independently audited. Roughly a year later the same organization reversed course on quality grounds — its chief executive said cost had become too predominant an evaluation factor and the result was lower quality — and committed to always keeping a human available to customers who want one. This is the contact-centre domain's cleanest benefit-then-cost arc: the deflection numbers and the walk-back come from the same deployment.

klarnabankab2024GroundingVendorSave

Klarna Bank AB (2024, February 27). Klarna AI assistant handles two-thirds of customer service chats in its first month (press release via PR Newswire). https://www.prnewswire.com/news-releases/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month-302072740.html

https://www.prnewswire.com/news-releases/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month-302072740.html

Appears in: PAN framework development

Grounds: domain grounding: customer service and contact centres (copilots, chatbots, QA); model org: klarna_ai_assistant

ivanova2025GroundingTrade pressSave

Ivanova, I. (2025, May 9). Klarna plans to hire humans again, as new landmark survey reveals most AI projects fail to deliver. Fortune. https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/

https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/

Appears in: PAN framework development

Grounds: domain grounding: customer service and contact centres (copilots, chatbots, QA); model org: klarna_ai_assistant

EmpiricalThe lesson the benefit-then-cost arc carries is that deflection is not resolution. A published deflection numb…

The lesson the benefit-then-cost arc carries is that deflection is not resolution. A published deflection number reports how many contacts the AI handled, not whether it handled them well, and a figure that is impressive on cost can hide a quality cost that only shows up later — which is what the organization's own reversal described. The survey backdrop sharpens it: most customers say they would rather not meet AI in service and fear it makes reaching a human harder, and industry analysts expect a large share of organizations to abandon plans to reduce their customer-service workforce with AI. The governable reading is to measure resolution and repeat contact against deflection rather than counting deflection as a win by itself, and to protect the path to a human as the safety valve a deflection-maximizing design tends to erode.

ivanova2025GroundingTrade pressSave

Ivanova, I. (2025, May 9). Klarna plans to hire humans again, as new landmark survey reveals most AI projects fail to deliver. Fortune. https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/

https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/

Appears in: PAN framework development

Grounds: domain grounding: customer service and contact centres (copilots, chatbots, QA); model org: klarna_ai_assistant

gartner2025bGroundingTrade pressSave

Gartner, Inc. (2025, June 10). Gartner Predicts 50% of Organizations Will Abandon Plans to Reduce Customer Service Workforce Due to AI (poll of 163 service leaders). https://www.theregister.com/software/2025/06/11/half_of_firms_set_to_abandon_plans_to_ditch_customer_service/502135

https://www.theregister.com/software/2025/06/11/half_of_firms_set_to_abandon_plans_to_ditch_customer_service/502135

Appears in: PAN framework development

Grounds: domain grounding: customer service and contact centres (copilots, chatbots, QA); model org: klarna_ai_assistant

EmpiricalA video platform ran an unintended natural experiment on automated content moderation. When the pandemic sent …

A video platform ran an unintended natural experiment on automated content moderation. When the pandemic sent its human reviewers home, the platform said it would rely more on automated removal and deliberately chose over-enforcement rather than let harmful content stay up. The result, from the platform's own transparency reporting, was that automated removals more than doubled in a single quarter (to about 11.4 million videos), appeals roughly doubled, and the reinstatement rate on appeal jumped from about 25 percent to about 50 percent. The platform also withheld strikes where no human had reviewed the removal, treating the automated decision as provisional. The doubling of the reinstatement rate is the finding: it is direct evidence that the automation was making roughly twice the rate of catchable errors, and that the human review and appeals path was the loop catching them.

youtubegoogle2020GroundingVendorSave

YouTube / Google (2020, August 25). Responsible policy enforcement during Covid-19. Official YouTube blog. https://blog.youtube/inside-youtube/responsible-policy-enforcement-during-covid-19/

https://blog.youtube/inside-youtube/responsible-policy-enforcement-during-covid-19/

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: youtube_covid_enforcement

EmpiricalThe lesson the natural experiment carries is that the human review and appeals path is the error-correction lo…

The lesson the natural experiment carries is that the human review and appeals path is the error-correction loop for automated enforcement, not an optional add-on. Automated moderation makes errors at scale, and a doubling of the reinstatement rate when human review thinned is a measurement of those errors — they were always being made at that rate, and were visible only because the appeals queue surfaced them. Two things follow. Over-enforcement versus under-enforcement is a chosen trade-off: with review capacity cut, the organization decided which error to make, and that was a governance decision. And proactive removal acts before anyone sees the content, so an over-broad takedown is invisible unless appealed — and some removals are irreversible, as when automated systems destroyed documentation of war crimes with archival access declined, leaving no correction loop at all.

youtubegoogle2020GroundingVendorSave

YouTube / Google (2020, August 25). Responsible policy enforcement during Covid-19. Official YouTube blog. https://blog.youtube/inside-youtube/responsible-policy-enforcement-during-covid-19/

https://blog.youtube/inside-youtube/responsible-policy-enforcement-during-covid-19/

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: youtube_covid_enforcement

humanrightswatch2020GroundingAdvocacySave

Human Rights Watch (2020, September 10). 'Video Unavailable': Social Media Platforms Remove Evidence of War Crimes. https://www.hrw.org/report/2020/09/10/video-unavailable/social-media-platforms-remove-evidence-war-crimes

https://www.hrw.org/report/2020/09/10/video-unavailable/social-media-platforms-remove-evidence-war-crimes

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: youtube_covid_enforcement

EmpiricalA platform enforces its content standards with automated classifiers at a scale no human team could match, bac…

A platform enforces its content standards with automated classifiers at a scale no human team could match, backed by a layered correction structure: an internal appeals process, and above it an external oversight board that issues binding decisions on the individual cases it takes and non-binding policy recommendations to the platform. In one year the board overturned the platform's original decision in around 90 percent of the cases it decided, and the platform reported implementing, in progress on, or already aligned with the large majority of the board's cumulative recommendations. This is the moderation domain's most built-out, institutionalized correction structure — layered appeals rising to an independent-adjacent external body that publishes its reasons.

metaplatformsGroundingVendorSave

Meta Platforms (quarterly). Community Standards Enforcement Report. Meta Transparency Center. https://transparency.meta.com/reports/community-standards-enforcement/

https://transparency.meta.com/reports/community-standards-enforcement/

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: meta_content_enforcement

oversightboard2024GroundingReferenceSave

Oversight Board (2024, June 27). 2023 Annual Report Shows Board's Impact on Meta. https://www.oversightboard.com/news/2023-annual-report-shows-boards-impact-on-meta/

https://www.oversightboard.com/news/2023-annual-report-shows-boards-impact-on-meta/

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: meta_content_enforcement

EmpiricalThe reach of the correction structure is the governable limit. The roughly 90 percent overturn rate is measure…

The reach of the correction structure is the governable limit. The roughly 90 percent overturn rate is measured on selected cases — the board chooses emblematic disputes to set precedent, so the figure is evidence that escalated decisions were often wrong, not a random error rate, and the overwhelming majority of automated enforcement decisions never reach the board at all. The board is funded through a platform-established trust, which makes it independent-adjacent rather than fully independent, and its policy recommendations are non-binding. The honest reading is that this correction structure is real and genuinely better than most, and its reach is bounded to the tiny fraction of cases that escalate — so the governing question is whether the correction reaches the scale of the enforcement it is meant to check.

oversightboard2024GroundingReferenceSave

Oversight Board (2024, June 27). 2023 Annual Report Shows Board's Impact on Meta. https://www.oversightboard.com/news/2023-annual-report-shows-boards-impact-on-meta/

https://www.oversightboard.com/news/2023-annual-report-shows-boards-impact-on-meta/

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: meta_content_enforcement

EmpiricalA media outlet published AI-drafted finance explainers under a human-sounding staff byline without disclosing …

A media outlet published AI-drafted finance explainers under a human-sounding staff byline without disclosing to readers that the articles were machine-written. When the practice came to light, the outlet's own audit found it had to issue corrections on a majority of the AI-written articles — on the order of 41 of 77. A byline implies a human review that the reader trusts, and a correction rate that high is a direct measurement that the review the byline implied was not actually performed before publication. A later and sharper case saw another outlet publish articles under entirely fabricated author personas presented as real people, so the failure ran from undisclosed AI drafting to invented human bylines.

bonifacic2023GroundingTrade pressSave

Bonifacic, I. (2023, January 25). CNET had to correct most of its AI-written articles. Engadget. https://www.engadget.com/cnet-corrected-41-of-its-77-ai-written-articles-201519489.html

https://www.engadget.com/cnet-corrected-41-of-its-77-ai-written-articles-201519489.html

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: cnet_ai_drafting

harrisondupre2023GroundingInvestigativeSave

Harrison Dupré, M. (2023, November 27). Sports Illustrated Published Articles by Fake, AI-Generated Writers. Futurism. https://futurism.com/sports-illustrated-ai-generated-writers

https://futurism.com/sports-illustrated-ai-generated-writers

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: cnet_ai_drafting; model org: sports_illustrated_advon

EmpiricalEditorial AI moves the failure from a takedown to a publication, but the governable structure is the same as i…

Editorial AI moves the failure from a takedown to a publication, but the governable structure is the same as in moderation: the byline is the accountability object, and it stands for a review that either happened or did not. Two things are owed to the reader — disclosure that AI was involved, and an editorial check that actually took place — and this deployment gave neither, publishing under a staff byline that implied both. When a large share of AI-drafted articles needs correction, the review was not performed, and the byline misrepresented who did the work. The governable reading is that a human byline on machine-drafted content is a claim about review and disclosure, and a high correction rate is the evidence that the claim was false.

bonifacic2023GroundingTrade pressSave

Bonifacic, I. (2023, January 25). CNET had to correct most of its AI-written articles. Engadget. https://www.engadget.com/cnet-corrected-41-of-its-77-ai-written-articles-201519489.html

https://www.engadget.com/cnet-corrected-41-of-its-77-ai-written-articles-201519489.html

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: cnet_ai_drafting

EmpiricalA large automaker deployed in-line AI inspection at production scale: camera and acoustic systems that detect …

A large automaker deployed in-line AI inspection at production scale: camera and acoustic systems that detect defects during assembly and feed real-time flags to the line worker via a smart device, on a line running on the order of a thousand-plus vehicles a day at a takt of under a minute per station. The system has been established as a company standard and is being extended to suppliers. The governing design is that the AI flags and a human on the line responds — the inspection is wired into a resourced response loop, including the ability to stop the line, so the benefit runs through the human response the flag triggers rather than through the model alone. The documented facts here are the system's function, the worker-interaction model, the scale, and the standardization; the deployment's benefit is reported through corporate and trade channels, and defect-rate deltas from a primary source are not public.

bmwgrouppressclub2025GroundingReferenceSave

BMW Group PressClub (2025, April 28). Artificial intelligence as a quality booster (GenAI4Q pilot, Plant Regensburg). https://www.press.bmwgroup.com/global/article/detail/T0449729EN/artificial-intelligence-as-a-quality-booster?language=en

https://www.press.bmwgroup.com/global/article/detail/T0449729EN/artificial-intelligence-as-a-quality-booster?language=en

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: bmw_aiqx_inspection

metrologyandqualitynews2026GroundingTrade pressSave

Metrology and Quality News (2026, July 6). BMW Group Advances Use of Physical AI in Production (AIQX, Plant Spartanburg). https://metrology.news/bmw-group-advances-use-of-physical-ai-in-production/

https://metrology.news/bmw-group-advances-use-of-physical-ai-in-production/

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: bmw_aiqx_inspection

leanenterpriseinstituteGroundingReferenceSave

Lean Enterprise Institute. Automatic Line Stop (Lean Lexicon). https://www.lean.org/lexicon-terms/automatic-line-stop/

https://www.lean.org/lexicon-terms/automatic-line-stop/

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: audi_press_shop_inspection; model org: bmw_aiqx_inspection

EmpiricalThe lesson the deployment carries is that an in-line inspection AI is only as good as the human-response loop …

The lesson the deployment carries is that an in-line inspection AI is only as good as the human-response loop it triggers, and that loop is the governable object. When the AI flags a defect, a resourced response — a worker with the time to check the flag and the authority to stop the line — is what turns a detection into a caught defect; without it, the flag is just a decision no one acts on. This is why the failure modes in this domain are matters of the loop's calibration rather than the model's raw accuracy: too many false alarms and operators stop responding, too much trust and they stop checking. The honest boundary is that the benefit is reported through corporate and trade channels, and no named manufacturer has publicly attributed a shipped-defect escape to its AI inspection, so the response loop is drawn as the resourced strength and its calibration as the thing to govern, not as a claim about defects that did or did not ship.

bmwgrouppressclub2025GroundingReferenceSave

BMW Group PressClub (2025, April 28). Artificial intelligence as a quality booster (GenAI4Q pilot, Plant Regensburg). https://www.press.bmwgroup.com/global/article/detail/T0449729EN/artificial-intelligence-as-a-quality-booster?language=en

https://www.press.bmwgroup.com/global/article/detail/T0449729EN/artificial-intelligence-as-a-quality-booster?language=en

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: bmw_aiqx_inspection

leanenterpriseinstituteGroundingReferenceSave

Lean Enterprise Institute. Automatic Line Stop (Lean Lexicon). https://www.lean.org/lexicon-terms/automatic-line-stop/

https://www.lean.org/lexicon-terms/automatic-line-stop/

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: audi_press_shop_inspection; model org: bmw_aiqx_inspection

EmpiricalA peer-reviewed heavy-industry predictive-maintenance case study achieved a large, measured reduction in false…

A peer-reviewed heavy-industry predictive-maintenance case study achieved a large, measured reduction in false alarms — on the order of 90 percent — through a closed operator-feedback loop: the maintenance crews investigated the alerts, labeled which were real, and the model retrained on those labels, so the false-alarm rate fell sharply over successive rounds. This is the industrial-QA domain's best-measured quantitative benefit, and it comes from an anonymized study site rather than a named-manufacturer press release, which is the pattern across this domain — the peer-reviewed magnitudes are at anonymized or smaller sites, while the named deployments report their benefit through corporate and trade channels.

hermansa2021aGroundingPeer-reviewedSave

Hermansa, M., Kozielski, M., Michalak, M., Szczyrba, K., Wróbel, Ł., & Sikora, M. (2021). Sensor-Based Predictive Maintenance with Reduction of False Alarms — A Case Study in Heavy Industry. Sensors, 22(1), 226. https://doi.org/10.3390/s22010226 https://pmc.ncbi.nlm.nih.gov/articles/PMC8749854/

doi.org/10.3390/s22010226

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance)

EmpiricalThe lesson the case carries is that the same closed feedback loop that produced the benefit is the thing that …

The lesson the case carries is that the same closed feedback loop that produced the benefit is the thing that can break it, because the loop depends on the crews continuing to engage with the alerts — investigating them, labeling them, responding — and that engagement fails in two opposite directions. Alert fatigue: if too many false alarms arrive before the loop has tuned them down, crews stop trusting the alerts and stop responding, so the feedback the model needs to improve never arrives and the loop stalls. Automation bias: if crews defer to the alerts and stop applying their own judgment, the labels the model retrains on become an echo of its own calls rather than an independent check. Either way the loop degrades, so the measured benefit is contingent on the loop staying calibrated — enough trust that crews respond, enough independence that their labels still carry real judgment.

romeo2025aGroundingAcademicSave

Romeo, G., & Conti, D. (2025). Exploring automation bias in human-AI collaboration: a review and implications for explainable AI. AI & Society. https://doi.org/10.1007/s00146-025-02422-7 https://link.springer.com/article/10.1007/s00146-025-02422-7

doi.org/10.1007/s00146-025-02422-7

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: heavy_industry_pdm

wittbold2026GroundingVendorSave

Wittbold, K. (2026, June 18). Why Your Team Has Stopped Trusting Their Predictive Maintenance Alerts. Augury blog. https://www.augury.com/blog/machine-health/why-your-team-has-stopped-trusting-their-predictive-maintenance-alerts/

https://www.augury.com/blog/machine-health/why-your-team-has-stopped-trusting-their-predictive-maintenance-alerts/

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: heavy_industry_pdm

EmpiricalMachine-learning automated visual inspection of filled injectable drug products flags particulate and cosmetic…

Machine-learning automated visual inspection of filled injectable drug products flags particulate and cosmetic defects that manual inspection or fixed-rule cameras would otherwise judge. In this safety-critical, regulated manufacturing setting the error trade-off is asymmetric and deliberate: a false accept — a missed defect in an injectable that reaches a patient — is a patient-safety failure, while a false reject — scrapping a good vial — is a cost, so the system is tuned to over-reject rather than risk a miss. Because the inspection sits inside a validated pharmaceutical quality process, the AI cannot simply be switched on; it must be qualified within that process, and a regulator is actively developing the framework for how AI in drug manufacturing should be validated and monitored.

veillon2023aGroundingPeer-reviewedSave

Veillon, R., Shabushnig, J., Aabye-Hansen, L., et al. (2023). Applying Machine Learning to the Visual Inspection of Filled Injectable Drug Products. PDA Journal of Pharmaceutical Science and Technology, 77(5), 376-401. https://doi.org/10.5731/pdajpst.2022.012796 https://journal.pda.org/content/77/5/376

doi.org/10.5731/pdajpst.2022.012796

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance)

usfda2023GroundingGovernmentSave

U.S. FDA, CDER/OPQ (2023). Discussion Paper: Artificial Intelligence in Drug Manufacturing. Docket FDA-2023-N-0487. https://www.fda.gov/media/165743/download

https://www.fda.gov/media/165743/download

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: pharma_avi_inspection

EmpiricalTwo governable surfaces follow from putting AI inside a regulated inspection. First, qualification: an AI in a…

Two governable surfaces follow from putting AI inside a regulated inspection. First, qualification: an AI in a validated quality process is not simply deployed but must be qualified and monitored for drift, and because the regulator's AI-specific framework is still developing, the qualification of the model's behavior over time is an emerging, not-yet-settled check rather than a solved one. Second, the human backstop: the manual inspector is what catches the false accepts the over-reject tuning is meant to avoid, so if inspectors come to defer to the AI and stop scrutinizing, that backstop erodes exactly where it matters most — the missed defect the asymmetric tuning was designed to prevent. The governable reading is that the over-reject tuning lowers the visible risk without removing it, and the qualification and the human backstop are what keep the residual risk covered.

veillon2023aGroundingPeer-reviewedSave

Veillon, R., Shabushnig, J., Aabye-Hansen, L., et al. (2023). Applying Machine Learning to the Visual Inspection of Filled Injectable Drug Products. PDA Journal of Pharmaceutical Science and Technology, 77(5), 376-401. https://doi.org/10.5731/pdajpst.2022.012796 https://journal.pda.org/content/77/5/376

doi.org/10.5731/pdajpst.2022.012796

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance)

usfda2023GroundingGovernmentSave

U.S. FDA, CDER/OPQ (2023). Discussion Paper: Artificial Intelligence in Drug Manufacturing. Docket FDA-2023-N-0487. https://www.fda.gov/media/165743/download

https://www.fda.gov/media/165743/download

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: pharma_avi_inspection

EmpiricalA state's statewide dropout early-warning system used ensemble machine learning to label every grade 6 to 9 st…

A state's statewide dropout early-warning system used ensemble machine learning to label every grade 6 to 9 student's risk of not graduating on time and delivered the label to school staff through dashboards for about a decade. An independent, decade-scale audit found the system was wrong roughly 74 percent of the time when it predicted a student would not graduate, produced higher false-alarm rates for Black and Hispanic students, and that the deployer's own internal equity research had gone unpublished — while a survey of districts found administrators reporting no training on how to interpret a 'high risk' label. The state stopped publishing the dashboards in 2023 and said it was evaluating the system's future. The deployment is the education domain's clearest case of a risk label whose error and group disparity entered how students were seen rather than the help they received.

feathers2023GroundingInvestigativeSave

Feathers, T. (2023, April 27). False Alarm: How Wisconsin Uses Race and Income to Label Students 'High Risk'. The Markup (with Chalkbeat). https://themarkup.org/machine-learning/2023/04/27/false-alarm-how-wisconsin-uses-race-and-income-to-label-students-high-risk

https://themarkup.org/machine-learning/2023/04/27/false-alarm-how-wisconsin-uses-race-and-income-to-label-students-high-risk

Appears in: PAN framework development

Grounds: domain grounding: dropout and chronic-absenteeism prediction; domain grounding: schools and education AI (early warning, tutoring, proctoring, grading); model org: wisconsin_dews

knowles2015aGroundingPeer-reviewedSave

Knowles, J.E. (2015). Of needles and haystacks: Building an accurate statewide dropout early warning system in Wisconsin. Journal of Educational Data Mining, 7(3), 18-67. https://doi.org/10.5281/zenodo.3554725 https://jedm.educationaldatamining.org/index.php/JEDM/article/view/JEDM082

doi.org/10.5281/zenodo.3554725

Appears in: PAN framework development

Grounds: domain grounding: schools and education AI (early warning, tutoring, proctoring, grading)

wisconsindepartmentofpublici2023GroundingGovernmentSave

Wisconsin Department of Public Instruction. WISEdash for Districts: Dropout Early Warning System (DEWS) Dashboards (including the October 12, 2023 retirement notice). https://dpi.wi.gov/wisedash/districts/about-data/dews

https://dpi.wi.gov/wisedash/districts/about-data/dews

Appears in: PAN framework development

Grounds: domain grounding: schools and education AI (early warning, tutoring, proctoring, grading); model org: wisconsin_dews

EmpiricalThe lesson the case carries is that a risk label is only as good as the intervention it triggers and the train…

The lesson the case carries is that a risk label is only as good as the intervention it triggers and the training of the human who reads it. A label that is wrong most of the time, delivered to staff with no guidance on interpreting it, imports the model's error and its group disparity into how students are perceived rather than into a resourced response — the flag becomes a lens on the student rather than a trigger for help. Set against this, a large district's transparent, low-tech on-track indicator, built on interpretable research and paired with real intervention, accompanied a rise in graduation to a record level. The contrast locates the benefit in the intervention the indicator makes legible enough for staff to act on well, not in the sophistication of the prediction — an interpretable indicator that drives help can outperform an opaque model that only labels.

allensworth2007GroundingAcademicSave

Allensworth, E.M., & Easton, J.Q. (2007). What Matters for Staying On-Track and Graduating in Chicago Public Schools. University of Chicago Consortium on School Research. https://consortium.uchicago.edu/publications/what-matters-staying-track-and-graduating-chicago-public-schools-focus-students

https://consortium.uchicago.edu/publications/what-matters-staying-track-and-graduating-chicago-public-schools-focus-students

Appears in: PAN framework development

Grounds: domain grounding: schools and education AI (early warning, tutoring, proctoring, grading); model org: cps_freshman_ontrack

feathers2023GroundingInvestigativeSave

Feathers, T. (2023, April 27). False Alarm: How Wisconsin Uses Race and Income to Label Students 'High Risk'. The Markup (with Chalkbeat). https://themarkup.org/machine-learning/2023/04/27/false-alarm-how-wisconsin-uses-race-and-income-to-label-students-high-risk

https://themarkup.org/machine-learning/2023/04/27/false-alarm-how-wisconsin-uses-race-and-income-to-label-students-high-risk

Appears in: PAN framework development

Grounds: domain grounding: dropout and chronic-absenteeism prediction; domain grounding: schools and education AI (early warning, tutoring, proctoring, grading); model org: wisconsin_dews

EmpiricalA public university required students to pan their webcam around their home before an online exam, using remot…

A public university required students to pan their webcam around their home before an online exam, using remote-proctoring software that flags suspected cheating from the video. A federal court held that the pre-exam room scan was an unreasonable search under the Fourth Amendment — a first-of-its-kind ruling that a routine proctoring practice violated a student's constitutional rights in their own home. Separately, peer-reviewed measurement of automated proctoring found the software produced more face-detection failures, more red flags, and higher priority scores for darker-skinned and Black students, with no corresponding difference in actual cheating. The deployment is the education domain's clearest case of surveillance-based integrity AI whose costs — a rights violation and a demographic burden of suspicion — are each independently established.

ogletreev2022aGroundingRegulatorySave

Ogletree v. Cleveland State University, No. 1:21-cv-00500 (N.D. Ohio, August 22, 2022); via Higher Ed Dive and Future of Privacy Forum analyses. https://fpf.org/blog/federal-court-deems-universitys-use-of-room-scans-within-the-home-unconstitutional/

https://fpf.org/blog/federal-court-deems-universitys-use-of-room-scans-within-the-home-unconstitutional/

Appears in: PAN framework development

Grounds: domain grounding: schools and education AI (early warning, tutoring, proctoring, grading)

Topics: privacy-security

yoderhimes2022aGroundingPeer-reviewedSave

Yoder-Himes, D.R., Asif, A., Kinney, K., et al. (2022). Racial, skin tone, and sex disparities in automated proctoring software. Frontiers in Education, 7, 881449. https://doi.org/10.3389/feduc.2022.881449 https://www.frontiersin.org/journals/education/articles/10.3389/feduc.2022.881449/full

doi.org/10.3389/feduc.2022.881449

Appears in: PAN framework development

Grounds: domain grounding: schools and education AI (early warning, tutoring, proctoring, grading)

EmpiricalThe lesson the case carries is that surveillance-based integrity AI is not a free default: it carries a rights…

The lesson the case carries is that surveillance-based integrity AI is not a free default: it carries a rights cost that can be independently adjudicated and a demographic burden that can be measured, and both are owed a reckoning before the surveillance is imposed, not after a court or an audit finds the harm. A room scan of a student's home was held to be an unreasonable search, so the surveillance has a rights dimension a court can rule on regardless of the integrity goal. And because the software flags darker-skinned and Black students more often with no more actual cheating, and a flag is an accusation the student must answer, a disparate flag rate is a disparate burden of suspicion. The governable surfaces are the proportionality of the surveillance to the integrity problem it is trying to solve, and the measured flag rate by group.

yoderhimes2022aGroundingPeer-reviewedSave

Yoder-Himes, D.R., Asif, A., Kinney, K., et al. (2022). Racial, skin tone, and sex disparities in automated proctoring software. Frontiers in Education, 7, 881449. https://doi.org/10.3389/feduc.2022.881449 https://www.frontiersin.org/journals/education/articles/10.3389/feduc.2022.881449/full

doi.org/10.3389/feduc.2022.881449

Appears in: PAN framework development

Grounds: domain grounding: schools and education AI (early warning, tutoring, proctoring, grading)

ogletreev2022aGroundingRegulatorySave

Ogletree v. Cleveland State University, No. 1:21-cv-00500 (N.D. Ohio, August 22, 2022); via Higher Ed Dive and Future of Privacy Forum analyses. https://fpf.org/blog/federal-court-deems-universitys-use-of-room-scans-within-the-home-unconstitutional/

https://fpf.org/blog/federal-court-deems-universitys-use-of-room-scans-within-the-home-unconstitutional/

Appears in: PAN framework development

Grounds: domain grounding: schools and education AI (early warning, tutoring, proctoring, grading)

Topics: privacy-security

EmpiricalA parcel carrier's route-optimization system is a documented operations-research success: it re-optimizes deli…

A parcel carrier's route-optimization system is a documented operations-research success: it re-optimizes delivery routes across the fleet and was reported to save on the order of 100 million miles and about 10 million gallons of fuel a year, a genuine and peer-reviewed efficiency gain. The same system that computes the efficient route also dictates it to the driver and monitors adherence through vehicle telematics, so the efficiency is enforced through workplace surveillance — the optimization and the monitoring are one system, and the driver's discretion over how to run the route is what it replaces. The benefit is real and measured in miles and fuel; the cost is the driver autonomy the enforcement removes and the surveillance the enforcement requires.

holland2017GroundingAcademicSave

Holland, C., Levis, J., Nuggehalli, R., Santilli, B., & Winters, J. (2017). UPS Optimizes Delivery Routes. Interfaces, 47(1), 8-23. https://doi.org/10.1287/inte.2016.0875

doi.org/10.1287/inte.2016.0875

Appears in: PAN framework development

Grounds: domain grounding: logistics, dispatch and scheduling (route optimization, warehouse, workforce scheduling); model org: ups_orion_routing

levy2023GroundingAcademicSave

Levy, K. (2023). Data Driven: Truckers, Technology, and the New Workplace Surveillance. Princeton University Press. https://press.princeton.edu/books/hardcover/9780691175300/data-driven

https://press.princeton.edu/books/hardcover/9780691175300/data-driven

Appears in: PAN framework development

Grounds: domain grounding: logistics, dispatch and scheduling (route optimization, warehouse, workforce scheduling); model org: ups_orion_routing

EmpiricalThe lesson the case carries is that an optimization which manages the worker executing it couples the efficien…

The lesson the case carries is that an optimization which manages the worker executing it couples the efficiency gain to a cost the efficiency metric does not see: the worker's autonomy, and the surveillance required to enforce the plan. The system measures miles and fuel, not whether the pace it sets is feasible for a person or whether the monitoring it requires is proportionate — so the governable surfaces are whether the optimization internalizes the human executing it, meaning a route that is feasible and humane rather than merely optimal on paper, and whether the surveillance that enforces it is governed rather than treated as a free byproduct of routing. An optimization is a success on its own terms and can still externalize a cost onto the worker that never appears in the miles-and-fuel number it reports.

levy2023GroundingAcademicSave

Levy, K. (2023). Data Driven: Truckers, Technology, and the New Workplace Surveillance. Princeton University Press. https://press.princeton.edu/books/hardcover/9780691175300/data-driven

https://press.princeton.edu/books/hardcover/9780691175300/data-driven

Appears in: PAN framework development

Grounds: domain grounding: logistics, dispatch and scheduling (route optimization, warehouse, workforce scheduling); model org: ups_orion_routing

holland2017GroundingAcademicSave

Holland, C., Levis, J., Nuggehalli, R., Santilli, B., & Winters, J. (2017). UPS Optimizes Delivery Routes. Interfaces, 47(1), 8-23. https://doi.org/10.1287/inte.2016.0875

doi.org/10.1287/inte.2016.0875

Appears in: PAN framework development

Grounds: domain grounding: logistics, dispatch and scheduling (route optimization, warehouse, workforce scheduling); model org: ups_orion_routing

EmpiricalA warehouse operation's algorithmic management pairs a genuine, peer-reviewed human-robot picking benefit — ro…

A warehouse operation's algorithmic management pairs a genuine, peer-reviewed human-robot picking benefit — robots and workers collaborating to raise throughput, documented in the operations-research literature — with a documented injury-productivity trade-off. When the algorithm sets the pace of the physical work, a federal safety regulator cited the operation for exposing workers to ergonomic hazards, and a legislative inquiry tied the speed the system demands to warehouses it described as uniquely dangerous. The productivity gain and the worker-injury risk are therefore coupled: the same pace that raises units per hour is the pace regulators and the inquiry connected to injury. The benefit is real and the injury cost is separately documented, one in the OR literature and one in safety-inspection findings and a legislative report.

allgor2023GroundingAcademicSave

Allgor, R., Cezik, T., & Chen, D. (2023). Algorithm for Robotic Picking in Amazon Fulfillment Centers Enables Humans and Robots to Work Together Effectively. INFORMS Journal on Applied Analytics, 53(4). https://doi.org/10.1287/inte.2022.1143

doi.org/10.1287/inte.2022.1143

Appears in: PAN framework development

Grounds: domain grounding: logistics, dispatch and scheduling (route optimization, warehouse, workforce scheduling); model org: amazon_fulfillment_management

ussenatecommitteeonhealth2024GroundingGovernmentSave

U.S. Senate Committee on Health, Education, Labor, and Pensions (2024, December 15). The Injury-Productivity Trade-off: How Amazon's Obsession with Speed Creates Uniquely Dangerous Warehouses. https://www.help.senate.gov/imo/media/doc/amazon_investigation.pdf

https://www.help.senate.gov/imo/media/doc/amazon_investigation.pdf

Appears in: PAN framework development

Grounds: domain grounding: logistics, dispatch and scheduling (route optimization, warehouse, workforce scheduling); model org: amazon_fulfillment_management

usdepartmentoflabor2023aGroundingRegulatorySave

U.S. Department of Labor, OSHA (2023, January 18 and February 1). Federal safety inspections at Amazon warehouse facilities find company exposed workers to ergonomic, struck-by hazards (national news releases). https://www.osha.gov/news/newsreleases/osha-national-news-release/20230201

https://www.osha.gov/news/newsreleases/osha-national-news-release/20230201

Appears in: PAN framework development

Grounds: domain grounding: logistics, dispatch and scheduling (route optimization, warehouse, workforce scheduling)

Topics: ai-safety

EmpiricalThe lesson the case carries is that when an algorithm sets the pace of physical work, the productivity metric …

The lesson the case carries is that when an algorithm sets the pace of physical work, the productivity metric it optimizes — units per hour — cannot see the cost the pace imposes on the body executing it. The injury shows up in safety-inspection data and a legislative inquiry, not on the throughput dashboard, so a productivity number can rise while the cost accumulates unrecorded on the metric that reports success. The governable question is whether the pace-setting internalizes the worker's safety, treating a sustainable rate as part of what 'optimal' means, or externalizes it as an injury the metric never records. Ethnographic research describes this algorithmic management as a 'game' whose rules the worker cannot change, which is what makes the pace a management decision the organization owns rather than a fact of the work.

cheon2025GroundingAcademicSave

Cheon, E., & Erickson, I. (2025). Fulfillment of the Work Games: Warehouse Workers' Experiences with Algorithmic Management. Proceedings of the ACM on Human-Computer Interaction (CSCW). https://doi.org/10.1145/3757409 https://arxiv.org/abs/2508.09438

doi.org/10.1145/3757409

Appears in: PAN framework development

Grounds: domain grounding: logistics, dispatch and scheduling (route optimization, warehouse, workforce scheduling); model org: amazon_fulfillment_management

Topics: algorithmic-fairness

ussenatecommitteeonhealth2024GroundingGovernmentSave

U.S. Senate Committee on Health, Education, Labor, and Pensions (2024, December 15). The Injury-Productivity Trade-off: How Amazon's Obsession with Speed Creates Uniquely Dangerous Warehouses. https://www.help.senate.gov/imo/media/doc/amazon_investigation.pdf

https://www.help.senate.gov/imo/media/doc/amazon_investigation.pdf

Appears in: PAN framework development

Grounds: domain grounding: logistics, dispatch and scheduling (route optimization, warehouse, workforce scheduling); model org: amazon_fulfillment_management

EmpiricalA federal asylum agency uses dialect-recognition AI to estimate an applicant's country or region of origin fro…

A federal asylum agency uses dialect-recognition AI to estimate an applicant's country or region of origin from a short speech sample, as one input into the credibility assessment of their claimed origin. The tool's reliability is limited: government-reported recognition is around 80 percent for Arabic — roughly a 20 percent error rate — and computational linguists judge separating some closely related language varieties close to hopeless. The agency's own caseworkers describe the tool as only a rough compass, too imprecise to resolve the hard cases, and its outputs as clues rather than determinations. Used honestly as one clue among several it is defensible; the documented risk is that an imprecise output acquires more authority than its accuracy supports, in a determination where the state's tool is set against the applicant's own account of who they are.

lulamae2022aGroundingInvestigativeSave

Lulamae, J. (2022, September 5). The BAMF's controversial dialect recognition software: new languages and an EU pilot project. AlgorithmWatch; with Beck, J. (2026), Verfassungsblog legal analysis (https://doi.org/10.59704/b22636dc94f60b29). https://verfassungsblog.de/dialect-recognition-software-dias-law/

doi.org/10.59704/b22636dc94f60b29)

Appears in: PAN framework development

Grounds: domain grounding: immigration and asylum AI (casework tools, triage, agency governance)

scheel2024aGroundingPeer-reviewedSave

Scheel, S. (2024). Epistemic domination by data extraction: questioning the use of biometrics and mobile phone data analysis in asylum procedures. Journal of Ethnic and Migration Studies, 50(9), 2289-2308. https://doi.org/10.1080/1369183X.2024.2307782 https://pmc.ncbi.nlm.nih.gov/articles/PMC11034547/

doi.org/10.1080/1369183X.2024.2307782

Appears in: PAN framework development

Grounds: domain grounding: immigration and asylum AI (casework tools, triage, agency governance)

EmpiricalTwo governable surfaces follow from putting a low-reliability signal into a high-stakes credibility determinat…

Two governable surfaces follow from putting a low-reliability signal into a high-stakes credibility determination. First, whether the tool's documented imprecision actually bounds the weight it carries: a rough compass treated as one is honest, but the same output can harden into a credibility finding it cannot support once a phrase like the software indicates a particular origin enters the record and confronts the applicant. Second, whether the applicant can see and contest the signal: in asylum determinations the person with the most at stake and the most knowledge of their own origin is often unable to see or challenge the AI's estimate, so the correction that would catch an error is severed on exactly the side that holds the truth. The governable reading is that reliability must bound authority, and the affected person must be able to contest a signal used against them.

scheel2024aGroundingPeer-reviewedSave

Scheel, S. (2024). Epistemic domination by data extraction: questioning the use of biometrics and mobile phone data analysis in asylum procedures. Journal of Ethnic and Migration Studies, 50(9), 2289-2308. https://doi.org/10.1080/1369183X.2024.2307782 https://pmc.ncbi.nlm.nih.gov/articles/PMC11034547/

doi.org/10.1080/1369183X.2024.2307782

Appears in: PAN framework development

Grounds: domain grounding: immigration and asylum AI (casework tools, triage, agency governance)

lulamae2022aGroundingInvestigativeSave

Lulamae, J. (2022, September 5). The BAMF's controversial dialect recognition software: new languages and an EU pilot project. AlgorithmWatch; with Beck, J. (2026), Verfassungsblog legal analysis (https://doi.org/10.59704/b22636dc94f60b29). https://verfassungsblog.de/dialect-recognition-software-dias-law/

doi.org/10.59704/b22636dc94f60b29)

Appears in: PAN framework development

Grounds: domain grounding: immigration and asylum AI (casework tools, triage, agency governance)

EmpiricalA government's immigration-enforcement triage algorithm identifies and recommends people for enforcement actio…

A government's immigration-enforcement triage algorithm identifies and recommends people for enforcement actions — returns, bail conditions, casework — drawing on sensitive data including detention, health, vulnerability, and location-monitoring records. Uncovered through roughly a year of freedom-of-information litigation, its training materials show an asymmetric override design: officials must record a justification for rejecting a recommendation but not for accepting one. That design builds a rubber-stamping incentive into the workflow — accepting the algorithm is frictionless, overriding it requires work — so the human in the loop is nominal rather than a real check. It is the corpus's clearest documented instance of automation bias engineered into an agency workflow, in one of the highest-stakes enforcement settings a state operates.

privacyinternational2024bGroundingAdvocacySave

Privacy International (2024, October 17). Automating the hostile environment: uncovering the secretive Home Office algorithm at the heart of immigration enforcement (IPIC); with the 2025 ICO complaint and the primary FOI trail. https://privacyinternational.org/news-analysis/5452/automating-hostile-environment-uncovering-secretive-home-office-algorithm-heart

https://privacyinternational.org/news-analysis/5452/automating-hostile-environment-uncovering-secretive-home-office-algorithm-heart

Appears in: PAN framework development

Grounds: domain grounding: immigration and asylum AI (casework tools, triage, agency governance); model org: home_office_ipic

Topics: privacy-security

privacyinternational2024aGroundingAdvocacySave

Privacy International (2024, October 17). Automating the hostile environment: uncovering the secretive Home Office algorithm at the heart of immigration enforcement (IPIC); with the 2025 ICO complaint and the primary FOI trail. https://www.whatdotheyknow.com/request/identify_and_prioritise_immigrat_3

https://www.whatdotheyknow.com/request/identify_and_prioritise_immigrat_3

Appears in: PAN framework development

Grounds: domain grounding: immigration and asylum AI (casework tools, triage, agency governance)

Topics: privacy-security

EmpiricalThe lesson the case carries is that nominal human oversight is not real oversight. An asymmetric override — wh…

The lesson the case carries is that nominal human oversight is not real oversight. An asymmetric override — where accepting the algorithm's recommendation is frictionless and rejecting it requires a recorded justification — engineers automation bias into the process by making deference the path of least resistance, so the claim that a human makes the final decision can be true and empty at once. Two governable surfaces follow. Whether the review is genuinely symmetric: the official as free and as prompted to reject as to accept, so an error is as likely to be caught as waved through. And whether the affected person is told the AI is used and can contest it: applicants are frequently not told, which severs the correction on the side that could challenge the recommendation, so the one check that survives the asymmetric override — the person it is about — is cut out too.

privacyinternational2024cGroundingAdvocacySave

Privacy International (2024, October 17). Automating the hostile environment: uncovering the secretive Home Office algorithm at the heart of immigration enforcement (IPIC); with the 2025 ICO complaint and the primary FOI trail. https://privacyinternational.org/press-release/5640/privacy-international-issues-complaint-uk-regulator-regarding-deployment-two

https://privacyinternational.org/press-release/5640/privacy-international-issues-complaint-uk-regulator-regarding-deployment-two

Appears in: PAN framework development

Grounds: domain grounding: immigration and asylum AI (casework tools, triage, agency governance); model org: home_office_ipic

Topics: privacy-security

privacyinternational2024aGroundingAdvocacySave

Privacy International (2024, October 17). Automating the hostile environment: uncovering the secretive Home Office algorithm at the heart of immigration enforcement (IPIC); with the 2025 ICO complaint and the primary FOI trail. https://www.whatdotheyknow.com/request/identify_and_prioritise_immigrat_3

https://www.whatdotheyknow.com/request/identify_and_prioritise_immigrat_3

Appears in: PAN framework development

Grounds: domain grounding: immigration and asylum AI (casework tools, triage, agency governance)

Topics: privacy-security

EmpiricalA large automaker developed an in-house deep-learning system to detect hairline cracks in pressed sheet-metal …

A large automaker developed an in-house deep-learning system to detect hairline cracks in pressed sheet-metal parts, trained on several terabytes of images drawn from seven presses at its home plant plus several sister plants, in development since mid-2016 and tested for series deployment. The documented change is a generational replacement: the system takes over an inspection duty previously performed by manual visual checks plus fixed-rule camera systems, rather than augmenting a human inspector's judgment on each part. The record — a reprint of the manufacturer's own press material with its CIO quoted — documents the development lineage, the data scale, and what the system replaced; it publishes no quantitative defect-rate figures, so the deployment's benefit magnitude is a corporate claim, not an audited measurement.

EmpiricalThe governance shape of this deployment is inheritance rather than assistance: by replacing the manual visual …

The governance shape of this deployment is inheritance rather than assistance: by replacing the manual visual check and the fixed-rule camera generation, the learned system inherits the whole inspection duty for the defect class it covers, so there is no per-part human judgment running alongside it to catch what it misses. Its training data is pooled across presses and plants, which means one model's blind spots are correlated across every line it inspects. The failure regime is mechanism-level — drift as dies wear and parts change, complacency over an inspection nobody re-performs — because no named manufacturer, including this one, has publicly attributed a shipped-defect escape to its AI inspection.

EmpiricalA rail service provider operates sensor-based predictive maintenance on a high-speed fleet under a priced avai…

A rail service provider operates sensor-based predictive maintenance on a high-speed fleet under a priced availability contract: roughly 300 sensors per train read at five-minute intervals (on the order of a million readings per train-year), overlaid with human-written failure reports, maintained against a promise that refunds the full fare if a journey is delayed more than fifteen minutes. The documented results are only one noticeably delayed journey in 2,300 (by five minutes) and discovered failure signatures such as an engine-temperature pattern preceding failure by three days. The fleet outcome is documented in a vendor-side trade case study; the analytics' own precision is not published, so the model-level figures remain unstated while the operational outcome is on the record.

rcrwirelessnews2016GroundingTrade pressSave

RCR Wireless News (2016, September 12). Case study: Siemens reduces train failures with Teradata Aster (Renfe Velaro E predictive maintenance). https://www.rcrwireless.com/20160912/big-data-analytics/siemens-train-teradata-tag31-tag99

https://www.rcrwireless.com/20160912/big-data-analytics/siemens-train-teradata-tag31-tag99

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: siemens_renfe_velaro_pdm

EmpiricalThe governance shape of this deployment is uptime-as-contract: the party that operates and tunes the analytics…

The governance shape of this deployment is uptime-as-contract: the party that operates and tunes the analytics is the service provider who pays for misses under the refund promise, so the incentive to prevent a delay is priced into the same organization that holds the model levers - an alignment the domain's other deployments lack. Two documented dependencies temper it: the failure signatures were discovered by joining sensor streams to human-written failure reports, so the discovery loop runs on documentation crews write for their own purposes; and a continuous sensor stream is the analytics' only view of the machine, so a failing sensor and a failing train arrive looking the same until someone goes and looks.

rcrwirelessnews2016GroundingTrade pressSave

RCR Wireless News (2016, September 12). Case study: Siemens reduces train failures with Teradata Aster (Renfe Velaro E predictive maintenance). https://www.rcrwireless.com/20160912/big-data-analytics/siemens-train-teradata-tag31-tag99

https://www.rcrwireless.com/20160912/big-data-analytics/siemens-train-teradata-tag31-tag99

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance); model org: siemens_renfe_velaro_pdm

EmpiricalA large urban school district operationalized a transparent ninth-grade indicator - course credits earned plus…

A large urban school district operationalized a transparent ninth-grade indicator - course credits earned plus no more than one core-course failure - from consortium research showing it predicts high-school graduation with about 85 percent accuracy, and wired it to school-level attention rather than to an opaque score. District graduation rates subsequently rose to record highs. The indicator is a rule anyone can read: a teacher can explain to a student exactly why they are off-track and exactly what would change it, so the contest-and-correction loop that opaque early-warning deployments sever is open by construction.

allensworth2007GroundingAcademicSave

Allensworth, E.M., & Easton, J.Q. (2007). What Matters for Staying On-Track and Graduating in Chicago Public Schools. University of Chicago Consortium on School Research. https://consortium.uchicago.edu/publications/what-matters-staying-track-and-graduating-chicago-public-schools-focus-students

https://consortium.uchicago.edu/publications/what-matters-staying-track-and-graduating-chicago-public-schools-focus-students

Appears in: PAN framework development

Grounds: domain grounding: schools and education AI (early warning, tutoring, proctoring, grading); model org: cps_freshman_ontrack

EmpiricalThe documented limits are as instructive as the result. The indicator's accuracy and the district's graduation…

The documented limits are as instructive as the result. The indicator's accuracy and the district's graduation rise are associational at district scale - no randomized trial assigns schools to use it - and the benefit mechanism runs through the intervention, not the flag: an indicator wired to attention still depends on the attention being resourced, and the research base's central finding is that what predicted graduation was a condition schools could act on (freshman-year course performance), not a fixed trait of the student. The rule's power is that it points at something changeable, and the district's practice is what changed it.

allensworth2007GroundingAcademicSave

Allensworth, E.M., & Easton, J.Q. (2007). What Matters for Staying On-Track and Graduating in Chicago Public Schools. University of Chicago Consortium on School Research. https://consortium.uchicago.edu/publications/what-matters-staying-track-and-graduating-chicago-public-schools-focus-students

https://consortium.uchicago.edu/publications/what-matters-staying-track-and-graduating-chicago-public-schools-focus-students

Appears in: PAN framework development

Grounds: domain grounding: schools and education AI (early warning, tutoring, proctoring, grading); model org: cps_freshman_ontrack

EmpiricalA storied sports outlet published product-review articles under entirely fabricated author personas - invented…

A storied sports outlet published product-review articles under entirely fabricated author personas - invented names, AI-generated headshots, fictional biographies - produced by a third-party content contractor, with the AI involvement disclosed to no reader. An investigation surfaced the personas by reading the public site; the articles were then deleted rather than corrected, the outlet attributed the content to the contractor, and the parent company's chief executive was subsequently fired. The corroborated record documents the fabrication, the deletion, the contractor attribution, and the executive consequence.

harrisondupre2023GroundingInvestigativeSave

Harrison Dupré, M. (2023, November 27). Sports Illustrated Published Articles by Fake, AI-Generated Writers. Futurism. https://futurism.com/sports-illustrated-ai-generated-writers

https://futurism.com/sports-illustrated-ai-generated-writers

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: cnet_ai_drafting; model org: sports_illustrated_advon

npr2023GroundingTrade pressSave

NPR (2023, November 28). Sports Illustrated is accused of posting articles by writers created by AI https://www.npr.org/2023/11/28/1215693615/sports-illustrated-is-accused-of-posting-articles-by-writers-created-by-ai

https://www.npr.org/2023/11/28/1215693615/sports-illustrated-is-accused-of-posting-articles-by-writers-created-by-ai

Appears in: PAN framework development

Grounds: dossier corroboration: Sports Illustrated / AdVon fabricated AI authors; model org: sports_illustrated_advon

EmpiricalThe governance failure ran across an organizational seam: the drafting, the bylines, and the personas were pro…

The governance failure ran across an organizational seam: the drafting, the bylines, and the personas were produced by a contractor, and the outlet's editorial function demonstrably did not operate across that boundary - the fabrication was discovered by outside investigation, not by any internal check, and the accountability afterward ran through contract and employment rather than through any editorial process. A byline makes two claims to the reader - that a person produced this, and that the outlet's review stands behind it; this deployment fabricated the first and vacated the second, and the deletion afterward removed the evidence rather than correcting the record.

harrisondupre2023GroundingInvestigativeSave

Harrison Dupré, M. (2023, November 27). Sports Illustrated Published Articles by Fake, AI-Generated Writers. Futurism. https://futurism.com/sports-illustrated-ai-generated-writers

https://futurism.com/sports-illustrated-ai-generated-writers

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: cnet_ai_drafting; model org: sports_illustrated_advon

npr2023GroundingTrade pressSave

NPR (2023, November 28). Sports Illustrated is accused of posting articles by writers created by AI https://www.npr.org/2023/11/28/1215693615/sports-illustrated-is-accused-of-posting-articles-by-writers-created-by-ai

https://www.npr.org/2023/11/28/1215693615/sports-illustrated-is-accused-of-posting-articles-by-writers-created-by-ai

Appears in: PAN framework development

Grounds: dossier corroboration: Sports Illustrated / AdVon fabricated AI authors; model org: sports_illustrated_advon

EmpiricalA parcel firm's customer-facing support chatbot, after a system update, was prompted by a customer into sweari…

A parcel firm's customer-facing support chatbot, after a system update, was prompted by a customer into swearing and into composing a poem calling its own operator the worst delivery firm in the world. The firm attributed the behavior to the update and disabled the AI element immediately. The documented governance facts are exactly two: the update preceded the behavior, and the off switch worked - the firm learned of the incident from the customer's viral post rather than from any release gate, but the disablement was immediate once it knew.

EmpiricalThe failure shape is a change-management regression, not a wrong policy or a deflection metric: guardrails tha…

The failure shape is a change-management regression, not a wrong policy or a deflection metric: guardrails that had held in production stopped holding after a change, publicly, within hours, in a channel that talks to anyone. What the deployment lacked was a release gate between the update and the public - the constraint layer's behavior after the change was tested by a customer with a prompt, not by the firm with a suite - and the discovery path ran through screenshots of one conversation going viral.

ConceptualThe volume's ethics chapter names ethics washing as addressing ethical concerns superficially, to gain public …

The volume's ethics chapter names ethics washing as addressing ethical concerns superficially, to gain public trust, while making no substantive change to practice. It identifies three forms this takes: ethical statements that are vague or go unenforced, ethics boards constituted with limited authority, and adoption of frameworks that carry no accountability mechanism. The chapter offers this as a taxonomy of forms, not as a measurement of how often each occurs.

an2026aAcademicSave

An, R., & Lindsey, M. A. (2026). Ethical Foundations of AI in Social Work. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_2

doi.org/10.1007/978-3-032-18443-6_2

Appears in: AI in Social Work (Springer, 2026)

Topics: ai-ethics, social-work

EmpiricalThe volume's child-welfare chapter reports that a widely deployed screening score predicts whether a child wil…

The volume's child-welfare chapter reports that a widely deployed screening score predicts whether a child will be placed out of the home within two years, which is the system's own future response rather than the maltreatment the worker is deciding about. The chapter treats the gap between the modelled target and the decision's actual question as a design property of the deployment, not as a defect in the model's accuracy.

zhang2026bAcademicAcademicSave

Zhang, L., & Denby-Brinson, R. (2026). AI in Child Welfare and Family Services. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_4

doi.org/10.1007/978-3-032-18443-6_4

Appears in: AI in Social Work (Springer, 2026)

Grounds: model org: allegheny_afst; model org: eckerd_florida_rsf_origin; model org: illinois_rapid_safety_feedback; model org: oregon_safety_at_screening

Topics: algorithmic-fairness, child-welfare, social-work

EmpiricalThe same chapter reports that workers who distrust a screening tool may still follow it, because organizationa…

The same chapter reports that workers who distrust a screening tool may still follow it, because organizational and policy pressure makes following the tool the defensible act. Deference on this account is produced by where accountability sits, not only by how much the worker trusts the output.

zhang2026bAcademicAcademicSave

Zhang, L., & Denby-Brinson, R. (2026). AI in Child Welfare and Family Services. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_4

doi.org/10.1007/978-3-032-18443-6_4

Appears in: AI in Social Work (Springer, 2026)

Grounds: model org: allegheny_afst; model org: eckerd_florida_rsf_origin; model org: illinois_rapid_safety_feedback; model org: oregon_safety_at_screening

Topics: algorithmic-fairness, child-welfare, social-work

EmpiricalHow a model scores on data held back from its own training and how it scores at a different site are different…

How a model scores on data held back from its own training and how it scores at a different site are different quantities. The volume's disability chapter reports a named pair where the second is materially lower than the first. A figure quoted without saying which of the two it is does not tell a reader what the model will do in their setting.

wang2026aAcademicSave

Wang, J., & Begg, M. D. (2026). AI for Physical, Cognitive, and Developmental Challenges. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_6

doi.org/10.1007/978-3-032-18443-6_6

Appears in: AI in Social Work (Springer, 2026)

Topics: disability, human-ai-interaction, social-work

EmpiricalThe volume's substance-use chapter reports that nearly all models in the field it reviews are developed and ev…

The volume's substance-use chapter reports that nearly all models in the field it reviews are developed and evaluated on a single dataset. Where that holds, a reported ceiling describes that sample rather than a portable capability, and it should be read as the best case observed on one collection of records, not as what the tool will do elsewhere.

saba2026AcademicAcademicSave

Saba, S., & Leibowitz, G. (2026). AI in Substance Use and Addiction Prevention. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_11

doi.org/10.1007/978-3-032-18443-6_11

Appears in: AI in Social Work (Springer, 2026)

Grounds: model org: odmap_overdose_spike_alerts; model org: woebot_health_app

Topics: social-work, substance-use

EmpiricalDiscrimination statistics such as AUROC, precision, F1 and lift describe how a model separates cases on a labe…

Discrimination statistics such as AUROC, precision, F1 and lift describe how a model separates cases on a labelled sample. They are not per-interaction rates at which an error is adopted, written into a record, or corrected. The volume's research chapter treats these as different quantities, and they must never be entered into a diagram as though one substitutes for the other.

yang2026eAcademicSave

Yang, Y., Huang, J., & An, R. (2026). AI in Social Work Research. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_21

doi.org/10.1007/978-3-032-18443-6_21

Appears in: AI in Social Work (Springer, 2026)

Topics: research-methods, social-work, social-work-research

ConceptualThe volume's poverty chapter uses the targeting literature's paired vocabulary for the two directions in which…

The volume's poverty chapter uses the targeting literature's paired vocabulary for the two directions in which a system steering a scarce resource fails: an inclusion error reaches someone the program did not intend to reach, and an exclusion error leaves out someone it did intend to reach. The chapter reports reducing both as the stated aim of machine-learning-assisted eligibility and proxy-means targeting. It is a narrative review and measures neither rate itself; the accuracy results it summarises belong to its sources.

zeng2026AcademicSave

Zeng, Y., & Singletary, J. E. (2026). AI in Combating Poverty and Economic Inequality. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_3

doi.org/10.1007/978-3-032-18443-6_3

Appears in: AI in Social Work (Springer, 2026)

Topics: algorithmic-fairness, poverty, social-work

ConceptualTwo chapters of the volume describe the same gap from opposite ends. The governance chapter names an ethical c…

Two chapters of the volume describe the same gap from opposite ends. The governance chapter names an ethical capacity gap: many social workers have not been trained in data science or AI oversight, which it argues leaves them ill-prepared to question or interpret the algorithmic outputs they are nonetheless answerable for. The literacy chapter's professional-development framework assigns audiences by tier, placing sanctioned-tool lists, ethics review boards and vendor bias-mitigation terms with agency leadership while the skill to audit a decision and advocate for a misclassified client is taught at the tier below; it also reports professional-body guidance placing the duty to train on the employer rather than on the individual practitioner. Both are arguments about where authority sits relative to capacity. Neither reports a measured rate of either.

yang2026aAcademicSave

Yang, F., & Liechty, J. M. (2026). Building AI Literacy and Competency in Social Work. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_23

doi.org/10.1007/978-3-032-18443-6_23

Appears in: AI in Social Work (Springer, 2026)

Topics: ai-literacy, social-work, workforce

huang2026bAcademicSave

Huang, J., Yang, F., & Lee, J. (2026). Ethical Challenges and AI Governance in Social Work. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_22

doi.org/10.1007/978-3-032-18443-6_22

Appears in: AI in Social Work (Springer, 2026)

Topics: ai-ethics, ai-governance, social-work

ConceptualThe volume's sexual and partner violence chapter draws a scope line around predictive risk modelling and resta…

The volume's sexual and partner violence chapter draws a scope line around predictive risk modelling and restates it twice: such models are used in research contexts to study population-level risk factors, they are not intended for individual-level decision-making in clinical or legal settings, and they are not intended for practitioners to screen or label individuals directly in real-world settings. The chapter treats the distance between that declared scope and a case-level use as a central governance danger rather than as a modelling defect, and reports no measurement of how often the line is crossed.

fang2026aAcademicSave

Fang, Y., & Postmus, J. L. (2026). AI in Sexual and Domestic Partner Violence. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_8

doi.org/10.1007/978-3-032-18443-6_8

Appears in: AI in Social Work (Springer, 2026)

Topics: gender-based-violence, social-work

ConceptualThe volume's school social work chapter names function creep, a term it takes from Koops (2021), as the long-t…

The volume's school social work chapter names function creep, a term it takes from Koops (2021), as the long-term risk that data collected for a beneficial purpose such as identifying mental health needs is later repurposed for an entirely different function; it names student discipline and sharing with external law-enforcement agencies as its examples. The chapter's stated concern is the absence of specific, renewed consent for the new use and the erosion of trust that follows, not the volume of data held. It offers this as a risk argument and reports no incidence rate.

huang2026aAcademicSave

Huang, J., & Stone, S. (2026). AI in School Social Work. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_10

doi.org/10.1007/978-3-032-18443-6_10

Appears in: AI in Social Work (Springer, 2026)

Topics: school-social-work, social-work, youth

EmpiricalAutomated content moderation fails in two directions at once, and which direction it favors is a governance ch…

Automated content moderation fails in two directions at once, and which direction it favors is a governance choice rather than a technical default. The volume's LGBTQIA+ chapter documents both halves landing on the same population: identity terms such as 'trans', 'queer' and 'nonbinary' have been flagged as inappropriate content while overt hate speech aimed at that population evades detection. The platform natural experiment already in this registry shows the same choice made explicitly rather than by default: with human review capacity withdrawn, the deployer said it would over-enforce rather than let harmful content stay up. The chapter is a peer-reviewed secondary synthesis and supplies the direction and the vocabulary, never a magnitude; the removal and reinstatement figures in this registry come from the platform's own transparency reporting under a separate claim and are not restated here. No outcome for the people whose content is moderated is computed anywhere in this Lab.

downey2026AcademicSave

Downey, D. L., & Jenkins, D. A. (2026). AI in Supporting LGBTQIA+ Populations. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_7

doi.org/10.1007/978-3-032-18443-6_7

Appears in: AI in Social Work (Springer, 2026)

Topics: lgbtqia, social-work

youtubegoogle2020GroundingVendorSave

YouTube / Google (2020, August 25). Responsible policy enforcement during Covid-19. Official YouTube blog. https://blog.youtube/inside-youtube/responsible-policy-enforcement-during-covid-19/

https://blog.youtube/inside-youtube/responsible-policy-enforcement-during-covid-19/

Appears in: PAN framework development

Grounds: domain grounding: content moderation and editorial AI (trust & safety, newsroom AI); model org: youtube_covid_enforcement

EmpiricalIn a school communication-monitoring deployment the flagged-content archive is itself the exposure pathway. Ev…

In a school communication-monitoring deployment the flagged-content archive is itself the exposure pathway. Every flag writes a durable record of a student's most sensitive writing, including disclosures of sexual orientation, and the investigative record already in this registry documents that archive being released unredacted through links that required no login, students who risked being outed after writing about sexual orientation or gender identity, and a district that discontinued the vendor in 2023 following an outing incident. The volume's LGBTQIA+ chapter supplies the mechanism that record instantiates: a store built to protect people is also the means by which they can be exposed, which is why practitioners are documented deliberately omitting sexual-orientation and gender-identity data from client information systems even at a cost to record completeness. This claim states institutional exposure only - what the record holds and which hand-offs it can travel along - and never a per-student outcome. Students sit outside these dynamics by construction and nothing about them is computed on any diagram.

downey2026AcademicSave

Downey, D. L., & Jenkins, D. A. (2026). AI in Supporting LGBTQIA+ Populations. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_7

doi.org/10.1007/978-3-032-18443-6_7

Appears in: AI in Social Work (Springer, 2026)

Topics: lgbtqia, social-work

bryanandlurye2025GroundingInvestigativeSave

Bryan and Lurye, Schools use AI to monitor kids. An investigation found security risks (The Christian Science Monitor, AP and Seattle Times Education Reporting Collaborative, 2025) https://www.csmonitor.com/USA/Education/2025/0312/ai-surveillance-schools-gaggle

https://www.csmonitor.com/USA/Education/2025/0312/ai-surveillance-schools-gaggle

Grounds: model org: gaggle_school_monitoring

EmpiricalCrisis-text triage classifiers read language, so the language they read least well is where they fail. The vol…

Crisis-text triage classifiers read language, so the language they read least well is where they fail. The volume's mental-health chapter reports that sarcasm, cultural idiom and code-switching confuse these models, that benchmark studies show markedly lower accuracy on African American English and other underrepresented dialects, and that the operational cost runs in both directions: a false positive can trigger an unwanted welfare check, a false negative leaves a texter waiting in silence. This is benchmark evidence about a class of classifier, not a measurement of any deployment in this registry. The underlying benchmark study is not held in this repository's reference snapshot, so no magnitude is carried, and the finding must never be attached to any named service's own published or unpublished figures. Nothing here is a fairness metric and nothing here is computed about any texter.

yang2026cAcademicSave

Yang, Y., & Traube, D. (2026). AI in Mental Health Services. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_9

doi.org/10.1007/978-3-032-18443-6_9

Appears in: AI in Social Work (Springer, 2026)

Topics: mental-health, social-work

ConceptualThe workplace chapter of the social work volume argues that algorithmic management — end-to-end digitalised ta…

The workplace chapter of the social work volume argues that algorithmic management — end-to-end digitalised task allocation, workflow organization, performance evaluation, scheduling and income distribution — is adopted for operational efficiency and cost reduction, and that the same systems erode frontline autonomy and make it hard for a worker to understand or question a decision about their work. The chapter reports no original measurement, so this is a documented direction and an argued mechanism, never a magnitude.

guo2026AcademicAcademicSave

Guo, P., & Hong, P. Y. P. (2026). AI in the Evolving Workplace. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_18

doi.org/10.1007/978-3-032-18443-6_18

Appears in: AI in Social Work (Springer, 2026)

Grounds: model org: amazon_fulfillment_management; model org: fortune500_agent_copilot; model org: klarna_ai_assistant

Topics: social-work, workforce

ConceptualThe volume's ethics chapter names five redress mechanisms as the operational answer to a diffuse responsibilit…

The volume's ethics chapter names five redress mechanisms as the operational answer to a diffuse responsibility gap: social impact assessment before deployment, audit trails documenting data inputs and decision processes, appeal mechanisms letting a person challenge an AI-driven decision, liability frameworks allocating responsibility by role, and ethics oversight committees seating practitioners, clients and technologists. It is a normative framework proposal, and no effect size is attached to any of the five.

an2026aAcademicSave

An, R., & Lindsey, M. A. (2026). Ethical Foundations of AI in Social Work. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_2

doi.org/10.1007/978-3-032-18443-6_2

Appears in: AI in Social Work (Springer, 2026)

Topics: ai-ethics, social-work

ConceptualThe health-care chapter renders the NASEM Assistance category as reach as much as throughput: post-discharge t…

The health-care chapter renders the NASEM Assistance category as reach as much as throughput: post-discharge texting that checks whether a patient obtained their medication and alerts a worker when something is wrong, and round-the-clock operation treated as timely aid, extend assistance beyond the clinic visit. The chapter states this as a practitioner expectation it argues for, and attaches no measured coverage, uptake or outcome figure to it.

ji2026AcademicSave

Ji, M., Xu, S., Yang, F., & Bachman, S. S. (2026). AI in Health Care and Health Disparities. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_5

doi.org/10.1007/978-3-032-18443-6_5

Appears in: AI in Social Work (Springer, 2026)

Topics: health-care, health-disparities, social-work

EmpiricalThe LGBTQIA+ chapter documents a governance trade-off practitioners already make: social work professionals in…

The LGBTQIA+ chapter documents a governance trade-off practitioners already make: social work professionals intentionally omit sexual-orientation and gender-identity data from client information systems to protect people from exposure, forced outing or violence. The chapter frames this as a considered deviation from data-completeness norms rather than a recording error, reports it from the literature it reviews, and gives no prevalence figure.

downey2026AcademicSave

Downey, D. L., & Jenkins, D. A. (2026). AI in Supporting LGBTQIA+ Populations. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_7

doi.org/10.1007/978-3-032-18443-6_7

Appears in: AI in Social Work (Springer, 2026)

Topics: lgbtqia, social-work

ConceptualMinimisation carries a cost the protective case usually leaves out: a record deliberately kept thinner is also…

Minimisation carries a cost the protective case usually leaves out: a record deliberately kept thinner is also a record that supports less verification, so minimising trades exposure against the evidence the correction loop itself runs on. The duty that travels with it is purpose limitation — consent obtained for one purpose does not cover reuse of that data to train a model for another — which is how the data-protection regulation states the two together. Direction only: none of these sources measures the size of either cost.

downey2026AcademicSave

Downey, D. L., & Jenkins, D. A. (2026). AI in Supporting LGBTQIA+ Populations. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_7

doi.org/10.1007/978-3-032-18443-6_7

Appears in: AI in Social Work (Springer, 2026)

Topics: lgbtqia, social-work

an2026aAcademicSave

An, R., & Lindsey, M. A. (2026). Ethical Foundations of AI in Social Work. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_2

doi.org/10.1007/978-3-032-18443-6_2

Appears in: AI in Social Work (Springer, 2026)

Topics: ai-ethics, social-work

europeanparliamentandcouncil2016GroundingRegulatorySave

European Parliament and Council of the European Union (2016). Regulation (EU) 2016/679 (General Data Protection Regulation), Article 5 — purpose limitation and data minimisation. https://eur-lex.europa.eu/eli/reg/2016/679/oj

https://eur-lex.europa.eu/eli/reg/2016/679/oj

Appears in: Evidence addition (2026)

Grounds: privacy law: purpose limitation and data minimisation (GDPR)

ConceptualThe disability chapter separates two implementation-stage failures that a single drift vocabulary blurs togeth…

The disability chapter separates two implementation-stage failures that a single drift vocabulary blurs together: concept drift, where the statistical properties of the data a deployed model processes change over time; and covariate shift, where the distribution of input features in the deployment environment differs from the distribution in the training data. They are named and defined as distinct mechanisms with different remedies, and neither is quantified.

wang2026aAcademicSave

Wang, J., & Begg, M. D. (2026). AI for Physical, Cognitive, and Developmental Challenges. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_6

doi.org/10.1007/978-3-032-18443-6_6

Appears in: AI in Social Work (Springer, 2026)

Topics: disability, human-ai-interaction, social-work

pham2025GroundingPeer-reviewedSave

Pham, T.M.T., Premkumar, K., Naili, M., & Yang, J. (2025). Time to Retrain? Detecting Concept Drifts in Machine Learning Systems. In ICSE 2025. https://doi.org/10.48550/arXiv.2410.09190 https://arxiv.org/abs/2410.09190

doi.org/10.48550/arXiv.2410.09190

Appears in: PAN framework development

Grounds: domain grounding: industrial operations and QA (visual inspection, predictive maintenance)

ConceptualThe disability chapter is the one setting in the volume where the AI is the assistance itself rather than a de…

The disability chapter is the one setting in the volume where the AI is the assistance itself rather than a decision aid steering an institutional decision about a person, which is the case that most tests the operator-network boundary. The resolution the framework already carries holds: model institutional propagation, keep clinical and operations staff in the operator network, and record the assisted person's own outcome externally. No outcome for a served person is computed from any diagram here.

wang2026aAcademicSave

Wang, J., & Begg, M. D. (2026). AI for Physical, Cognitive, and Developmental Challenges. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_6

doi.org/10.1007/978-3-032-18443-6_6

Appears in: AI in Social Work (Springer, 2026)

Topics: disability, human-ai-interaction, social-work

ConceptualThe older-adults chapter supplies the harder substitution case: where virtual contact substitutes for face-to-…

The older-adults chapter supplies the harder substitution case: where virtual contact substitutes for face-to-face contact, or a monitoring device substitutes for a person looking, the party who stops looking can be an unpaid family caregiver rather than an employee the organization can train, roster or audit. Substitution drawn on a staff link assumes an authority relationship that does not exist in that case, and no lever in the catalogue reaches that person.

shen2026AcademicSave

Shen, J., & Jennings, S. (2026). AI in Serving Older Adults. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_12

doi.org/10.1007/978-3-032-18443-6_12

Appears in: AI in Social Work (Springer, 2026)

Topics: older-adults, social-work

ConceptualThe housing chapter names consequences that sit outside an operator-network model entirely rather than being m…

The housing chapter names consequences that sit outside an operator-network model entirely rather than being merely unmodelled attributes of people: spatial stigma and the disinvestment or gentrification pressure that follows an area being algorithmically labelled as declining on superficial visual indicators, and a landlord's willingness to rent. These are place- and market-level effects, and nothing in a governance diagram computes them.

shin2026AcademicSave

Shin, J. C., & Foster, K. A. (2026). AI in Communities at Risk and Housing Stability. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_13

doi.org/10.1007/978-3-032-18443-6_13

Appears in: AI in Social Work (Springer, 2026)

Topics: community, housing, social-work

EmpiricalThe advocacy chapter's systematic review screened 7,715 records, included 415 articles and described 80 of the…

The advocacy chapter's systematic review screened 7,715 records, included 415 articles and described 80 of them in detail, reporting cluster sizes across five method-application buckets. Those are counts of what has been built and published: the review reports no pooled effectiveness estimate and no risk-of-bias appraisal, and its own limitations section calls for randomised trials or rigorous observational studies to become standard practice. A count of applications measures activity, never effect.

alba2026AcademicSave

Alba, C., & McCoy, H. (2026). AI in Advocacy and Social Justice. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_16

doi.org/10.1007/978-3-032-18443-6_16

Appears in: AI in Social Work (Springer, 2026)

Topics: advocacy, social-justice, social-work

ConceptualThe volume's conclusion proposes four distinct evaluation layers for AI in social work: technical performance,…

The volume's conclusion proposes four distinct evaluation layers for AI in social work: technical performance, distributive impact, procedural fairness and experiential legitimacy. It is a normative proposal, not an empirical finding, and the chapter reports no original measurement of any layer. Read against this Lab: pathway and lever direction sit in the technical-performance layer, the external equity surface is distributive impact and is recorded rather than derived, and procedural fairness and experiential legitimacy have no surface here at all.

yang2026dAcademicSave

Yang, Y., An, R., & Lindsey, M. A. (2026). Conclusion: Challenges and Opportunities for AI in Social Work. In R. An & M. A. Lindsey (Eds.), Artificial Intelligence in Social Work: Bridging Technology and Humanity. Springer. https://doi.org/10.1007/978-3-032-18443-6_24

doi.org/10.1007/978-3-032-18443-6_24

Appears in: AI in Social Work (Springer, 2026)

Topics: ai-governance, social-work

Why a governance lab at all

The Lab draws on the sociotechnical simulation. The headline finding: the same AI, in three modeled office cultures, let errors stick at very different rates, roughly 75%, 20%, and 16%[]. The governance context, not the model, drove that gap. That is the pattern you play against here.

Environmental figures use published data current as of early 2026 to show scale, not the measured footprint of any real deployment.