PAN Lab example
Los Angeles County Project AURA
Caught at the gate: a child-abuse risk model that never shipped
A proprietary model scores every abuse-and-neglect referral 1 to 1,000 from cross-agency records — but it is tested against past cases before anyone wires it to a live investigation. Modeled on Los Angeles County's Project AURA. The retrospective test is the whole story: at a high-risk cut it caught 171 of the worst-outcome children and flagged 3,829 who came to no harm, a false-positive rate of about 95.6%. The question this round asks is not how to run the tool safely — it is which controls decide, before go-live, whether it runs at all.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the AURA-class pre-deployment risk scorer network: 5 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 6 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
The documented event of this case is an evaluation, so it is modeled as one: a retrospective validation read the administrative record, set predictions against what had actually happened, and produced 171 true positives against 3,829 false ones. Both of its pathways are drawn present rather than empty because both actually happened - the read happened and the finding was acted on, ending the pilot before the score ever routed a live investigation. That exercised check is the one fact separating this record from the deployments that found out in production.
- assumed
This models the pre-deployment predictive-risk pattern documented in the Los Angeles County Project AURA case file — not a reconstruction of the actual tool.
- assumed
AURA was tested retrospectively against past cases and never used on a single live referral; this diagram models the shape it would have had if deployed, with the score anchoring investigation decisions — an authored modeling choice, not a record of live operation.
- baseline
The documented retrospective test produced 171 true positives against 3,829 false positives (about a 95.6% false-positive rate); the baseline treats the score as anchoring investigation attention while generating a false-positive volume that finite investigator capacity could not absorb.
- assumed
Peer pathways are authored on both signs, and this shape's defining absence is that both inhibiting checks start closed: no independent model ever cross-checked the opaque proprietary scorer, and no internal challenge met its rankings before the tool was shelved. Levers open them.
- assumed
The feature loop from accumulated cross-agency administrative records into the score is present at baseline, reflecting the case documentation's account of a model computed from administrative history collected for other purposes.
- assumed
The documented harm here is a false-positive-volume (saturation) problem, not a measured demographic disparity; the risk factors the model weighted are features, not an equity outcome. This Lab models institutional propagation, not demographics, and estimates no differential harm to served people.
What this example does not show
- AURA was never used on a live case; this scenario rehearses the shape it would have had, so the harm shown is counterfactual — the failure the tool would have produced, and the gate that stopped it.
- The documented harm here is a false-positive-volume problem, not a measured demographic disparity — the risk factors the model weighted are model features, not an equity outcome. The Lab models institutional propagation only; any harm to children and families is documented in the case file and measured outside any diagram like this one.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
In a retrospective test against historical outcomes, Los Angeles County's Project AURA — a proprietary risk model built by SAS — correctly flagged 171 of the highest-risk children but produced 3,829 false positives, a false-positive rate of about 95.6% that DCFS's own public-affairs director confirmed on the record, and the county shelved the tool in 2017 without ever using it on a live case.
empirical- Investigative The Imprint (Daniel Heimpel), Uncharted Waters: Data Analytics and Child Protection in Los Angeles (2015) https://imprintnews.org/featured/uncharted-waters-data-analytics-and-child-protection-in-los-angeles/10867
- Advocacy Child Protective Services Defense, Predictive Analytics in Child Welfare - Helping Hand, or Racial Bias? (Part 2) (2015) https://childprotectiveservicesdefense.com/predictive-analytics-child-welfare-helping-hand-racial-bias-2.html
- Advocacy NCCPR (Richard Wexler), Los Angeles County quietly drops its first child welfare predictive analytics experiment (2017) https://www.nccprblog.org/2017/05/los-angeles-county-quietly-drops-its.html
- Advocacy WitnessLA (Richard Wexler), LA County Nixes Alarmingly Unreliable Predictive Analytics Foster Care Scheme - For Now (2017) https://witnessla.com/op-ed-la-county-nixes-alarming-predictive-analytics-scheme-for-foster-care-for-now/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Escalate checks — State-feedback vigilance
- Require sign-off — Conformity assessment gate
- Gate vendor updates — Vendor quality gate
- Review on schedule — Oversight cadence & retrospectives
- Understand the system — Understand the system
- Review the riskiest first — Risk-tiered oversight
- Pause AI on alarms — Deployment circuit-breaker
- Upgrade model — Improve the model
- Check with a second model — Cross-model verification
- Assign a challenger — Structured dissent
- Store less data — Data minimization
Documented case histories
- Los Angeles County Project AURA
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- Douglas County Decision Aide
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Eckerd Rapid Safety Feedback: origin and spread
- Illinois Rapid Safety Feedback
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- Sistema Alerta Niñez (Chile)
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Oregon Safety at Screening
- What Works for Children's Social Care ML pilots
- New Zealand MSD Predictive Risk Modelling
- Gladsaxe model