PAN Lab example
What Works for Children's Social Care ML pilots
The bar it never cleared: a child-welfare prediction pilot
A government-funded evidence centre built prediction models to forecast whether a child's case would escalate, set a public success bar before it started, and scored every model against it. Modeled on England's What Works for Children's Social Care machine-learning pilots. None of the models cleared the bar, so none reached a caseworker. The rare case where the question 'does it actually work?' was asked first, in public, and answered honestly.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the WWCSC-class pre-deployment prediction pilot network: 4 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the pre-deployment evaluation pattern documented in the What Works for Children's Social Care machine-learning pilots case file — not a reconstruction of the actual research programme or its models.
- assumed
The model-to-decision coupling is drawn at the low end and is this shape's central counterfactual: the documented models were a research and feasibility build that never entered live casework, and the intended use was decision-support at a point of high social-worker discretion with documented low willingness to defer. The scenario rehearses the deployment decision the pre-registered evaluation ultimately answered in the negative.
- baseline
The pre-registered-evaluation node carries a real pathway of its own: the documented case scored its models against a published success bar before any go-live, so the check enters the dynamics rather than sitting inertly on the diagram. Whether the models cleared that bar is recorded in the case file, not computed here.
- assumed
The record-to-model loop is present at baseline: the training labels derive from past intervention decisions, so the model learns recorded practice rather than underlying risk — the feedback-loop concern the companion ethics review flagged. This is a documented risk in the data, not a measured disparity in these models.
- assumed
No demographic disparity figures were published for this programme; any differential harm to children and families here is a documented bias risk in the feedback loop, not a measured outcome. The Lab models institutional propagation only, and any such harm is documented in the case file and measured outside any diagram like this one.
What this example does not show
- The models here were never deployed in live casework; this scenario rehearses the deployment decision the evaluation ultimately answered in the negative. Bias enters this shape as a documented risk in the feedback loop, not as a measured disparity in these models; the Lab models institutional propagation only, and any differential harm to children and families is documented in the case file and measured outside any diagram like this one.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
None of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities cleared the pre-specified 65% average-precision success bar; the best single model reached only about 42% average precision and, at an operating point, missed roughly 79% of the children whose cases actually escalated.
empirical- Academic Clayton and Sanders, Can Machine Learning Save Children at Risk? (Significance, Royal Statistical Society) (2022) https://academic.oup.com/jrssig/article/19/6/22/7072840
- Trade press Community Care (Turner), 'No evidence' machine learning works well in children's social care, study finds (2020) https://www.communitycare.co.uk/2020/09/10/evidence-machine-learning-works-well-childrens-social-care-study-finds/
- Government evaluation ChildHub (Terre des hommes), Machine learning in children's services: does it work? (library record) (2020) https://childhub.org/en/child-protection-online-library/machine-learning-childrens-services-does-it-work
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Escalate checks — State-feedback vigilance
- Require sign-off — Conformity assessment gate
- Review on schedule — Oversight cadence & retrospectives
- Gate vendor updates — Vendor quality gate
- Pause AI on alarms — Deployment circuit-breaker
- Upgrade model — Improve the model
- Review the riskiest first — Risk-tiered oversight
- Mark AI-written records — Provenance labeling
- Store less data — Data minimization
- Assign a challenger — Structured dissent
- Understand the system — Understand the system
Documented case histories
- What Works for Children's Social Care ML pilots
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- Douglas County Decision Aide
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Eckerd Rapid Safety Feedback: origin and spread
- Illinois Rapid Safety Feedback
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- Sistema Alerta Niñez (Chile)
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Oregon Safety at Screening
- Los Angeles County Project AURA
- New Zealand MSD Predictive Risk Modelling
- Gladsaxe model