Skip to content

PAN Lab example

What Works for Children's Social Care ML pilots

The bar it never cleared: a child-welfare prediction pilot

A government-funded evidence centre built prediction models to forecast whether a child's case would escalate, set a public success bar before it started, and scored every model against it. Modeled on England's What Works for Children's Social Care machine-learning pilots. None of the models cleared the bar, so none reached a caseworker. The rare case where the question 'does it actually work?' was asked first, in public, and answered honestly.

Stylized model of a documented deploymentChild welfare & family services

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the WWCSC-class pre-deployment prediction pilot network: 4 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the pre-deployment evaluation pattern documented in the What Works for Children's Social Care machine-learning pilots case file — not a reconstruction of the actual research programme or its models.

  • assumed

    The model-to-decision coupling is drawn at the low end and is this shape's central counterfactual: the documented models were a research and feasibility build that never entered live casework, and the intended use was decision-support at a point of high social-worker discretion with documented low willingness to defer. The scenario rehearses the deployment decision the pre-registered evaluation ultimately answered in the negative.

  • baseline

    The pre-registered-evaluation node carries a real pathway of its own: the documented case scored its models against a published success bar before any go-live, so the check enters the dynamics rather than sitting inertly on the diagram. Whether the models cleared that bar is recorded in the case file, not computed here.

  • assumed

    The record-to-model loop is present at baseline: the training labels derive from past intervention decisions, so the model learns recorded practice rather than underlying risk — the feedback-loop concern the companion ethics review flagged. This is a documented risk in the data, not a measured disparity in these models.

  • assumed

    No demographic disparity figures were published for this programme; any differential harm to children and families here is a documented bias risk in the feedback loop, not a measured outcome. The Lab models institutional propagation only, and any such harm is documented in the case file and measured outside any diagram like this one.

What this example does not show

  • The models here were never deployed in live casework; this scenario rehearses the deployment decision the evaluation ultimately answered in the negative. Bias enters this shape as a documented risk in the feedback loop, not as a measured disparity in these models; the Lab models institutional propagation only, and any differential harm to children and families is documented in the case file and measured outside any diagram like this one.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • None of the 32 machine-learning models What Works for Children's Social Care built across four English local authorities cleared the pre-specified 65% average-precision success bar; the best single model reached only about 42% average precision and, at an operating point, missed roughly 79% of the children whose cases actually escalated.

    empirical
    • Academic Clayton and Sanders, Can Machine Learning Save Children at Risk? (Significance, Royal Statistical Society) (2022) https://academic.oup.com/jrssig/article/19/6/22/7072840
    • Trade press Community Care (Turner), 'No evidence' machine learning works well in children's social care, study finds (2020) https://www.communitycare.co.uk/2020/09/10/evidence-machine-learning-works-well-childrens-social-care-study-finds/
    • Government evaluation ChildHub (Terre des hommes), Machine learning in children's services: does it work? (library record) (2020) https://childhub.org/en/child-protection-online-library/machine-learning-childrens-services-does-it-work

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Child welfare & family services domain page.

Levers available here and the patterns behind them

Documented case histories