Skip to content

PAN Lab example

New Zealand MSD Predictive Risk Modelling

Halted before it ran: a national child-risk model

A predictive risk model would have scored every child's likelihood of a substantiated maltreatment finding by age five, computed from linked benefit and child-protection records, and handed the number to frontline social workers. Modeled on New Zealand's MSD predictive risk modelling tool. The distinctive fact: it never ran. The control that mattered fired before deployment — a layered ethical and privacy review, and finally a minister who refused to authorize a two-year study that would have scored about 60,000 newborns and watched whether high-risk predictions came true.

Stylized model of a documented deploymentChild welfare & family services

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the NZ-PRM-class halted-before-deployment risk model network: 4 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the halted-before-deployment predictive-risk-modelling pattern documented in the New Zealand MSD PRM case file — not a reconstruction of the actual tool.

  • baseline

    The tool was never operationally deployed; the accuracy figures it is cited with come from development data and were never field-validated. This diagram is a rehearsal of the shape the design would have taken, not a record of a system that ran.

  • assumed

    The feedback loop from the record into the model is present at baseline: the training target ('substantiated maltreatment') is itself an agency decision, and prior substantiations were among the model's predictors, so the record's decision patterns re-enter the score — the proxy-and-feedback concern reviewers and critics flagged.

  • assumed

    The ethical-and-privacy-review node carries the documented ex-ante review structure — an independent ethical review, a dedicated Maori ethical review, and a Privacy Impact Assessment discussed with the Privacy Commissioner — a real pre-deployment gate rather than an operational per-case reviewer.

  • assumed

    Peer pathways are authored on both signs: screeners anchor each other informally, one national model would homogenize its blind spots across every child, and a low peer/supervisory second-read pathway survives while the independent second-model check starts closed — no such check was ever built.

  • baseline

    The review-to-practice check pathway is drawn live at baseline, unlike the dormant checks that mark most maps in this set: the documented pre-deployment gate — layered ethical and privacy review, and finally an accountable authority's refusal to sign off — actually held, and the tool never scored a live case.

  • assumed

    Reviewers and critics flagged that both the outcome and the predictors could embed existing bias against Maori and benefit-receiving families; that concern is documented in the case file. This Lab models institutional propagation, not demographics, and estimates no differential harm to served people.

What this example does not show

  • This tool was never operationally deployed. The diagram is a rehearsal of the shape its design would have taken, not a record of a system that ran; the accuracy figures it is often cited with come from development data and were never field-validated.
  • The concern reviewers and critics raised was demographic — that both the outcome and the predictors could embed existing bias against Maori and benefit-receiving families. The Lab models institutional propagation, not demographics, and estimates no differential harm to served people; that concern is documented in the case file and measured outside any diagram like this one.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • New Zealand's Ministry of Social Development commissioned a child-maltreatment risk-modelling tool that, on a 2012 development sample of 57,986 children and 132 selected variables, reported an area under the ROC curve of 76% and a top risk decile in which 47.8% had a substantiated maltreatment finding by age five; those figures come from development data rather than field performance, the tool was never operationally deployed, and a proposed two-year study that would have scored about 60,000 newborns was halted by the incoming Social Development Minister, who annotated the briefing papers 'Not on my watch! These are children not lab rats.'

    empirical
    • Academic Vaithianathan, Maloney, Putnam-Hornstein, Jiang, Children in the Public Benefit System at Risk of Maltreatment: Identification Via Predictive Modeling (American Journal of Preventive Medicine, 2013) https://csda.aut.ac.nz/__data/assets/pdf_file/0019/11926/children-in-the-public-benefit-system-at-risk-of-maltreatment1.pdf
    • Investigative NZ Herald, Anne Tolley scraps 'lab rat' study on children (2015) https://www.nzherald.co.nz/nz/anne-tolley-scraps-lab-rat-study-on-children/C7GIGYW2467HG327FKXFRJDPEM/
    • Investigative Otago Daily Times, Call to stop child abuse risk modelling study (2015) https://www.odt.co.nz/news/national/call-stop-child-abuse-risk-modelling-study
    • Investigative Mordaunt, Child protection workers are under pressure in NZ. Can predictive modelling help? (The Conversation, 2026) https://theconversation.com/child-protection-workers-are-under-pressure-in-nz-can-predictive-modelling-help-278298

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Child welfare & family services domain page.

Levers available here and the patterns behind them

Documented case histories