Skip to content

PAN Lab example

Allegheny Family Screening Tool

The score and the screener: a child-welfare risk tool

A risk score built from old records lands in front of call screeners, who decide — except at the top of the scale, where the design makes the default call: above 17 with a child 16 or younger, screen-in is mandatory unless a supervisor signs a written justification, covering about a quarter of referrals. Modeled on Allegheny's Family Screening Tool (AFST). Watch two loops: memory, where today's decisions become tomorrow's inputs, and deference, where drift on the discretionary calls layers on top of the default the design already mandates.

Stylized model of a documented deploymentChild welfare & family services

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the AFST-class human-in-the-loop risk tool network: 4 components and 9 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the human-in-the-loop risk-tool pattern documented in the Allegheny Family Screening Tool (AFST) case file — not a reconstruction of the actual tool.

  • assumed

    Peer pathways are authored on both signs: screeners anchor each other informally, while the documented supervisory-review structure gives this shape a real inhibiting second-read pathway.

  • baseline

    The feedback loop from decisions into future scores is present at baseline, reflecting the case documentation's account of administrative-data training.

  • assumed

    The screening-supervisors node carries a real, lighter review pathway of its own — the documented supervisory-review structure over the screeners — so it enters the dynamics. Whether that review reduces documented disparity is recorded in the case file, not computed here.

  • assumed

    Screener discretion is real but anchored: the score shifts, not replaces, judgment.

What this example does not show

  • Bias propagates here the way failures do — through framings, records, and retrieval, as institutional workflow propagation. The Lab models no demographics and estimates no differential harm to served people; that harm is documented in the case files and measured outside any diagram like this one.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • Evaluation evidence on the Allegheny Family Screening Tool found that screener overrides of the tool's recommendations reduced racial disparity in screen-in rates relative to the tool alone.

    empirical
    • Government evaluation Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf
    • Academic Centre for Social Data Analytics (AUT), AFST evaluation summary https://csda.aut.ac.nz/news-and-events/2019/allegheny-family-screening-tool-evaluation-improved-decision-accuracy,-reduced-disparities
    • Academic Rittenhouse, Algorithms, Humans and Racial Disparities in Child Protective Services https://krittenh.github.io/katherine-rittenhouse.com/Rittenhouse_Algorithms.pdf
  • In the documented AFST evaluation, screener overrides of the tool — roughly a third of its recommendations — cut screen-in disparity from about 20% to 9% relative to the tool acting alone.

    empirical
    • Academic Rittenhouse, Algorithms, Humans and Racial Disparities in Child Protective Services https://krittenh.github.io/katherine-rittenhouse.com/Rittenhouse_Algorithms.pdf
    • Government evaluation Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf
    • Academic Centre for Social Data Analytics (AUT), AFST evaluation summary https://csda.aut.ac.nz/news-and-events/2019/allegheny-family-screening-tool-evaluation-improved-decision-accuracy,-reduced-disparities
    • Peer-reviewed Stapleton, L., Lee, M. H., Qing, D., Wright, M., Chouldechova, A., Holstein, K., Wu, Z. S., & Zhu, H. (2022). Imagining new futures beyond predictive systems in child welfare: A qualitative study with impacted stakeholders. 2022 ACM Conference on Fairness Accountability and Transparency, 1162–1177. https://doi.org/10.1145/3531146.3533177

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Child welfare & family services domain page.

Levers available here and the patterns behind them

Documented case histories