PAN Lab example
Accelerated Safety Analysis Protocol (ASAP Tool)
The score nobody sees: choosing who gets a second look in child protection
A city child-protection agency scores every open investigation at day 10 with a model it built itself, and ranks them. The top of that ranking fills a quality-assurance review list of about 3,000 cases a year, against about 50,000 investigations — so the model, not any person, decides who receives a second, harder look. The twist is who can see it. Families are not told. Their attorneys are not told. The caseworkers running the investigation are not told. And the transparency register states that the reviewers working the flagged list are not shown the scores either: they get cases, not numbers. The extra interviews, collateral contacts, referrals and consults that follow are documented into the same records the model scores from, and the agency told state auditors that because no score decides anything, there would be no basis for a complaint. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. With every tool the Lab currently offers, no affordable combination brings this system inside the win condition at those settings. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Concealed severe-harm quality-assurance allocator network: 10 components and 18 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 10 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the concealed post-hoc quality-assurance allocation pattern documented in the case file — a stylized network, not a reconstruction of the deployment's internal architecture.
- baseline
The heavy workload and limited staffing are both derived, not defaulted. Workload: the transparency register records a quality-assurance review channel of about 3,000 cases a year against about 50,000 investigations, a 16:1 ratio of work to second-look capacity, on a statutory stream that must respond to every hotline report (29.0% of investigations indicated in 2024, 31.1% in 2025). Capacity: the model replaced no working adjudication process — the register states no score is used determinatively and every decision remains human — and the public record is silent on how the review list was filled before May 2018, so the higher 'displaced a working process' value is not supported and is not asserted.
- baseline
The two ingest baselines are separated on documented weight. Administrative case history is drawn at full strength because the model trains on and scores from it exclusively, daily, across every active investigation, and was retrained in 2025 on 815,284 training and 205,241 test observations from 2012-2022 closed investigations. The state hotline and case register is drawn at a substantial level because it is a documented but partial feature family — the count of hotline calls on a family and their duration in minutes — that also determines which cases exist to be scored.
- baseline
The model-to-reviewer channel is drawn at full strength because the ranking is the sole documented determinant of which cases the quality-assurance unit reviews, with no competing selection channel in the record. The withholding of the scores themselves is drawn as structure rather than as a weaker coupling: it is the score-stripping review list, the case-file-is-the-whole-basis read at full strength, and the score-to-record write held empty.
- baseline
The quality-assurance peer check is drawn faint by arithmetic from documented coverage: the second look is a real review of case practice and documentation where it lands, and it lands on about 3,000 of roughly 50,000 investigations, roughly 6 per cent. It is not drawn empty, because the review is documented to happen.
- baseline
Binding framing ruling — the oversight graph is blocked, not inert, and is drawn that way. The model-facing review carries a live inbound, drawn faint (annual self-filed register, a state audit that obtained technical documentation and bias-testing results, and the external stakeholder advisory group of data scientists, legal advocates, affected individuals and contract providers that the agency told auditors reviews the model at development and at every revision) and a live outbound check, drawn faint (the audit's documented commitments to keep evaluation and update logs and to formalise a policy). The advisory group, the quarterly internal usage reports, the claimed 2024 governance policy and the bias-testing results are all agency accounts with no published methodology and no independent replication, and are labelled as such wherever they appear.
- baseline
The inspector-general channel is drawn as a record read, faint, and a cross-case policy review, empty, matching the two documented access states: case history after a child fatality did reach it in 1 of the 18 such matters in 2025 (13 of 16 blocked in 2024, 19 of 25 in 2023), while it reports that the same state confidentiality provisions prevent it from conducting policy reviews of systemic-unfairness allegations including racial disparities, and unfounded-report and alternative-response records are closed to it entirely.
- assumed
Labelled inference, per the binding ruling: the inspector general's May 2026 report never mentions this algorithm or any algorithm. Connecting its documented record-access severance to this model is a topological inference — strengthened, but not proven, by the model's 2025 expansion into the alternative-response record class the report says is prohibited to that office. Nothing on this network should be read as the inspector general having found anything about the model.
- baseline
The score-to-record write is drawn empty and the pathway is kept on the map rather than omitted. The register states scores are not shared with staff in the quality-assurance unit or the investigative unit, and the agency reports that after 2025 the only identifier the pipeline retains is the case number, so the record accumulates the flag's consequences without carrying the flag. Keeping the pathway visible is what lets a disclosure or provenance remedy act on it.
- baseline
The staff-to-record write is drawn at full strength as the loop's write side: the directed interviews, collateral contacts, referrals and consults become new activity in the same records that supply the model's prior-involvement variables, and the 2025 retrain undertaken to address drift arising in part from policy and practice changes is the agency's own evidence that operator behaviour reaches the model. Scoped per the ruling: the agency has never studied the effect of the flag-triggered review on downstream case outcomes such as removals, which is a narrower statement than 'no monitoring' — it told auditors it produces quarterly internal reports on the model's use in the quality-assurance programme.
- baseline
The monoculture self-loop is drawn at full strength because one model scores every child in every active investigation citywide every day, with the case taking the highest child score, and the agency's own technical audit states the training data likely included at least some implicit and systemic biases and that geographic variables may act as partial proxies for race — so a skew repeats across the whole caseload rather than scattering case by case.
- baseline
The model-side independent check is drawn empty because the record contains no second model, no external replication and no published precision or recall; the advocacy analysis is agnostic about what auditing was done rather than affirmative that none was, and the agency kept no logs of performance evaluations or model updates as of the 2019-2022 audit fieldwork. The agency's claim that the model beat experienced caseworkers with fewer false positives and more equity across race and ethnicity is an agency account with no published methodology and no independent verification, and a comparison against people is not a second model.
- assumed
Version discipline, per the binding ruling: the 279-variable feature set, the 2013-2014 severe-harm training cohort and the 18-month outcome window describe the earlier audited version; the 24-month window, the day-10 daily cadence and the 2012-2022 retrain describe the version in the current register. No parameter here mixes the two. The canonical selection rate used throughout is about 3,000 of about 50,000 investigations a year, not the roughly equivalent 200-300 cases a month figure.
- assumed
Four elements are deliberately absent because the evidence does not document them, and each absence is a real difference from a catalogue sibling rather than a gap. No enforcement node: no record automatically drives an action here, and the agency states decisions are made by staff and never directly by a machine. No external boundary or egress pathway: inputs are exclusively agency administrative data, the register lists no vendor, and no data leaving the governed system appears anywhere in the record. No guardrail and no automated output screen: nothing sits between the model and the people acting on its output, and the withholding of the scores is concealment rather than a check. No store-to-store replication: the record classes the inspector general cannot reach are access categories, not copies.
- assumed
Served families and children are not in these dynamics and no outcome for them is computed here. A flag, a review slot or a directed interview is an institutional signal, never a person. The systemwide disparity figures in the case file — families reported and children removed at several times the rate of white families and children — describe the child-welfare system as a whole and are not measured properties of this model's flag distribution; no flag-level demographic breakdown has been published.
What this example does not show
- The oversight around this deployment is blocked rather than empty, and the diagram draws it that way. A state audit obtained the technical documentation and bias-testing results and secured commitments; the agency files the tool in a public register each year; and it told auditors it maintains an external stakeholder advisory group of data scientists, legal advocates, people affected by the system and contract providers that reviews the model at development and at every revision. Those channels run on this network. What the record does not contain is any instance of an oversight body changing the model's operation.
- Every accuracy, equity and bias-audit claim here belongs to the agency. Its statement that the model outperformed experienced caseworkers with fewer false positives and more equity across race and ethnicity, its race and ethnicity bias-testing results, its claimed 2024 AI governance policy, its quarterly internal usage reports and its post-2025 case-number-only identifier are all agency accounts with no published methodology, no public figures and no independent replication. No public precision or recall exists for this model in any form.
- The inspector general's May 2026 report documents that five state confidentiality provisions deny, limit or delay its access to the agency's records, and that it obtained the full history in 1 of the 18 child fatalities with prior agency involvement reported to it in 2025. That report never mentions this algorithm or any algorithm. Drawing its severed record access onto this network is an inference, strengthened but not proven by the model's 2025 expansion into the alternative-response record class the report says is prohibited to that office.
- Two precision points the sources require. First, the never-conducted study is specifically of downstream case outcomes, including removals — not of all monitoring, since the agency reports quarterly internal reports on the model's use in the review programme. Second, no two model versions are mixed here: the 279 variables, the 2013-2014 training cohort and the 18-month outcome window describe the earlier audited version, while the 24-month window and the 2012-2022 retrain describe the current one. The selection rate used throughout is about 3,000 of about 50,000 investigations a year.
- A city council package passed unanimously in November 2025 legislated an independent office to audit and monitor city algorithmic tools and investigate complaints, and press coverage tied it directly to the investigation of this model. Its implementation status was unverified as of 2026-07-20, so it is not drawn as a working oversight channel on this network.
- Families and children are not modeled here, and no outcome for them is computed. The Lab models institutional propagation, not demographics, and estimates no differential harm to served people. The systemwide disparities recorded in the case file — families reported and children removed at several times the rate of white families and children — describe the child-welfare system as a whole, are not measured properties of this model's flag distribution, and are measured outside any diagram like this one.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Since May 2018, New York City's Administration for Children's Services has scored every open child-protection investigation at day 10 with an in-house machine-learning model (documented as the ASAP Tool / Severe Harm Predictive Model), rank-ordering cases by predicted likelihood of substantiated physical or sexual abuse within 24 months to fill a quality-assurance review worklist of about 3,000 of roughly 50,000 investigations a year; the LL35 register states scores are not shared with staff in the QA unit or the investigative unit, families and their attorneys are not told when a case is flagged, ACS told state auditors there would be 'no basis for a complaint' about a predictive model on an individual case, and the agency's own internal audit acknowledged the training data likely included implicit and systemic biases, that geographic variables may act as partial proxies for race, and that flag predictions are more likely to be incorrect than correct.
empirical- Government New York City Office of Technology and Innovation, Agency Compliance Reporting of Algorithmic Tools Calendar Year 2025 final version dated March 27 2026, Department of Social Services entry for the Homebase Risk Assessment Questionnaire (2026) https://www.nyc.gov/assets/oti/downloads/pdf/reports/LL35%20Report%202025%20-%20Final%20-%202026-03-27.pdf
- Investigative Lecher, The NYC Algorithm Deciding Which Families Are Under Watch for Child Abuse (The Markup, 2025) https://themarkup.org/investigations/2025/05/20/the-nyc-algorithm-deciding-which-families-are-under-watch-for-child-abuse
- Government New York State Comptroller, New York City Office of Technology and Innovation: Artificial Intelligence Governance (Report 2021-N-10) (2023) https://www.osc.ny.gov/files/state-agencies/audits/pdf/sga-2023-21n10.pdf
The May 2026 'Access Denied' report by the NYC Department of Investigation — the Charter-mandated inspector general for ACS — documents that five provisions of NY Social Services Law, as applied by state OCFS, routinely deny, limit, or delay DOI's access to ACS child-welfare records, with unfounded-report and CARES records prohibited entirely; DOI was barred from the full case history in 17 of the 18 child fatalities with prior ACS involvement reported to it in 2025 (13 of 16 in 2024; 19 of 25 in 2023). The report does not mention the algorithm — the linkage is a topological inference — but in 2025 ACS expanded the model to score CARES alternative-response cases, the record class state law prohibits DOI from accessing; the NYC Council's GUARD Act (passed unanimously November 25, 2025) legislated an Office of Algorithmic Data Accountability whose implementation status remains unverified as of mid-2026.
empirical- Government New York City Department of Investigation, Access Denied: Challenges to DOI's Oversight of the City's Child Welfare System (Release 10-2026) (2026) https://www.nyc.gov/assets/doi/reports/pdf/2026/10ACSReport.Release05.05.2026FINAL.pdf
- Government New York City Office of Technology and Innovation, Agency Compliance Reporting of Algorithmic Tools Calendar Year 2025 final version dated March 27 2026, Department of Social Services entry for the Homebase Risk Assessment Questionnaire (2026) https://www.nyc.gov/assets/oti/downloads/pdf/reports/LL35%20Report%202025%20-%20Final%20-%202026-03-27.pdf
- Trade press StateScoop, New York City Council passes landmark AI oversight package (2025) https://statescoop.com/ny-city-council-passes-landmark-ai-oversight-package/
The ACS severe-harm model trains on and scores exclusively from the agency's own administrative records (including hotline call counts and durations, prior involvement, and geography), and the intervention its flags trigger — extra interviews, collateral contacts, service referrals, consults, and QA documentation follow-ups — writes new activity into those same records, enriching the prior-involvement features of any future report on the family; ACS has never studied what effect the flag-triggered extra review has on downstream case outcomes, including whether a child is ultimately removed from the home (while telling state auditors it produces quarterly internal reports tracking the model's use in the QA program), kept no logs of model performance evaluations or updates as of the 2019-2022 state-audit fieldwork, and its claim that the model outperformed experienced caseworkers with fewer false positives and more race/ethnicity equity is an agency self-claim with no published methodology or independent verification.
empirical- Investigative Lecher, The NYC Algorithm Deciding Which Families Are Under Watch for Child Abuse (The Markup, 2025) https://themarkup.org/investigations/2025/05/20/the-nyc-algorithm-deciding-which-families-are-under-watch-for-child-abuse
- Government New York State Comptroller, New York City Office of Technology and Innovation: Artificial Intelligence Governance (Report 2021-N-10) (2023) https://www.osc.ny.gov/files/state-agencies/audits/pdf/sga-2023-21n10.pdf
A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.
empirical- Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
- Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Review the riskiest first — Risk-tiered oversight
- Mark AI-written records — Provenance labeling
- Gate record entries — Human-in-the-loop write gating
- Escalate checks — State-feedback vigilance
- Understand the system — Understand the system
- Vet connections — Connection authorization
- Store less data — Data minimization
- Assign a challenger — Structured dissent
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Peer sharing rules — Peer-edge governance
- Upgrade model — Improve the model
Documented case histories
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- Douglas County Decision Aide
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Eckerd Rapid Safety Feedback: origin and spread
- Illinois Rapid Safety Feedback
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- Sistema Alerta Niñez (Chile)
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Oregon Safety at Screening
- Los Angeles County Project AURA
- What Works for Children's Social Care ML pilots
- New Zealand MSD Predictive Risk Modelling
- Gladsaxe model