PAN Lab example
Colorado Family Safety and Risk Assessments
Two instruments, one decision, and a review that reached the legislature
Every screened-in child welfare referral in one state runs through two instruments at once: a structured safety determination a worker builds from field observation, and an actuarial risk scale computed largely from the family's stored history. Modeled on a statewide safety and risk assessment suite and the ombudsman-commissioned review a 2024 act required - its shape, not the real instruments. The two fail differently. Agreement on the safety judgment ran 94, 91 and 96 percent on abuse vignettes and 63 and 59 percent on hazardous-living and basic-needs vignettes, so the subjectivity concentrates where poverty is easiest to read as neglect. The actuarial scale climbs with any prior contact: the audit found it weights prior reports in a way that inflates scores for incidents that occurred long ago, that this disproportionately impacts minority families and perpetuates systemic bias, and that a decades-old domestic-violence incident is coded as recent. Meanwhile 68 percent of workers complete assessments away from the field, so the record the scale reads is reconstructed after the decisions it drove, and race and ethnicity are recorded inconsistently enough that the disparity analysis the legislature ordered came back limited. The oversight relay genuinely fired: complaints became testimony, testimony became a statute and an appropriation, and the statute became a 176-page review with 50 recommendations delivered to named committees on the deadline. Through mid-2026 that relay has changed statute, budget and public information, and not the instruments. So the question this board poses is which moves reach the graph, when the review has already reached the legislature. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. With every tool the Lab currently offers, no affordable combination brings this system inside the win condition at those settings. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Colorado-tools-audit-class dual-instrument assessment suite network: 8 components and 24 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 6 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the dual-instrument, statutory-audit pattern documented in the Colorado safety and risk tools case file. It is not a reconstruction of the state's actual instruments or record system, and not a claim about any family's case. Every node and edge is drawn from the audit's own description of the deployment; components the audit does not document - an external data feed, a downstream automatic action system, a queue, an automated output screen, an ungoverned egress path - are deliberately absent rather than drawn for shape.
- baseline
The two model nodes and their asymmetric feeds are the derivation's load-bearing choice. The audit documents two coupled instruments with two different error mechanisms feeding one decision: the actuarial scale reads stored history (prior reports, substantiations and removals are named scored inputs, and the audit found that weighting inflates scores for incidents that occurred long ago and disproportionately impacts minority families), while the consensus safety instrument reads the worker (five threshold criteria per danger item, with agreement at 94, 91 and 96 percent on the abuse vignettes against 63 and 59 percent on the hazardous-living and basic-needs vignettes). So only the actuarial channel carries a record-to-model edge, and the safety channel carries the strongest operator-to-model edge on the board.
- baseline
Pathway strengths are derived from documented frequency, coverage and mandate rather than from a template. Full strength is reserved for the three universal or verbatim-documented flows: the record-into-score feed the audit quotes, the worker judgment that constitutes the safety determination, and the write of every assessment into the record. The adoption pathways sit at a substantial level because 64 percent of the workforce and 52 percent of legal respondents called the instruments consistently the primary basis for a removal decision, which is most decisions rather than all. The discovery read and the review's case-file read sit faint because both are documented as thin: no clear sharing process for 76 percent of legal respondents, and 16 case files against the 65 or more of a standard federal review.
- baseline
Three checks are drawn at zero because the record documents them as not performed, not because the shape needed a gap. The instrument re-test: the ombudsman's 2023 brief records that the safety instrument has had no validity study since its 1999 inception while the 2014 review covered the distinct risk instrument, and the audit asks for updated validity and reliability studies on any redesign. The challenge from counsel: 94 percent of surveyed legal professionals report no involvement in instrument implementation or oversight. The legislative response: through 2026-07-20 no follow-up legislation, formal agency response or instrument change has been identified.
- baseline
The oversight relay is drawn exactly as far as the record carries it, per the binding framing ruling for this case: it has changed statute, budget and public information, and not the instruments. The legislative tier's inbound read is live because the act ordered a demographic and consistency read of the statewide record and the report returned to the named committees by the 1 March 2026 deadline; the review's check on practice is live at one step because the 176-page report, its 50 recommendations and four public information sessions exist. The return leg onto practice is drawn at zero. A decade earlier the state auditor's October 2014 performance audit flagged overlapping weaknesses without producing a redesign, the documented precedent for a review that reached the record and not the instruments.
- baseline
The heavy workload against limited staffing comes from documented workload against documented staffing. Both instruments are completed for every screened-in referral in 64 agencies on three clocks (24 hours for an unsafe determination, 14 days for record entry, 30 days for the risk level), and 68 percent of workforce respondents complete assessments away from the field, naming technology access at 30 percent, limited time during visits at 26 percent, instrument complexity at 24 percent, environment or client availability at 21 percent and travel at 4 percent. Capacity stays at the competent-human default because these instruments structure a judgment people were already making rather than replacing a working process, and the differential-response randomised trial behind the practice model found the structured track maintained safety at lower follow-up cost.
- baseline
The single-scale self-loop encodes correlated rather than idiosyncratic error: one actuarial scale scores every screened-in referral in all 64 counties, so a weighting choice inside it reaches the whole caseload. That is the structural mechanism behind the audit's finding that historical risk factors disproportionately elevate scores for families of color. It is a mechanism claim; no rate of correlated error is measured here.
- assumed
Served families and children, and any removal, reunification or safety outcome for them, are not in these dynamics. The Lab reads institutional propagation only, so a score, a determination or an approval on this map is an institutional signal and never a person. The audit's disparity findings are structural and qualitative findings about the instrument's mechanism, not measured outcome gaps: no administrative-data disparity analysis is published, because race and ethnicity are documented inconsistently in the record. No differential harm is computed from anything on this diagram.
- assumed
Survey figures carried into this derivation are non-probability, open-link perception signals with small valid strata (workforce 160 responses, of which 128 directly used the instruments; legal professionals 31 valid of 48 submitted; lived experience 15 valid of 22 submitted, all female parents, 12 of 15 from the metro region). Inter-rater agreement was measured on five hypothetical vignettes rather than on paired ratings of live cases, and the 16-file case review is illustrative by explicit design rather than a statewide rate. Where the audit's main body and an appendix table disagree on real-time completion, the main-body framing is used.
- assumed
No vendor is asserted for the actuarial scale: the audit names none, and its structured-decision-making-plus-actuarial characterisation is the audit citing a federal information gateway. The evaluator is a paid contractor of the ombudsman and states its report does not reflect that office's views, and the same office both raised the 2023 concerns and procured the review, so the relay is not fully arms-length. The instruments are regression-era structured and actuarial instruments codified in state rule; nothing here is machine learning or generative, and no model identifiers appear anywhere in this content.
- assumed
This deployment is distinct from the county-level predictive-analytics work the audit itself names as a separate phenomenon in two counties, with a third using a consulting tool. The Atlas models one of those separately; it shares no oversight relay with this statewide statutory suite, and nothing here should be read across to it.
What this example does not show
- Served families and children, and any removal, reunification or safety outcome for them, are not modeled here. The Lab reads institutional propagation only, so a score, a determination or an approval on this map is an institutional signal and never a person. The audit's disparity findings - that historical risk factors disproportionately elevate scores for families of color - are structural, qualitative findings about the instrument's mechanism, not measured outcome gaps: no administrative-data disparity analysis has been published, because race and ethnicity are documented inconsistently in the record. No differential harm is computed from these dynamics.
- The oversight relay is modeled exactly as far as the record supports and no further. It has produced a statute, an appropriation, a public 176-page review with 50 recommendations, and four public information sessions - and, through 2026-07-20, no enacted follow-up legislation, no formal agency response and no change to the instruments, which remain in statewide operation. Any reading that the review fixed or changed the instruments is premature by construction, and the return leg onto practice is drawn at zero for exactly that reason.
- The survey evidence is non-probability and open-link, with small valid strata (lived experience 15 valid responses, all female parents, 12 of 15 from the metro region; legal professionals 31 valid). Every percentage from those strata is a noisy perception signal rather than a population rate. Inter-rater agreement was measured on five hypothetical vignettes rather than paired ratings of live cases, and the 16-file case review traded breadth for depth by explicit design, so its proportions are illustrative rather than statewide rates.
- No vendor is asserted for the actuarial scale: the audit names none, and the structured-decision-making-plus-actuarial characterisation is the audit citing a federal information gateway. The evaluator is a paid contractor of the ombudsman, which both raised the original 2023 concerns and procured the review, and that dual role is carried here as a disclosed feature of the relay rather than an accusation. The instruments are regression-era structured and actuarial instruments codified in state rule; nothing here is machine learning or generative.
- This cell is distinct from the county-level predictive-analytics work the audit itself names as a separate phenomenon in two counties, with a third using a consulting tool. The Atlas models one of those separately, and it shares no oversight relay with this statewide statutory suite.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Colorado's legislature built a distinctive oversight topology — a standing independent statutory ombudsman (created 2010; an independent judicial-department agency since 2016 under C.R.S. 19-3.3) was handed a funded, nine-criteria audit mandate (HB 24-1046, signed May 28, 2024; $109,392 appropriation) over the statewide Family Safety and Family Risk Assessment tool suite used by 64 county-administered agencies — and the resulting ICF audit (176 pages, dated February 27, 2026, released March 2, 2026, with 50 recommendations across nine directives) found the tools well-aligned with policy on paper but inconsistently implemented: only 32% of surveyed caseworkers complete assessments in real time, vignette agreement fell to 59–63% on neglect and poverty scenarios versus 91–96% on abuse scenarios, and the audit recommends redefining safety and risk with observable behavior-based criteria and revising or replacing the actuarial risk tool.
empirical- Government evaluation ICF Incorporated, Colorado Safety and Risk Assessment Tools Audit: Evaluation of the Colorado Family Risk Assessment and Colorado Family Safety Assessment Tools (for the Office of the Colorado Child Protection Ombudsman, 2026) https://coloradocpo.org/wp-content/uploads/2026/02/Colorado-Family-Safety-Risk-Assessment-Tools-Audit-HB-24-1046_March-2-2026-FINAL-508.pdf
- Government Steffen, Audit Finds Issues with Colorado's Child Welfare Safety and Risk Assessment Tools (Office of the Colorado Child Protection Ombudsman news release via Pagosa Daily Post, 2026) https://pagosadailypost.com/2026/03/03/audit-finds-issues-with-colorados-child-welfare-safety-and-risk-assessment-tools/
- Government Colorado General Assembly, HB24-1046: Child Welfare System Tools, bill page (2024) https://leg.colorado.gov/bills/hb24-1046
- Reference Colorado Revised Statutes, Section 19-3.3-102: Office of the Child Protection Ombudsman Established, via Justia (2024) https://law.justia.com/codes/colorado/title-19/article-3-3/section-19-3-3-102/
Colorado's statutory audit loop fired end to end on the way up — the Child Protection Ombudsman's complaint-stream observations became testimony to the 2023 Child Welfare System Interim Study Committee (including its brief stating the Family Safety Assessment had never been validated since its 1999 inception), the testimony became HB 24-1046 requiring the ombudsman rather than the child welfare agency to procure a third-party audit, and ICF's competitively procured audit returned to named legislative committees by the March 1, 2026 statutory deadline with four public information sessions following — but as of July 2026 the loop has produced statute, budget, a public audit artifact, and public information only: no follow-up legislation has been enacted, no formal CDHS response has been identified, and the tools remain in statewide operation unchanged, a decade after the state auditor's October 2014 performance audit flagged overlapping assessment and documentation weaknesses without producing redesign.
empirical- Government Colorado General Assembly, House Bill 24-1046, enrolled act (2024) https://content.leg.colorado.gov/sites/default/files/documents/2024A/bills/2024a_1046_enr.pdf
- Government Child Protection Ombudsman of Colorado, Independent Audit of Colorado's Family Safety and Risk Assessment Tools, special initiative page (2026) https://coloradocpo.org/special-initiative/independent-audit-of-colorados-family-safety-and-risk-assessment-tools/
- Government Steffen, Audit Finds Issues with Colorado's Child Welfare Safety and Risk Assessment Tools (Office of the Colorado Child Protection Ombudsman news release via Pagosa Daily Post, 2026) https://pagosadailypost.com/2026/03/03/audit-finds-issues-with-colorados-child-welfare-safety-and-risk-assessment-tools/
- Government Office of Colorado's Child Protection Ombudsman, brief to the Colorado Child Welfare System Interim Study Committee, Hearing One, June 27, 2023 (2023) https://content.leg.colorado.gov/sites/default/files/images/office_of_colorados_child_protection_ombudsman_info_brief.pdf
- Government evaluation Colorado Office of the State Auditor, Child Welfare Performance Audit, October 2014 (2014) https://content.leg.colorado.gov/sites/default/files/documents/audits/1303p_-_child_welfare_performance_audit_october_2014_final_rev_11-3-14.pdf
- Government evaluation ICF Incorporated, Colorado Safety and Risk Assessment Tools Audit: Evaluation of the Colorado Family Risk Assessment and Colorado Family Safety Assessment Tools (for the Office of the Colorado Child Protection Ombudsman, 2026) https://coloradocpo.org/wp-content/uploads/2026/02/Colorado-Family-Safety-Risk-Assessment-Tools-Audit-HB-24-1046_March-2-2026-FINAL-508.pdf
The ICF audit found Colorado's actuarial Family Risk Assessment 'heavily weights prior reports, which inflate risk scores for incidents that occurred long ago, and disproportionately impacts minority families and perpetuates systemic bias' — with historical risk factors that 'disproportionately elevate scores for families of color' and a DV history item that codes any past incident, even decades old, as recent, penalizing survivors — while the record substrate degrades the oversight signal meant to catch exactly this: 68% of surveyed caseworkers cannot complete assessments in real time (documentation is back-filled into Trails after decisions), and race and ethnicity are so inconsistently documented that the audit says the disproportionality analysis requested by the legislature was limited; these are structural and qualitative audit findings, as no administrative-data disparity analysis has been published.
empirical- Government evaluation ICF Incorporated, Colorado Safety and Risk Assessment Tools Audit: Evaluation of the Colorado Family Risk Assessment and Colorado Family Safety Assessment Tools (for the Office of the Colorado Child Protection Ombudsman, 2026) https://coloradocpo.org/wp-content/uploads/2026/02/Colorado-Family-Safety-Risk-Assessment-Tools-Audit-HB-24-1046_March-2-2026-FINAL-508.pdf
A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.
empirical- Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
- Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Child welfare & family services domain page.
Levers available here and the patterns behind them
- Keep prompts neutral — Framing and mirroring reduction
- Mark AI-written records — Provenance labeling
- Vet connections — Connection authorization
- Gate record entries — Human-in-the-loop write gating
- Understand the system — Understand the system
- Review on schedule — Oversight cadence & retrospectives
- Assign a challenger — Structured dissent
- Peer sharing rules — Peer-edge governance
- Review the riskiest first — Risk-tiered oversight
- Check with a second model — Cross-model verification
- Upgrade model — Improve the model
Documented case histories
- The audit that reached the legislature before it reached the tools: Colorado's safety and risk instruments
- Allegheny Family Screening Tool
- Allegheny Hello Baby
- Douglas County Decision Aide
- The score nobody sees: New York City's concealed severe-harm QA algorithm
- Eckerd Rapid Safety Feedback: origin and spread
- Illinois Rapid Safety Feedback
- The vendor's ledger: Family-Match, the eharmony-derived adoption matcher the states kept coming back to
- ProKid (Netherlands)
- Insight Bristol / Think Family Database
- Hackney / Xantura Early Help Profiling
- Sistema Alerta Niñez (Chile)
- The map, not the score: place-based risk terrain and the records it concentrates
- The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
- US Birth Match
- Oregon Safety at Screening
- Los Angeles County Project AURA
- What Works for Children's Social Care ML pilots
- New Zealand MSD Predictive Risk Modelling
- Gladsaxe model