Domain Atlas / Child welfare & family services
The score nobody sees: New York City's concealed severe-harm QA algorithm
Since May 2018, New York City's Administration for Children's Services has scored every open child-protection investigation at day 10 with an in-house machine-learning model (documented as the ASAP Tool / Severe Harm Predictive Model), rank-ordering cases by predicted likelihood of substantiated physical or sexual abuse within 24 months to fill a quality-assurance review worklist of about 3,000 of roughly 50,000 investigations a year; the LL35 register states scores are not shared with staff in the QA unit or the investigative unit, families and their attorneys are not told when a case is flagged, ACS told state auditors there would be 'no basis for a complaint' about a predictive model on an individual case, and the agency's own internal audit acknowledged the training data likely included implicit and systemic biases, that geographic variables may act as partial proxies for race, and that flag predictions are more likely to be incorrect than correct.[3]
What happened
The Administration for Children's Services (ACS) built its severe-harm risk algorithm in house — the Office of Research Analytics / Division of Policy, Planning, and Analysis; the city's LL35 transparency register lists Vendor(s): None — and has used it since May 2018. The register documents it as the Accelerated Safety Analysis Protocol (ASAP) Tool; the 2023 New York State Comptroller audit calls it the Severe Harm Predictive Model (SHM). Every day, at day 10 of each open child-protection investigation, the model computes each child's likelihood of a substantiated physical or sexual abuse allegation within the following 24 months; the case inherits the maximum score across its children, and cases are rank-ordered. The top slice fills the Quality Assurance unit's review capacity — about 3,000 cases a year out of roughly 50,000 investigations, some 6 percent (The Markup's reporting gives the equivalent monthly figure of 200 to 300). The architecture is post-hoc: the score never gates screening, response, or substantiation. It allocates an after-the-fact second-scrutiny channel over cases that are already open. The version described in ACS's internal technical audit used 279 variables — including the number and duration in minutes of hotline calls, prior ACS involvement, geography (borough, zoning area, community district), mother's age, sibling count and ages, and caretaker physical and mental-health indicators — trained on 2013-2014 cases that ended in severe harm; in 2025 ACS retrained the model on closed investigations from 2012-2022 (815,284 training and 205,241 test observations) to address data drift, and expanded it to also score CARES alternative-response cases. The Comptroller's audit described an 18-month severe-harm window in the earlier version; the current register says 24 months — the label window changed across versions, and the two versions' parameters should not be mixed.
What makes the deployment singular is who sees the output: no one. ACS tells neither families nor their attorneys when the algorithm flags a case — Brooklyn Defender Services' Nila Natarajan said nothing in discovery in individual ACS prosecutions "suggests or implies, and certainly doesn't lay out, use of any of these tools" — and it does not tell its own frontline caseworkers either. The register goes further: "Scores are not shared with staff in the QA unit or the investigative unit." The reviewers who work the model-selected worklist never see the model's numbers; the cases simply arrive. When a case is flagged, a senior specialist may convene an "action meeting" with investigators and direct additional steps — interviews with family members and collateral contacts, service referrals, consults with Family Court Legal Services and clinical experts (ACS spokesperson Marisa Kaufman) — and the QA team follows up with field offices on documentation and practice gaps. All of that activity writes into the same administrative records the model trains on and scores from, enriching the prior-involvement features of any future report on the family. ACS says it mixes current and past data to avoid relying on "historical patterns," but it has never studied what effect the flag-triggered extra review has on downstream case outcomes, including whether a child is ultimately removed from the home — though it told state auditors it produces quarterly internal reports tracking the model's use in the QA program, so the unmeasured question is specifically outcomes, not usage.
The agency's own internal technical audit, obtained by The Markup through a records request, acknowledged that the training data "likely included at least some implicit and systemic biases," that geographic variables "may act as partial proxies for race," and that because only a tiny number of cases lead to severe harm, predictions about the hundreds of flagged cases "are more likely to be incorrect than correct." Race and ethnicity labels are excluded from the features — the Comptroller found ACS was the only agency it sampled with detailed technical documentation, including race/ethnicity bias-testing results — but the same audit trail records what ACS could not show: no logs of model performance evaluations or updates were kept as of the 2019-2022 audit fieldwork, and ACS could not say how often the model had been revised or tested (it promised auditors it would keep formal logs going forward). ACS also claimed to auditors that head-to-head comparisons showed the model outperformed experienced caseworkers with fewer false positives and more equity across race and ethnicity — an agency self-claim with no published methodology and no independent verification — and that it maintains an external stakeholder advisory group (data scientists, legal advocates, individuals impacted by the city's child-welfare system, and contract providers) that reviews the model during development and whenever it is revised, likewise a claimed-but-unverifiable channel. Asked about public complaint mechanisms, ACS officials told the auditors that because no score is ever used in a determinative decision, "there would be no basis for a complaint related to the use of a predictive model on an individual case." The Electronic Frontier Foundation characterized the deployment as operating "in complete secrecy," with no mechanism for families to contest algorithmic flagging and no disclosed independent auditing — situating it alongside Allegheny County's contested family-screening tool — and ACS's Director of Predictive Analytics describes the models as flagging "high-risk cases for support," citing a 2024 AI governance policy grounded in "transparency, bias mitigation, and privacy safeguards" that has not been located as a public document. The system-level context is stark, though it describes the child-welfare system as a whole rather than the model's flag distribution, which has never been published demographically: Black families are reported to ACS at seven times the rate of white families, and Black children are thirteen times more likely to be removed from their homes.
The oversight record is dense, and blocked rather than inert. On May 5, 2026, Department of Investigation Commissioner Nadia Shihata released "Access Denied" (Release #10-2026), documenting that five provisions of New York Social Services Law, as applied by the state Office of Children and Family Services (OCFS), block DOI — ACS's Charter-mandated inspector general — from records essential to oversight: OCFS "routinely denies, limits, or delays" authorization, and unfounded-report and CARES records are prohibited to DOI entirely. In 2025 DOI was notified of 18 child fatalities where ACS had prior involvement within the last decade and was barred by state law from the full case history in 17 of them (13 of 16 in 2024; 19 of 25 in 2023); DOI says the same laws also prevent policy reviews of systemic-unfairness allegations such as racial disparities. The DOI report never mentions the algorithm — connecting it to the model is a topological inference, strengthened but not proven by the model's 2025 expansion into CARES cases, the very record class state law prohibits DOI from accessing. Meanwhile the city's legislature acted: on November 25, 2025, the Council unanimously passed the GUARD Act package (Council Member Jennifer Gutierrez), creating an independent Office of Algorithmic Data Accountability to audit and monitor city AI tools and investigate complaints, with press coverage tying the push directly to The Markup's May 2025 ACS revelations. As of mid-2026 the office's implementation status is unverified — a legislated but dormant check. The model remains operational: retrained, expanded, disclosed tersely in an annual register, scored daily against records its inspector general cannot fully read, by an agency whose stated position is that no complaint about it is possible.
The sociotechnical reading
The child-welfare cases in this Atlas mostly share a canonical seam: a risk score lands in front of a screener or caseworker, and the drama is what the human does with the number — defer to it, override it, drift toward it. This case deletes that seam on purpose, and its lesson begins there. Read as a system, the NYC deployment is a concealed allocator: the model's judgment never reaches any human AS a judgment. It reaches the organization as a worklist — cases, not scores, by the register's own words — and reaches the frontline as directives with the model's fingerprint stripped entirely, since caseworkers are never told a flag started the action meeting. Every operator in the chain is score-blind by design. The consequence is structural, not attitudinal: the override edge this domain's other networks argue about does not exist here. Nobody can discount a score they cannot see; families and attorneys cannot contest a flag that never appears in discovery; and the agency's position — that a non-determinative score gives 'no basis for a complaint' — converts the concealment into a doctrine. Score-blindness does not prevent deference. It perfects it: an allocation that arrives as fact, unaccompanied by the number that would let anyone weigh it.
The second fact is the ration. The QA unit can review about 3,000 of roughly 50,000 investigations a year. At that ratio the model is not advising a decision-maker — it IS the decision-maker for the only decision that matters in this channel: who receives compounding scrutiny. Threshold placement, not any human judgment, is the operative policy lever, and it moves silently with caseload arithmetic. And the scrutiny compounds through the record: flag-triggered interviews, collateral contacts, referrals, consults, and documentation follow-ups all write new activity into the same administrative store the model trains on and scores from daily — call counts, prior involvement, the features the agency's own audit flagged as carrying systemic bias and partial race proxies. The loop is unmeasured rather than refuted: ACS has never studied whether the flag-triggered review changes downstream case outcomes, including removals, even as its 2025 data-drift retrain documents that operator behavior does flow back into the model. The Lab draws this loop honestly: the score-to-record write is the case's defining absence, because scores are never written anywhere the people working the case would read them — so the model reaches its own future inputs only through the activity its flags trigger, which is exactly what makes the loop invisible to everyone inside it. The pathway is kept on the map rather than omitted, since it is precisely what a disclosure or provenance remedy would open.
The third fact is the oversight column, and the binding reading is blocked-but-not-inert. It would be too easy — and wrong — to draw a wholly dead graph. The state Comptroller did audit the tool, and found the paradox that defines self-governed deployments: ACS was the best-documented agency in the sample, with bias-testing results no peer could show, and it simultaneously kept no logs of evaluations or updates and could not say how often the model had been revised. The agency claims an external stakeholder advisory group reviews the model at each revision and that quarterly internal usage reports exist — claims recorded by auditors that no outside party has verified, carried here as the thin live check they are. What is severed is the deep read: the charter-mandated inspector general is statutorily blocked from the records — full access in 1 of 18 fatality matters in 2025 — and the model's 2025 expansion moved its input surface into the CARES record class the inspector general is prohibited from touching outright, a linkage the Atlas labels as its own topological inference since the oversight report never names the model. And the one response with enforcement design — the algorithmic-accountability office the Council legislated unanimously in November 2025 — remains unverified as operational: a check that exists on paper and has never fired. Net, across eight years: no documented instance of any oversight body changing the model's operation.
So the instruments that fit this cell are not the domain's usual ones. There is no visible score to calibrate a human against, no override channel to protect, no vendor to gate. The governable surfaces are the ones the concealment leaves: the hand-offs (make each hop carry its provenance, so the reviewers know they work an algorithmic allocation and the field knows a flag began the directive); the write side of the loop (gate and mark flag-triggered activity so the model stops reading the echo of its own allocations); the store's reconciliation (fund the outcome-grounding nobody has done); and the paper checks (put the claimed advisory review and the dormant audit office on a cadence the agency does not control). The distinct lesson the Atlas draws here: a model can capture an institution without ever showing anyone a number — concealment plus a capacity ration is a governance surface in itself, because when no operator, subject, or overseer can see the model's judgment, the only lever left inside the system is the threshold, and the only levers left outside it are the ones that govern the hand-offs, the writes, and the checks around a black box everyone is inside of. The honest boundary throughout: served families and children are not modeled in the paired Lab, which reads institutional propagation only; the system-level disparity figures are not flag-level measurements; every accuracy and equity claim in the record is an agency self-claim with no public methodology; and the 'more likely incorrect than correct' characterization is the agency's own internal audit's qualitative statement, not a measured rate.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library.