Domain Atlas / Child welfare & family services
The guardrail's blind side: DC's walled-off child-welfare chatbot that began writing into the case record
DC's Child and Family Services Agency published a 16-page pre-deployment AI Values Alignment Report (one of only three district-wide under Mayor's Order 2024-028) for CORA, a staff-facing policy chatbot launched June 16, 2025 inside its STAAND case-management system, whose centerpiece guardrail is a read-side exclusion — 'the chatbot does not draw from confidential case files' and has no access to STAAND case data; by December 2025 the agency's own tip sheets documented a CORA phone app that ingests case documents, photos of handwritten notes, and voice dictation and saves AI-drafted contact notes into the STAAND case record after worker approval — the write direction the no-case-files rule never governed, superseding the May 2025 report's dated statement that generated text is 'guidance, not text to be used by the employee as part of case documentation.'[3]
What happened
The Child and Family Services Agency (CFSA) built CORA — the Case Operations Resource Assistant — in house, through its Child Information System Administration, on a commercial low-code assistant platform hosted in a FedRAMP High government community cloud, with the platform vendor's services arm as implementation partner under a three-year District contract with option years. It launched June 16, 2025 inside STAAND, the agency's new federal-standard child-welfare information system, which replaced the 25-year-old FACES system. (CFSA's own materials render the STAAND acronym both as "Stronger Together Against Abuse and Neglect DC" and with an added "in"; both agency renderings exist, and the vendor's "Stand Together" rendering is the outlier.) CORA is not a risk scorer: it produces no scores, no per-family outputs, and no decision recommendations. It answers policy, procedure, and system questions from a curated corpus, gives step-by-step document guides, and identifies itself as non-human in every interaction; answers cite their source with a link to the document.
The governance work came first, and it was real. Under Mayor's Order 2024-028, CFSA filed and published a 16-page AI Values Alignment Report before deployment — dated May 22, 2025, the same date the tool appeared on the labor union's Labor Management Partnership meeting agenda — one of only three such reports published district-wide (alongside the city technology office's own assistant and a housing agency's rent registry). The report names its accountable individuals, describes a knowledge store limited to three open sources — the CFSA public website, the CFSA policy index, and STAAND training materials — and specifies the validation gate: every document is validated by a Deputy-Director-appointed subject-matter expert from the AI Steering Committee before entering the corpus, which the committee re-reviews annually. The centerpiece guardrail is structural: "The chatbot does not draw from confidential case files," and "CORA will not have access to the case data in STAAND." The agency classified the tool as "tangentially" rights-impacting — staff decisions it procedurally guides include health and risk assessments, service referrals, and child placement decisions, but "staff will remain substantively responsible for their decisions" — and named the harm pathway itself: inaccurate information contributing to "an inaccurate decision... that results in leaving a child in an abusive or neglectful home," rated "very low" given source curation and human responsibility. The report also conceded, in terms, what the architecture cannot do: asked whether humans can review and approve AI outputs before they are enacted, it answers "No." Validation is ex-ante (curation before content enters the source) and ex-post (quality assurance, monthly internal accuracy reports, a real-time answer-flagging channel) — never per-interaction. The agency's budget "funds 25,000 queries per month" — a spending ceiling, not a usage figure; no usage data has been published. The federal Administration for Children and Families was notified through an Advance Planning Document under child-welfare-system funding rules, and DC's AI Task Force won the Center for Public Sector AI's inaugural AI 50 Award — announced August 6, 2025 — with CORA cited as one of the District's three live AI projects.
Then the agency's own user documentation recorded the drift. The May 2025 report had stated that CORA generates "guidance, not text to be used by the employee as part of case documentation." By December 2025, CFSA's tip sheets — dated agency documents, which is what makes the drift trackable — describe something more: a CORA AI app on agency iPhones that lets a field social worker select an assigned referral or case, ingest documents, photos of handwritten notes, and voice-dictation transcripts, apply AI prompts (structure the note as Purpose, Content, Assessment, and Plan, summarize, improve tone, fix names, translate), and save the AI-shaped contact note into the STAAND case record, with workers "responsible for verifying the accuracy of their notes before approving them in final format." Case-specific content now enters the model through the operator — the direction the retrieval exclusion does not govern — and AI-drafted text flows out into the durable record that other workers, supervisors, and courts read later. The same December tip sheet shows an in-product warning on the Ask feature, in red text on the tool's own landing card: "The Ask Feature is under construction and answers may not be fully accurate. Consult your supervisor as needed" — a live accuracy disclaimer six months after launch, delegating per-query vigilance to workers and supervisors, with the STAAND Help Desk and training team as the designated recourse when guidance "does not make sense." The platform around the tool iterates fast — STAAND enhancement builds 2.1 through 3.1.1, eleven numbered builds, shipped between July 2025 and June 2026, roughly monthly — while the corpus's formal re-review runs annually. And the corpus includes STAAND tip sheets, which now include tip sheets about CORA itself: the tool's documentation recirculates through the tool.
What the record does not contain is any outside look at the running system. The oversight graph is dense but almost entirely internal: the AI Steering Committee (source validation, annual corpus review, downstream child-welfare metric monitoring), the AI Committee (monthly accuracy reports — internal and unpublished), IT change control, and a planned incident-response team. The external ties are thin and mostly pre-launch: the published report on the city register, the federal funding notification, the union briefing, and DC Council performance oversight, where no CORA-specific findings were located as of July 2026 — absence of findings, not clearance. There is no Office of the Inspector General (OIG) or auditor review, no published accuracy, usage, flag-rate, or escalation data, and — for the first time in three decades — no court monitor: the LaShawn A. class action, filed 1989, ended with court oversight concluding in 2021 and final closure after a data-validation period in 2022. Every quantitative outcome figure in the record is a vendor claim: the platform vendor's customer story reports 45 minutes saved per intake report, one to four hours per case, and feature delivery in three weeks instead of four months at a claimed 20x lower cost — unaudited marketing figures from agency interviews, in a story that describes the STAAND platform and its roadmap (voice dictation, AI training agents) but never names CORA, so CORA-specific claims rest on the agency's own documents alone. The tool remains live and expanding, with staff training calendars running through August 2026.
The sociotechnical reading
Every other child-welfare network in this Atlas is some variant of a scorer: a model emits a number or a flag about a family, and the drama is what the institution does with it. This case is the domain's first retriever, and its lesson is about a different governance object entirely — not the judgment but the plumbing. Read as a system, CORA's celebrated safety property is a statement about one direction of one edge: case files do not flow INTO the model. The agency built that wall carefully, published a governance report around it, and earned real credit for doing the homework before launch. The Atlas draws the wall honestly — a blocked read, latent at zero. And then it draws what the wall does not touch: within seven months, the documented growth of the system ran through the two directions the guardrail never mentioned. Case content flows into the model THROUGH THE OPERATOR — documents, photos of handwritten notes, voice dictation, tied to a named referral, uploaded by the worker the wall was never built to stop. And model output flows into the case record THROUGH THE WRITE PATH — AI-drafted contact notes, saved behind a single approval click into the permanent file that other workers, supervisors, and courts read for years. A guardrail that names one direction creates a governance shadow over the others; the May 2025 report's claim that generated text was 'guidance, not case documentation' was superseded by the agency's own December 2025 tip sheets, and that documented drift — not any adjudicated failure, for there is none — is the case's spine.
The second lesson is about verification architecture, and the agency stated it with unusual honesty: per-output human validation is impossible. Asked whether humans can review and approve outputs before they are enacted, the report answers 'No.' Everything, therefore, is batch-mode: an expert gate on documents entering the corpus (which makes error injection slow, centralized, and auditable — genuinely good design), an annual corpus re-review (against a platform shipping monthly builds — a decontamination clock running an order of magnitude slower than the thing it governs), and monthly internal accuracy reports (output-side, unpublished). Between the batches sits a disclaimer: 'answers may not be fully accurate. Consult your supervisor as needed,' in red text on the tool's own landing card. That sentence is the system's per-query verification plan — and it delegates vigilance to a workforce whose tool was deployed, by the agency's own rationale, to stop staff 'wading through voluminous documentation.' The consistency argument cuts both ways and the agency made only one cut: a chatbot that answers everyone uniformly does increase consistency — and it makes a single wrong validated document a correlated error, broadcast simultaneously to every operator in the building, in a system where the humans have been given both a reason to stop double-checking and a disclaimer that assumes they still will. The Lab draws this as the modelToModel uniform-broadcast loop and prices the deference levers accordingly.
The third lesson is the shape of the oversight column: a closed loop with a published preface. Knowledge authors, tool owners, operators, validators, and monitors are all units of one agency — the same organization writes the policies, encodes them, serves them, consumes them, and audits the result, down to the detail that the corpus now contains tip sheets about the tool itself. The external record is real but almost entirely pre-launch: the published report, the federal funding notification, the union briefing. Nothing outside the agency has examined the running system — no auditor, no published accuracy data, no council finding — and the deployment sits in the agency's first post-consent-decree era in three decades, so the historical backstop of a court monitor is gone precisely when the newest practice change shipped. This is not the concealed-allocator pathology of this domain's darkest cells; the agency told the public what it built, before it built it, in unusual detail. It is something subtler and more common: self-governance done diligently, whose diligence has never been tested from outside, and whose central safety claim was half-superseded by its own next tip sheet.
So the instruments that fit this cell are directional. Gate and mark the write path (write-gate, provenance-labels): machine-drafted text should enter a permanent case record bounded and labeled, because a court reading a contact note in two years cannot otherwise know a machine shaped it. Discipline the operator-side inflow (framing-hygiene, data-minimization): the do-not-enter-private-data warning, upgraded from a warning to a practice. Pin the connection surface (connection-auth): 'the assistant will not have access to the case data' should be an enforced property, not a sentence practice can drift past, and the vendor roadmap is already pushing. Hold the vigilance the disclaimer assumes (vigilance, deskilling-arrest), and put the batch-mode checks on a clock the agency does not control (oversight-cadence, peer-governance). The distinct lesson the Atlas draws here: a read-side guardrail is not a system guardrail — the honest question for any retrieval deployment is not 'what can the model see' but 'what can the system write, and who checks the direction nobody named.' The honest boundary throughout: served families and children are not modeled in the paired Lab; the agency's harm pathway and 'very low' rating are its own self-assessment; its equity analysis is self-assessed with no differential-impact testing documented anywhere; and every outcome figure is an agency or vendor claim, labeled at each use.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library.