Skip to content

PAN Lab example

CORA, the DC CFSA policy assistant

The guardrail that named one direction

A child-welfare agency built a staff-facing policy assistant into its new case-management system and did the governance homework first: a sixteen-page alignment report published on the city technology register before launch, one of three district-wide, naming its accountable officials and its checking arrangements. Its centrepiece safety claim is a direction - the chatbot does not draw from confidential case files, and has no access to the case data. Seven months later the agency's own tip sheet describes a phone app that ingests case documents, photographs of handwritten notes and voice dictation, and saves AI-drafted contact notes into the permanent case record. Read the two together and the question stops being what the model can see. The record it answers from is checked by a named expert before anything enters; the record it now writes into has no such gate. Nobody outside the agency has examined the running system, and the agency's own report states that outputs cannot be checked before they are acted on. Before you pick a target level: this board cannot be won under All Governance Targets. With every tool the Lab currently offers, no affordable combination brings this system inside the win condition at that setting. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore, Service Targets Only, and Service and Safety Targets can be won.

Stylized model of a documented deploymentChild welfare & family services

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the CFSA-CORA-class curated-corpus policy retriever network: 10 components and 19 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 8 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the curated-corpus policy-retriever pattern the case file documents, not a reconstruction of the tool, and it makes no claim about the tool's accuracy in either direction. The whole network is a fictional-register model of institutional propagation; a safe-looking or unsafe-looking reading here is a property of this model, never a measurement of the deployment.

  • baseline

    Ten components, all documented, none decorative. The record names a retrieval layer, a curated knowledge store, the three source-document families it is assembled from with their custodial offices, a separate permanent case record, all agency staff as one operator tier and supervisors plus help desk plus training team as another, a dense in-house committee tier, and an external notice tier that received a pre-deployment document. Three kinds are deliberately absent for evidence reasons: no enforcement component, because the agency states the tool makes no rights-impacting decisions and the record shows no downstream action system it drives; no worklist, because no queue or backlog is documented anywhere; and no external boundary with data-leaving pathways, because the deployment is documented as running in a government community cloud under a District contract, so no pathway out of the governed system is documented and drawing one would invent an exposure the record does not support.

  • baseline

    The guardrail is drawn as structure, in the direction the record actually states it. The case-record store carries a retrieval pathway into the assistant, empty, quoting the agency's own centrepiece claim that the chatbot does not draw from confidential case files and has no access to the case data; the same store carries a machine-write pathway out of the assistant, at a substantial level, because the agency's December 2025 tip sheet describes AI-drafted contact notes being saved into that record once a worker approves them. The staff-to-model pathway is drawn at a substantial level and marked privacy-sensitive for the same reason: documents, photographs of handwritten notes and voice dictation tied to a selected referral enter through the worker, which is the direction a retrieval exclusion does not reach. The May 2025 report is quoted with its date throughout because its guidance-not-case-documentation statement is superseded by the agency's own later tip sheet, and that documented drift is the case's spine.

  • baseline

    The corpus ingest and its gate are the MiDAS shape inverted, and both strengths follow from the documented cadence. The source-to-store replication runs at a substantial level (a curated assembly refreshed as documents change, not a continuous automatic feed), and the reconciliation pathway back runs at a substantial level rather than empty or full: a subject-matter expert appointed by a deputy director checks every document before ingestion, which is mandatory and strong, while the committee's formal re-read of the whole store runs annually against a platform that shipped eleven numbered builds between July 2025 and June 2026. A mandatory gate on a twelve-month clock, over a system on a monthly one, is what a substantial level records.

  • baseline

    The retrieval spine runs at full strength because the record makes the store the whole of the assistant's knowledge: its sources are limited to the agency website, the policy index and the system training materials, and every answer cites one of those documents. The model self-loop runs at a substantial level rather than the catalogue's usual low setting for a documented reason rather than a stylistic one: uniformity is this deployment's stated purpose, not a side effect. The agency's own rationale is that staff wade through voluminous documentation producing variations in work product and that a chatbot answering uniformly should increase consistency, which is the same mechanism, read from the failure side, as one wrong checked document reaching every desk at once.

  • assumed

    The adoption arm is drawn at a substantial level and deliberately not at full strength. Exposure is designed to be near total (the report answers the user question with all agency staff, and training is mandatory and folded into onboarding), but the only volume figure in the record is a budget of twenty-five thousand queries per month, which is a spending ceiling rather than observed usage, and no usage, flag-rate, help-desk or escalation figure has ever been published. Pinning a saturated adoption arm on an unobserved ceiling would be an authored number wearing an evidence label, so the conservative value is used and the gap is recorded here.

  • baseline

    The per-answer citation and the in-product notice are drawn as the catalogue's third bounded automated output check, which the fidelity spec names as the honest shape for a screen that catches only a fraction. Two artefacts ride on every answer before a worker acts: a link to the document the answer came from, and the notice on the tool's own landing card that the Ask feature is under construction, that answers may not be fully accurate, and to consult a supervisor as needed. It is drawn faint rather than higher because it is a wrapper on the output rather than a reading of the output's content, so what it catches is what a reader chooses to follow up, and no catch rate, click-through rate or flag rate has been published. Automated output checks are bounded by construction; nothing here should be read as a substitute for a check of the answer itself.

  • baseline

    Checking is batch-mode by the agency's own admission and the map draws it that way rather than asserting it in copy. Asked whether humans can review and approve outputs before they are enacted, the report answers no: checking happens before content enters the store and afterwards through quality assurance. So the two operator-side check pathways run faint each, a supervisor consult that opens when a worker decides to raise something and a committee cycle that turns real-time flags and interaction logs into monthly internal accuracy reports, and the only per-output human step on the write path is the approval click on a drafted note.

  • baseline

    The two review tiers are separated because the record separates them, and the outside-review pathway is drawn empty for a documented reason rather than a rhetorical one. The in-house tier is wired in by a strong observability pathway (audit logs required for every interaction, plus telemetry on outcomes, knowledge-source usage and feedback) and every seat in it belongs to the deploying agency. The external tier is wired in faintly by the one transparency event that fired: a sixteen-page alignment report filed and published on the city technology register before launch, one of three district-wide, alongside a federal funding notification and a union briefing. What the record does not carry is any examination of the running system: no auditor or inspector-general review, no published accuracy or usage figures, no council finding on this tool, and no court monitor, the agency's federal court oversight having ended in 2021 with final closure in 2022.

  • baseline

    A heavy workload matched by ample capacity is derived from two documented sentences rather than from the modal pair. Workload: the agency's stated rationale is that staff wade through voluminous documentation causing variations in work product, in a child-welfare agency whose documented role includes fielding more than a hundred calls a day about children's safety (a vendor-story figure, labelled as such). Capacity: the agency's own stated contingency if the tool is unavailable is full reversion to the websites and hardcopy guides staff used before, and staff remain substantively responsible for their decisions. A deployer certifying in writing that its manual process still works is the strongest available evidence for a high counterfactual floor, and it is what lets this network read net-negative if its assistive arms are governed away.

  • assumed

    Every quantitative figure in the record is agency-stated or vendor-claimed and is labelled at each use. The time-savings and delivery-speed figures, forty-five minutes per intake report, one to four hours per case, and three weeks against four months at a claimed twenty-times lower cost, are unaudited vendor customer-story claims, and that story never names this tool, so tool-specific claims rest on the agency's own documents alone. The query figure is a budget ceiling. No independent audit, evaluation, accuracy dataset or usage dataset exists. The agency's equity analysis is self-assessed and no differential-impact or demographic testing of the tool is documented anywhere, so no demographic parameter enters this model and none may be inferred from it.

  • assumed

    Served families and children are not in the dynamics. This Lab reads institutional propagation only; an answer, a draft or an approved note on this map is an institutional signal and never a person. The agency's report names the harm pathway in its own words, inaccurate information contributing to a decision that leaves a child in an abusive or neglectful home, and rates that risk very low given source curation and human responsibility. That rating is an agency self-assessment recorded in the case file, and nothing on this diagram computes, confirms or contradicts it.

What this example does not show

  • No outcome for a served family or child is modeled. The Lab reads institutional propagation only; the families and children in these cases are boundary-only. The agency's own named harm pathway, inaccurate information contributing to a decision that leaves a child in an abusive or neglectful home, and its own rating of that risk as very low, live in the case file and are never computed on this diagram.
  • This is a self-governed deployment and every number in its record is agency-stated or vendor-claimed. The time-savings and delivery figures, forty-five minutes per intake report, one to four hours per case, and three-week delivery at a claimed twenty-times lower cost, come from a vendor customer story that never names this tool, so tool-specific claims rest on the agency's own documents alone. The twenty-five thousand queries a month figure is a budget ceiling, not observed usage.
  • The governance report is dated evidence and is quoted with its date, because it is partly superseded by practice. Its statement that generated text is guidance rather than case documentation predates the December 2025 feature that saves AI-drafted notes into the case record, and that drift is the case's spine rather than an inconsistency in the sources.
  • No independent examination of the running system exists: no inspector-general or auditor review, no published accuracy or usage data, and no council finding as of mid-2026, in the agency's first era without a federal court monitor in three decades. Absence of findings is absence of evidence, not clearance, and nothing here should be read as either an endorsement or a criticism of the tool's accuracy.
  • The agency's equity analysis is self-assessed and no differential-impact or demographic testing of this tool is documented anywhere. No demographic parameter enters this model and none may be inferred from it.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • DC's Child and Family Services Agency published a 16-page pre-deployment AI Values Alignment Report (one of only three district-wide under Mayor's Order 2024-028) for CORA, a staff-facing policy chatbot launched June 16, 2025 inside its STAAND case-management system, whose centerpiece guardrail is a read-side exclusion — 'the chatbot does not draw from confidential case files' and has no access to STAAND case data; by December 2025 the agency's own tip sheets documented a CORA phone app that ingests case documents, photos of handwritten notes, and voice dictation and saves AI-drafted contact notes into the STAAND case record after worker approval — the write direction the no-case-files rule never governed, superseding the May 2025 report's dated statement that generated text is 'guidance, not text to be used by the employee as part of case documentation.'

    empirical
    • Government DC Child and Family Services Agency, AI Values Alignment Report: CCWIS Case Operations Resource Assistant (CORA) Use Case (2025) https://techplan.dc.gov/sites/default/files/dc/sites/itstrategicplan/publication/attachments/OCTO%20Submission%20AI%20Values%20Alignment%20Report%20-%20CORA%20Use%20Case%20FINAL%20FOR%20PUBLICATION.pdf
    • Government DC Child and Family Services Agency, AI Assisted Contact Notes using CORA (tip sheet) (2025) https://cfsa.dc.gov/sites/default/files/dc/sites/cfsa/page_content/attachments/AI%20Assisted%20Contact%20Notes%20Tip%20Sheet%20Using%20CORA_Dec%202025.pdf
    • Government DC Office of the Chief Technology Officer, DC's AI Values and Strategic Plan (AI Values Alignment Report register) (2026) https://techplan.dc.gov/page/dcs-ai-values-and-strategic-plan
  • CFSA's May 2025 AI Values Alignment Report concedes that per-output human validation of CORA is impossible — asked whether humans can review and approve AI outputs before they are enacted, it answers 'No,' with validation occurring before content enters the knowledge source (SME validation of every document, annual committee re-review of the corpus) and afterward via quality assurance (monthly internal accuracy reports and a real-time answer-flagging channel, none published); by December 2025 the tool's own landing card carried the in-product warning 'The Ask Feature is under construction and answers may not be fully accurate. Consult your supervisor as needed,' delegating per-query vigilance to workers and supervisors six months after launch.

    empirical
    • Government DC Child and Family Services Agency, AI Values Alignment Report: CCWIS Case Operations Resource Assistant (CORA) Use Case (2025) https://techplan.dc.gov/sites/default/files/dc/sites/itstrategicplan/publication/attachments/OCTO%20Submission%20AI%20Values%20Alignment%20Report%20-%20CORA%20Use%20Case%20FINAL%20FOR%20PUBLICATION.pdf
    • Government DC Child and Family Services Agency, How to Use CORA (tip sheet) (2025) https://cfsa.dc.gov/sites/default/files/dc/sites/cfsa/page_content/attachments/CORA%20AI%20Tip%20Sheet%20FINAL%20121925.pdf
  • CORA operates as a single-agency closed loop with no external examination of the running system: its knowledge authors, tool owners, operators, and oversight committees are all CFSA units; no OIG or auditor review, no published accuracy or usage data, and no DC Council finding exists as of mid-2026; the deployment sits in CFSA's first era without a court monitor in three decades (LaShawn A. oversight ended 2021, final closure 2022); the 25,000 queries per month figure is a budget ceiling, not observed usage; and every quantitative outcome figure (45 minutes saved per intake report, one to four hours per case, three-week feature delivery at a claimed 20x lower cost) is an unaudited vendor claim from a customer story that never names CORA, so CORA-specific claims rest on agency sources alone.

    empirical
    • Government DC Child and Family Services Agency, AI Values Alignment Report: CCWIS Case Operations Resource Assistant (CORA) Use Case (2025) https://techplan.dc.gov/sites/default/files/dc/sites/itstrategicplan/publication/attachments/OCTO%20Submission%20AI%20Values%20Alignment%20Report%20-%20CORA%20Use%20Case%20FINAL%20FOR%20PUBLICATION.pdf
    • Vendor Microsoft, Washington DC CFSA transforms its systems with Microsoft Dynamics 365 and AI (customer story) (2025) https://www.microsoft.com/en/customers/story/25302-washington-dc-cfsa-microsoft-copilot-studio
    • Government DC Child and Family Services Agency and Executive Office of the Mayor, Mayor Bowser Announces the End of Court Oversight of the DC Child and Family Services Agency (2021) https://cfsa.dc.gov/release/mayor-bowser-announces-end-court-oversight-dc-child-and-family-services-agency
  • Documented benefit-automation failures replicated determinations into downstream systems with no independent reconciliation against the source records — Michigan MiDAS actioned replicated flags and Robodebt reversed the onus onto recipients.

    empirical
    • Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
    • Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
    • Government Royal Commission into the Robodebt Scheme, Report (2023) https://robodebt.royalcommission.gov.au/publications/report
    • Investigative Law Society Journal, Crude, cruel and unlawful: Robodebt findings https://lsj.com.au/articles/crude-cruel-and-unlawful-robodebt-royal-commission-findings/
  • A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.

    empirical
    • Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
    • Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
  • Automated catch fractions cap out below completeness — around 84% balanced accuracy in optimistic settings versus about 55% on hard content and 9.3% recall in worst-case measurements.

    empirical
    • Preprint Faithful RAG leaderboard (arXiv:2505.04847) — FaithJudge with o3-mini-high reaches ~84% balanced accuracy / ~82% F1 on FaithBench (optimistic ceiling). https://arxiv.org/abs/2505.04847
    • Preprint 'The Illusion of Progress' (arXiv:2508.08285) — LLM-as-Judge Precision 0.736 / Recall 0.957 / F1 0.832 vs human consensus on QA. https://arxiv.org/abs/2508.08285
    • Preprint Mental-health chatbot detection (arXiv:2604.06216) — GPT judges 54.6% accuracy, 9.3% recall (miss 90.7% of hallucinations); traditional methods F1<0.30 on subjective content. https://arxiv.org/abs/2604.06216
    • Industry Datadog LLM-as-a-judge (2025) — detection F1 drops substantially from HaluBench to the harder RAGTruth; harder hallucinations are harder to catch.
    • Peer-reviewed Same detection-accuracy literature as catch_at_generation (FaithBench arXiv:2410.13210; arXiv:2508.08285); audit-time detection is bounded by the same hallucination-detection ceiling.

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Child welfare & family services domain page.

Levers available here and the patterns behind them

Documented case histories