Skip to content

PAN Lab example

CA-CDS Child Abuse Alerting

The mandated report: alerting that writes into another organisation

An emergency department screens every child under 13 with a five-item nurse questionnaire, and a rule-and-text trigger reads the chart as it fills: the screen result, a free-text scan of the chief complaint and nursing notes, the orders, the diagnoses. A hit raises a dashboard icon and a pop-up recommending a complete guideline workup. Modeled on the Child Abuse Clinical Decision Support system built at UPMC Children's Hospital of Pittsburgh and taken multi-site from there. What makes this network unlike the domain's other bedside alerts is where its product goes. The system's working end product is not the alert — it is a mandated report, assembled from the chart and carried across an organisational boundary by a clinician who is not permitted to withhold it once suspicion forms. It lands in a state child protection agency's record. The agency must screen it, cannot see the chart it came from, and owes nothing back: the statute that compels the crossing provides no return leg, so the sending record never learns what the receiving one decided. Meanwhile the check that would perfect what crosses — a pre-checked order set that produced 100% guideline compliance when used — was used on a minority of encounters, and most surveyed practitioners did not recognise the icon that says the system fired. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. Cost is not what blocks it. Inside the budget the failure regime can be brought to calm and the benefit reads clear both margins; the pathway gate is the only gate that fails, and it fails at any budget — the whole lever list at once, costing two and a half times what you have, still leaves twelve pathways open. They are the chart reads and alert surfaces the screening itself runs on, the mandated report, and the receiving agency's own reads and writes. Closing the first set would mean switching the screen off. The second may not be gated by anyone. The third belongs to another organisation. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.

Stylized model of a documented deploymentClinical decision support & deterioration alerting

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the CA-CDS-class mandated-report child-abuse alerting network: 11 components and 25 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 6 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    Topology. Eleven components, all documented, none decorative, split across two organisations: a health system (the CA-CDS deployments at UPMC Children's Hospital of Pittsburgh, the UPMC general emergency departments, the University of Wisconsin and Northwell Health) and the state child protection agency in each jurisdiction — a statutory class, deliberately unnamed, because no source documents which agency received this deployment's reports and the receiving side is parameterized from national statute summaries and NCANDS figures rather than from any named agency. Two record stores, one per organisation; four groups of staff (two per organisation); a guardrail (the physical-abuse order set, a named component of the CDS build); a worklist (the screened-in response queue, with a published 102-hour mean clock); and two reviewers, each with a documented inbound read. The agency's intake and case record and its prior-involvement and central-registry holdings are drawn as ONE store at this granularity: no source documents error localising differently between them, no reconciliation between them is documented, and no lever offered here reaches either — so both documented facts, the arrival of the report and the durability of what persists, are carried by the one agency store. Three absences are equally derived: no input source (the trigger reads only the hospital's own encounter record — no live telemetry and no external feed, which is what separates this org from the domain's bedside-alert siblings), no enforcement component (the agency screens every referral by human decision against statutory criteria; nothing is automatically actioned from a record), and no external boundary (no unsanctioned tool, ungoverned host or undocumented replication out of the governed system appears in the record).

  • baseline

    The defining edge set. The crossing is drawn as two forward legs and one drawn absence: the mandated report as record content crossing store-to-store into the agency record; the same report as a human act — a clinician ringing a statewide 24-hour line; and the reconciliation check on the crossing drawn empty, because the statutory scheme specifies what a report must contain, how it is screened, who investigates and which third parties are notified, and establishes no duty to return the screening decision or disposition to the reporting clinician's record — and no source in this record documents a record-level write-back onto the hospital chart. Where a state informs a reporter of a disposition, that is a communication to a person, not a write to the sending store. The forward legs are compelled: the duty attaches to the individual practitioner, no institutional policy relieves it in 17 states plus DC and the Virgin Islands, employers may not discourage a report in 12 states, and the receiving agency may not decline to screen.

  • baseline

    Heavy workload against very limited capacity. Workload: a screen that is universal by design over every child under 13 at every covered emergency department, feeding a receiving side whose published load is national and worsening — 2,107,473 referrals screened in against an estimated 2,292,000 screened out in FFY 2023, on 5,936 intake workers and 21,739 response workers completing 66 responses each per year at a 102-hour mean first-contact time. Capacity very limited: the documented manual counterfactual is why the system exists — identification roughly quadrupled when the trigger went live at two general emergency departments (P < .001), guideline compliance on the general-ED path ran 33-57% before implementation, and the subspecialty consult that anchors manual practice is documented at the tertiary hospital and absent from the general-ED path that carries most of the volume. The domain's best-known measured clinician variability (22.5% of white versus 52.9% of minority children reported, at a different institution, 1994-2000) is context for why a universal instrument exists, and is never a measurement of this deployment.

  • baseline

    Baselines. The chart-to-trigger read is at full strength (universal by design), and it carries the physician orders that are among the trigger's own sources; the alert and screen pathways to the two hospital classes are at a substantial level, bounded by the measured visibility gap (69% did not recognise the dashboard icon, 27% could not recall an alert, 54% did not know the screen result was viewable) rather than by any documented rejection (only 4.5% disagreed with the recommendations); the order-set check is faint because it is near-complete where exercised (100% and 96% full compliance when used) and exercised on a minority of encounters (43 uses in a seven-month trial, 23 uses in an entire implementation period, 58% reported non-use); the consult-team pathways run faint to substantial because the service is documented present at the tertiary site and absent from the general-ED path; the mandated-report crossing is at a substantial level — compelled once suspicion forms, carried on a documented fraction of encounters (17% of triggered children reported at one system, 8.8% at another, reports rising 0.6% to 0.9% of children, P = .03); the agency intake screen is at full strength because every referral must be screened; and the agency-side reads and writes are at a substantial level against a published 102-hour clock and 66 responses per worker per year. Where the LEG-2 re-derivation merged two pathways, the survivor keeps its own documented rung — no baseline was moved by the coarsening.

  • baseline

    Record error, never detection. The model here is a rule-and-text trigger over the chart — no learned risk score, no abuse-likelihood model — and every error semantic on this diagram is a defect of a record or a workflow artifact: an alert fired on a chart the specification did not intend ('burn' in a 3-month-old's chief complaint; 'the father speaks broken English' matching 'broke'; 70% of audited overtriggers from the free-text scan, 81% of 242 triggers judged appropriate), a screen never completed (68-80% completion), a report that crosses without the evaluation the guideline names (full compliance 33-80% across sites and periods). The deployment's published sensitivity, specificity and predictive values are measured against a hospital child protection team's chart assessment — concordance with an expert record judgment, not detection of abuse in any child — and, following the PAN org's deliberate exclusion, no parameter on this diagram is drawn from them; they appear only in the case file, with that framing. No abuse likelihood, per-family outcome or substantiation reading exists anywhere in this network.

  • baseline

    The monoculture and its check. The model self-loop is at a substantial level: one trigger specification was rebuilt into three commercial record platforms (Cerner at the originating hospital, Epic and Allscripts at the disseminated systems), so the string-match defect repeats wherever it runs — held below full because each platform build is a local re-implementation and the record documents sharp site divergence on nominally identical software (81% of respondents wanting continued use at one system against 3% at another). The model-side check is faint: silent-mode operation preceded every go-live and a chart review judged 195 of 242 triggers appropriate — a real, episodic check, and the source of the overtrigger examples.

  • baseline

    The evaluation team's view is one-directional by derivation. The second reviewer receives chart audits from the hospital record and nothing from the agency store, because no evaluation in this record follows a report across the boundary to what the receiving agency did with it — the deployment's effect on the receiving organisation is unmeasured. Its outbound check on practice is faint: silent-mode runs, pre- and post-implementation audits and a published barrier survey, with authority to measure and publish, never to stop an alert or a report.

  • assumed

    Where the record is silent, the conservative value. No source publishes an operator correction or override rate for any class here, a per-alert defect rate, a report-content defect rate, or record-hygiene measures for the agency store; those magnitudes are estimated within the qualitative rungs the documented figures support, matching the PAN org's own estimated-on-both-counts discipline. The nurse and practitioner writes are at a substantial level on documented completion and use figures; the agency-side reads and writes, including the prior-involvement and central-registry checks they carry, rest on the statutory descriptions and the federal collection figures alone.

  • assumed

    Served children and families are not in the dynamics. Children seen in an emergency department, families reported to an agency, and people named in a durable agency record are boundary populations recorded in the case file. No node, edge, baseline or lever here computes an outcome for any child or family; the harm and benefit of reporting itself — protection, disruption, differential exposure — are documented questions the case file carries and this diagram never answers. The one deployment-level equity measurement in the record is a null (screening rates did not differ by patient or hospital characteristics across 13 emergency departments) and it is stated as a null, not as evidence of equity.

What this example does not show

  • Served children and families are not modeled. No abuse likelihood, per-family outcome or substantiation reading exists anywhere on this diagram; the network propagates record error through two organisations' operators and stores. Whether any report protected or harmed the family it named is a question this Lab never computes — and the record's own evaluations do not answer it either, because none follows a report across the boundary.
  • The trigger's published performance figures — sensitivity 96.8%, specificity 98.5%, positive predictive value 26.5% — are measured against a hospital child protection team's chart assessment. They are concordance with an expert record judgment, not detection of abuse in any child, and no parameter on this diagram is drawn from them.
  • The receiving side is a statutory class, not a named agency. No source in this record documents which state or county agency received this deployment's reports, and the agency side here is parameterized from national statute summaries and federal FFY 2023 collection figures. The named agency-side systems elsewhere in this atlas — call-screening and prevention tools operating on reports that have already arrived — are separate deployments at a separate boundary, and nothing here is wired to them.
  • No litigation and no regulatory action against this system appears anywhere in the record. The governing legal frame is not AI regulation at all: it is the mandatory-reporting statutes of each state under the federal Child Abuse Prevention and Treatment Act, which compel the individual clinician to report suspicion and compel the state agency to screen and respond.
  • The one deployment-level equity measurement in this record is a null: screening rates did not differ by patient or hospital characteristics across 13 emergency departments. No source publishes trigger, report or compliance rates disaggregated by subpopulation for this deployment, so whether the compulsory crossing happens at different rates for different groups of children — the question this topology most needs answered — is unmeasured. The domain's best-known disparity measurement comes from a different institution and era and is context for why a universal instrument exists, never a finding about this system.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The Child Abuse Clinical Decision Support system (CA-CDS), built by a consortium at UPMC Children's Hospital of Pittsburgh, combines a five-item nurse-administered screen for every child under 13, a free-text scan of the chief complaint and nursing assessment, physician orders, and discharge diagnoses into a rule-and-text trigger — no learned risk score — that raises a dashboard icon and a pop-up recommending a guideline-aligned physical-abuse order set. Implementation measurably moved volume: identification roughly quadrupled at two general emergency departments (P < .001), and triggering roughly doubled at both disseminated health systems (2.4% to 3.5% of children at the University of Wisconsin on Epic; 1.1% to 1.9% at Northwell Health on Allscripts). Its published performance — sensitivity 96.8%, specificity 98.5%, positive predictive value 26.5% — was measured against the hospital child protection team's chart assessment as reference standard: a concordance with an expert record judgment, not detection of abuse in any child. The documented failure mode is a string match: 'burn' in a 3-month-old's chief complaint is a trigger and so is 'burning up with fever'; 'broke' in a young infant triggers on 'the father speaks broken English'; 70% (33 of 47) of audited overtriggers came from the free-text scan, and 81% (195 of 242) of triggers were judged appropriate on chart review.

    empirical
    • Academic Berger, Heineman and Fromkin, Using Computer Alert Systems in the Emergency Room to Screen for Child Abuse (Patient-Centered Outcomes Research Institute final research report, 2019) https://www.ncbi.nlm.nih.gov/books/NBK602609/
    • Academic Feldstein, Barata, McGinn, Berger and colleagues, Disseminating child abuse clinical decision support among commercial electronic health records: Effects on clinical practice (JAMIA Open, 2023) https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10101685/
    • Academic Rosenthal, Skrbin, Fromkin, Heineman, McGinn, Richichi and Berger, Integration of physical abuse clinical decision support at 2 general emergency departments (Journal of the American Medical Informatics Association, 2019) https://academic.oup.com/jamia/article-abstract/26/10/1020/5518587
  • The CA-CDS deployment's working end product is a mandated report that crosses a one-way organisational boundary. An individual clinician, discharging a personal statutory duty under each state's mandatory-reporting laws pursuant to the federal Child Abuse Prevention and Treatment Act, transmits the report to the state child protection agency: health-care workers are designated reporters in 46 states plus DC and five territories, no institutional policy relieves the individual duty in 17 states plus DC and the Virgin Islands, and the standard is suspicion with no burden of proof. Implementation raised reports to child protective services from 0.6% to 0.9% of children at one health system (P = .03), with 17% of triggered children reported against 0.26% of non-triggered. The receiving agency must screen every referral; in federal fiscal year 2023 states screened in 2,107,473 referrals against an estimated 2,292,000 screened out, on 5,936 intake workers and 21,739 investigation and alternative-response workers completing 66 responses each per year at a mean first-contact time of 102 hours, with 28 states reporting increases and citing staff shortages and turnover. The statutory scheme establishes no duty to return the screening decision or disposition to the reporting clinician's record, and no source documents a record-level write-back onto the hospital chart; the report joins a durable agency record whose prior-involvement and central-registry checks are standard elements of the investigation that follows the next report about the same family.

    empirical
    • Government Child Welfare Information Gateway, Mandatory Reporting of Child Abuse and Neglect (State Statutes, current through May 2023), Children's Bureau, ACYF, ACF, U.S. Department of Health and Human Services https://artifacts.childwelfare.gov/public/documents/mandatory-reporting-abuse-neglect.pdf
    • Government Child Welfare Information Gateway, Making and Screening Reports of Child Abuse and Neglect (State Statutes), Children's Bureau, ACYF, ACF, U.S. Department of Health and Human Services https://artifacts.childwelfare.gov/public/documents/making-screening-reports-child-abuse-neglect_0.pdf
    • Government U.S. Department of Health and Human Services, Administration for Children and Families, Children's Bureau, Child Maltreatment 2023 https://www.acf.hhs.gov/sites/default/files/documents/cb/cm2023.pdf
    • Academic Feldstein, Barata, McGinn, Berger and colleagues, Disseminating child abuse clinical decision support among commercial electronic health records: Effects on clinical practice (JAMIA Open, 2023) https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10101685/
  • The CA-CDS record measures both its check and the gap over it. The physical-abuse order set produced 100% full guideline compliance when used at the originating hospital (43 uses in a seven-month trial; partial compliance fell from 10% to 3%, P = .04) and 96% (22 of 23 uses) at Northwell Health — while 58% of surveyed practitioners reported not using it at all. Of 71 practitioners analysed across 19 UPMC general emergency departments in February 2020, 69% did not recognise the dashboard icon indicating a trigger, 54% did not know they could view the screen result, 27% could not recall seeing an alert, and 65% were uncertain which tests to order; only 4.5% (3 of 66) disagreed with the recommendations, 75% said the tool raised awareness, 72% discussed alerts face-to-face with the child's nurse, and 54% named lack of social work or ancillary support as a barrier. Full compliance with the guideline evaluation at the moment of reporting ran 80% to 75% at Wisconsin and 33% to 50% at Northwell, and site acceptance diverged sharply on nominally identical software: 81% of Wisconsin respondents wanted continued use against 3% at Northwell.

    empirical
    • Academic Suresh, Saladino, Fromkin, Heineman, McGinn, Richichi and Berger, Integration of physical abuse clinical decision support into the electronic health record at a Tertiary Care Children's Hospital (Journal of the American Medical Informatics Association, 2018) https://pubmed.ncbi.nlm.nih.gov/29659856/
    • Academic Feldstein, Barata, McGinn, Berger and colleagues, Disseminating child abuse clinical decision support among commercial electronic health records: Effects on clinical practice (JAMIA Open, 2023) https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10101685/
    • Academic Peterson, Yealy, Heineman and Berger, Barriers to Adoption of a Child-Abuse Clinical Decision Support System in Emergency Departments (Western Journal of Emergency Medicine, 2024) https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11610743/
  • The only deployment-level equity measurement for the CA-CDS is a null: screening rates did not differ by patient or hospital characteristics across 13 UPMC general emergency departments. No source publishes trigger, report, or compliance rates disaggregated by subpopulation for this deployment. The domain's best-known disparity measurement — 22.5% of white versus 52.9% of minority children reported for suspected abuse among 388 children under 3 with acute fractures, and skeletal surveys ordered at 8.75 adjusted odds for minority toddlers — comes from a different institution (an urban academic children's hospital in Philadelphia) and era (1994-2000), and is the measured clinician variability a universal, instrument-driven screen is meant to standardise: context for this case, never a measurement of this system.

    empirical
    • Academic Berger, Heineman and Fromkin, Using Computer Alert Systems in the Emergency Room to Screen for Child Abuse (Patient-Centered Outcomes Research Institute final research report, 2019) https://www.ncbi.nlm.nih.gov/books/NBK602609/
    • Academic Lane, Rubin, Monteith and Christian, Racial differences in the evaluation of pediatric fractures for physical abuse (JAMA, 2002) https://pubmed.ncbi.nlm.nih.gov/12350191/

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.

All of them in context on the Clinical decision support & deterioration alerting domain page.

Levers available here and the patterns behind them

Documented case histories