Skip to content

PAN Lab example

Wisconsin DEWS

Wrong most of the time and unevenly by race

An ensemble model scores every grade 6-9 student's risk of not graduating and hands a 'high risk' label to school staff through a dashboard. Modeled on a statewide deployment an audit found wrong ~74% of the time on non-graduation, with higher false-alarm rates for Black and Hispanic students, staff untrained on the label, and the internal equity research unpublished; it was withdrawn in 2023. A transparent, low-tech indicator paired with real intervention accompanied a record graduation rate elsewhere. So watch what a label wired to a dashboard, not to help, actually does: change how a student is seen.

Stylized model of a documented deploymentEducation AI

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Early-warning-class whose label became a lens network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    The dashboard roster is drawn as its own artifact, because it is the whole of what this system delivered to a school. The catalogue keeps that kind of mediator as a component precisely so that 'a label, not an intervention' can be a statement about structure rather than a turn of phrase: a name arrives on a roster, and whether anything is attached to it is a separate question the diagram can now show being answered or not. A heavy workload against very limited capacity - every student in four grade levels across a state, read by staff the survey found had been given no training to read it. That pairing is the case: the volume was total and the capacity to interpret it was never built.

  • baseline

    This models the early-warning pattern documented in the case file - not a reconstruction of the actual system. A statewide dropout early-warning ML scored every grade 6-9 student and delivered a 'high risk' label to staff for about a decade. An independent audit found it wrong ~74% of the time on its non-graduation predictions, with higher false-alarm rates for Black and Hispanic students, and a district survey found staff reporting no training on interpreting the label; the state stopped publishing the dashboards in 2023. The audit figures are the independent investigation's own analysis, entered as such.

  • assumed

    The group disparity is drawn on the record-to-model training pathway: the model learned from student data including race and income, and produced higher false-alarm rates for Black and Hispanic students. The disparity is a recorded external finding, never computed on this diagram. A risk label is not an intervention (drawn on the model's self-loop): it is a signal that only becomes help if a trained human acts on it and a resourced response exists - wired to a dashboard rather than help, it is a lens on the student.

  • baseline

    The two governable absences are the latent checks. The published accuracy-by-group equity check (empty independent model check) - the department ran internal equity research but did not publish it, so the disparity was an outside finding rather than a governed metric. And the training + intervention loop (empty oversight check) - staff untrained on the label, the label wired to a dashboard not a resourced response. The counter-case (a transparent, interpretable on-track indicator paired with real intervention accompanying a record graduation rate) locates the benefit in this loop: an interpretable indicator wired to help outperforms an opaque model that only labels.

  • assumed

    No student outcome is modeled here. This Lab reads institutional propagation only, and the students being scored are boundary-only. The audit's accuracy and disparity findings, the missing training, the unpublished equity research, and the interpretable-indicator counter-case live in the case file, and are never computed from anything in this diagram.

What this example does not show

  • No student outcome is modeled. The Lab reads institutional propagation only; the students being scored are boundary-only, and the audit's accuracy and disparity findings, the missing training, the unpublished equity research, and the interpretable-indicator counter-case live in the case file, never computed on this diagram.
  • The ~74%-wrong and racial false-alarm-disparity figures are the independent investigation's own analysis entered as such; the disparity is drawn as a recorded external finding on the training edge, not a computed harm, and the equity audit and the training/intervention loop are drawn as two latent checks.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A state's statewide dropout early-warning system used ensemble machine learning to label every grade 6 to 9 student's risk of not graduating on time and delivered the label to school staff through dashboards for about a decade. An independent, decade-scale audit found the system was wrong roughly 74 percent of the time when it predicted a student would not graduate, produced higher false-alarm rates for Black and Hispanic students, and that the deployer's own internal equity research had gone unpublished — while a survey of districts found administrators reporting no training on how to interpret a 'high risk' label. The state stopped publishing the dashboards in 2023 and said it was evaluating the system's future. The deployment is the education domain's clearest case of a risk label whose error and group disparity entered how students were seen rather than the help they received.

    empirical
    • Investigative Feathers, T. (2023, April 27). False Alarm: How Wisconsin Uses Race and Income to Label Students 'High Risk'. The Markup (with Chalkbeat). https://themarkup.org/machine-learning/2023/04/27/false-alarm-how-wisconsin-uses-race-and-income-to-label-students-high-risk
    • Peer-reviewed Knowles, J.E. (2015). Of needles and haystacks: Building an accurate statewide dropout early warning system in Wisconsin. Journal of Educational Data Mining, 7(3), 18-67. https://doi.org/10.5281/zenodo.3554725 https://jedm.educationaldatamining.org/index.php/JEDM/article/view/JEDM082
    • Government Wisconsin Department of Public Instruction. WISEdash for Districts: Dropout Early Warning System (DEWS) Dashboards (including the October 12, 2023 retirement notice). https://dpi.wi.gov/wisedash/districts/about-data/dews
  • The lesson the case carries is that a risk label is only as good as the intervention it triggers and the training of the human who reads it. A label that is wrong most of the time, delivered to staff with no guidance on interpreting it, imports the model's error and its group disparity into how students are perceived rather than into a resourced response — the flag becomes a lens on the student rather than a trigger for help. Set against this, a large district's transparent, low-tech on-track indicator, built on interpretable research and paired with real intervention, accompanied a rise in graduation to a record level. The contrast locates the benefit in the intervention the indicator makes legible enough for staff to act on well, not in the sophistication of the prediction — an interpretable indicator that drives help can outperform an opaque model that only labels.

    empirical
    • Academic Allensworth, E.M., & Easton, J.Q. (2007). What Matters for Staying On-Track and Graduating in Chicago Public Schools. University of Chicago Consortium on School Research. https://consortium.uchicago.edu/publications/what-matters-staying-track-and-graduating-chicago-public-schools-focus-students
    • Investigative Feathers, T. (2023, April 27). False Alarm: How Wisconsin Uses Race and Income to Label Students 'High Risk'. The Markup (with Chalkbeat). https://themarkup.org/machine-learning/2023/04/27/false-alarm-how-wisconsin-uses-race-and-income-to-label-students-high-risk

Where this connects

Institutional pressures in this domain

  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.

All of them in context on the Education AI domain page.

Levers available here and the patterns behind them

Documented case histories