PAN Lab example
A heavy-industry predictive-maintenance deployment
Ninety percent fewer false alarms while the crews still label
Sensor analytics forecast equipment failures; crews investigate the alerts and label which were real, and the model retrains on the labels. Modeled on a peer-reviewed case study that cut false alarms ~90% through exactly this closed feedback loop - the domain's best-measured benefit, from an anonymized site. But the loop is the vulnerability: it runs on crews engaging, and fails two ways - alert fatigue (crews stop labeling, feedback stalls) and automation bias (crews defer, labels echo the model). The measured benefit is contingent on the loop staying calibrated.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Predictive-maintenance-class whose benefit is a fragile loop network: 7 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
A third documented component: the record describes a stand-alone anomaly detector and a separate correction model built on the crews' feedback, and the measured reduction in false alarms belongs to the second, not the first. PP-152 still folded the two into one node with the loop drawn as an edge; a collision with the routing org made that visible, and drawing the correction model is the fix - a correction to the derivation rather than a tie-break. It matters for reading the case: the thing that suppressed the false alarms is made out of people having gone and looked, so it lasts exactly as long as they keep doing that.
- assumed
Two documented components come onto the board. The sensor stream is drawn as its own inbound source at full strength, because this model reads machines rather than anything a person wrote down - which is what makes it a different animal from every records-fed model in the catalogue, and why a failing sensor and a failing machine arrive looking the same. The per-alert feature attribution is drawn as a mediating artifact, because the record names it specifically as what kept crews willing to go and look: an alert that can say why is one a person can argue with, rather than merely obey or ignore, and both of those are how this loop dies. A heavy workload against limited capacity - crews who did go and look, which is the only reason the measured result exists at all. The label store is no longer marked privacy-sensitive; it holds alerts and dispositions about equipment.
- baseline
This models the predictive-maintenance pattern documented in the case file - not a reconstruction of the actual system. A peer-reviewed heavy-industry case study cut false alarms by roughly 90% through a closed operator-feedback loop: crews investigated alerts, labeled real-vs-false, and the model retrained on the labels. It is the domain's best-measured quantitative benefit, from an anonymized study site - drawn as the present feedback loop, at a low level, because the benefit is the loop, not the model alone. The figure is the case study's own peer-reviewed result, entered as such.
- assumed
The same closed loop is the vulnerability, drawn as the latent oversight check, empty at baseline: the loop depends on crews continuing to engage, and that fails two ways. Alert fatigue - too many false alarms before the loop tunes them down, so crews stop labeling and the feedback stalls at the noisiest point. Automation bias - crews defer and stop applying judgment, so the labels become an echo of the model's own calls and the retraining learns nothing. The measured benefit is contingent on keeping the loop in a narrow band: enough trust to respond, enough independence to still judge.
- assumed
The governable variable is the loop's calibration, not the model's raw accuracy - the measured ~90% is the reward for keeping the loop in its band, a contingent result rather than a permanent property. A deployment that reports the number without governing the loop is reporting a result it can lose, because the mechanism that produced it runs on human engagement that erodes in both directions if unmanaged.
- assumed
No equipment-failure or safety outcome is modeled here. This Lab reads institutional propagation only, and the equipment and the people it serves are boundary-only. The measured false-alarm reduction, the closed-loop mechanism, and the fatigue-and-over-trust dynamics live in the case file, and are never computed from anything in this diagram.
What this example does not show
- No equipment-failure or safety outcome is modeled. The Lab reads institutional propagation only; the equipment and the people it serves are boundary-only, and the measured false-alarm reduction, the closed-loop mechanism, and the fatigue-and-over-trust dynamics live in the case file, never computed on this diagram.
- The ~90% false-alarm reduction is the case study's own peer-reviewed result entered as such; the closed feedback loop is drawn PRESENT (the measured benefit) and the loop-calibration against alert fatigue and automation bias as a latent check, not a computed harm.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A peer-reviewed heavy-industry predictive-maintenance case study achieved a large, measured reduction in false alarms — on the order of 90 percent — through a closed operator-feedback loop: the maintenance crews investigated the alerts, labeled which were real, and the model retrained on those labels, so the false-alarm rate fell sharply over successive rounds. This is the industrial-QA domain's best-measured quantitative benefit, and it comes from an anonymized study site rather than a named-manufacturer press release, which is the pattern across this domain — the peer-reviewed magnitudes are at anonymized or smaller sites, while the named deployments report their benefit through corporate and trade channels.
empirical- Peer-reviewed Hermansa, M., Kozielski, M., Michalak, M., Szczyrba, K., Wróbel, Ł., & Sikora, M. (2021). Sensor-Based Predictive Maintenance with Reduction of False Alarms — A Case Study in Heavy Industry. Sensors, 22(1), 226. https://doi.org/10.3390/s22010226 https://pmc.ncbi.nlm.nih.gov/articles/PMC8749854/
The lesson the case carries is that the same closed feedback loop that produced the benefit is the thing that can break it, because the loop depends on the crews continuing to engage with the alerts — investigating them, labeling them, responding — and that engagement fails in two opposite directions. Alert fatigue: if too many false alarms arrive before the loop has tuned them down, crews stop trusting the alerts and stop responding, so the feedback the model needs to improve never arrives and the loop stalls. Automation bias: if crews defer to the alerts and stop applying their own judgment, the labels the model retrains on become an echo of its own calls rather than an independent check. Either way the loop degrades, so the measured benefit is contingent on the loop staying calibrated — enough trust that crews respond, enough independence that their labels still carry real judgment.
empirical- Academic Romeo, G., & Conti, D. (2025). Exploring automation bias in human-AI collaboration: a review and implications for explainable AI. AI & Society. https://doi.org/10.1007/s00146-025-02422-7 https://link.springer.com/article/10.1007/s00146-025-02422-7
- Vendor Wittbold, K. (2026, June 18). Why Your Team Has Stopped Trusting Their Predictive Maintenance Alerts. Augury blog. https://www.augury.com/blog/machine-health/why-your-team-has-stopped-trusting-their-predictive-maintenance-alerts/
Where this connects
Institutional pressures in this domain
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Industrial QA & operations AI domain page.
Levers available here and the patterns behind them
- Review the riskiest first — Risk-tiered oversight
- Escalate checks — State-feedback vigilance
- Keep skills sharp — Deskilling-arrest mandate
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Upgrade model — Improve the model