Skip to content

PAN Lab example

A heavy-industry predictive-maintenance deployment

Ninety percent fewer false alarms while the crews still label

In a peer-reviewed heavy-industry study, crews labeled a sensor model's alarms, and false alarms fell about 90 percent. The gain lasts while crews stay engaged.

See more

The system is a sensor-based anomaly detector, run by the plant's maintenance and reliability staff, that raises alarms on patterns that come before equipment failures. On its own it raised many false alarms, so a second model, built from the crews' labels of past alarms, corrects them.

The case study

The case is documented in a peer-reviewed study by Hermansa and five colleagues, published in the journal Sensors in 2021. It covers coal crushers at a power plant and converter gantries at a steelworks, in Europe. The case file describes the study site as anonymized.

The headline result is a cut in false alarms of 90.25 percent on average, credited to the correction model. It is the best-measured quantitative benefit among the Lab's industrial case files. In this field, the peer-reviewed figures come from anonymized or smaller sites. Named manufacturers mostly report their gains through company and trade channels, not independent measurement.

How the loop works

The anomaly detector raised alarms, and maintenance crews investigated them. The crews labeled which alarms were real failures and which were false. Those labels became training data. A correction model built from them screened the detector's alarms. The case study credits the cut in false alarms to that correction model.

Over successive rounds of labeling and retraining, the false-alarm rate fell. The labels taught the system to stop raising alarms on patterns that had not been failures. The benefit was not a property of either model alone. It came from the loop between the models and the crews.

The case study also showed each alarm with an explanation made with SHAP, a method that shows which readings drove an alarm. It used these explanations to keep the operators' trust through the change.

How the loop can fail

The loop runs on the crews' continued engagement, and that fails in two opposite directions. Both are well documented in the research literature.

The first is alert fatigue. If too many false alarms arrive before the loop has tuned them down, crews stop trusting the alarms. They stop investigating and labeling them, and the loop stalls at its worst point. The benefit is most fragile at the start, when the false-alarm rate is highest and trust is easiest to lose.

The second is automation bias. If crews come to treat an alarm as a confirmed failure, not a hypothesis to test, they stop applying their own judgment. Their labels then echo the detector's own calls. A model retrained on labels that agree with it learns nothing and can drift.

Augury, a company that sells machine-health monitoring, wrote in June 2026 that false alarms destroy maintenance teams' trust in predictive maintenance alerts. It wrote that once that trust is gone, it is very hard to get back. It gave the example of one compressor with more than 50 AI detections in a year, none of them real.

What to govern

The thing to govern is the loop's calibration: the balance between the crews trusting the alarms and judging them for themselves. It has to stay in a narrow range. Crews need enough trust to keep responding, and enough independence that their labels still carry real judgment. Too little trust and the loop stalls. Too much and it collapses into automation bias.

The measured 90 percent is the reward for keeping the loop in that range. It is contingent, not a permanent property of the models. A deployment that reports the number without governing the loop is reporting a result it can lose.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A pathway is closed when mistakes stop passing along it.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. Understand the system, the tool that lowers other tools' prices, is not offered here, so every tool costs its full price. The cheapest combinations cost 7 of this case's 10 budget units. All of them pair Escalate checks with Check with a second model.

Either of those two at its stronger setting is enough. At standard strength, they need one more tool: Review on schedule, Keep skills sharp, or Train the staff. Without Check with a second model, the cheapest combination costs 8 and uses Review on schedule at its stronger setting.

More tools are not better here. The service target asks that the maintenance models stay useful to the crews' work. Every tool at once, ignoring the budget, misses the Service Targets Only targets. Some of the tools take something from the service, and together they take it below its target.

Under Service and Safety Targets and under All Governance Targets, this case is not fully addressable with the available tools. Both levels ask you to close every failure pathway. Six failure pathways can still pass mistakes on, whatever you choose, even with every tool at once.

Five of them are the feedback loop's own pathways, which the measured benefit runs through. Three carry the crews' labels: into the detector's retraining, into the correction model, and into the maintenance record. The other two are the maintenance record used for retraining and the crews reading the alarm history. The sixth is the detector reading the sensors.

None of the tools offered for this case acts on these six pathways. The tools offered are the ones that act on whether crews keep investigating, which the sources put at the center of this case. That is a finding about the deployment, not a flaw in your choices.

Stylized model of a documented deploymentIndustrial QA & operations AI

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Predictive-maintenance-class whose benefit is a fragile loop network: 7 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 6 assumptions
  • assumed

    This example draws the correction model as its own part, apart from the anomaly detector. The sources describe a stand-alone detector and a separate correction model built on the crews' feedback. The measured cut in false alarms belongs to the correction model, not the detector. The part that cut the false alarms is made from crews having gone out and looked. So it lasts as long as they keep doing that.

  • assumed

    This example draws two more parts the sources document. One is the equipment sensor readings, which the detector reads in place of anything a person wrote down. The other is the explanation shown with each alarm. The case study used these explanations to keep the operators' trust. This example assumes the crews face more alarms than they can easily handle. The crews did investigate and label the alarms, and the measured result depends on that. The maintenance record is not treated as sensitive personal data, because it holds alarms and actions about equipment.

  • baseline

    This example follows the predictive maintenance pattern in the case file. It does not rebuild the actual system. A peer-reviewed case study in heavy industry cut false alarms by about 90 percent through a closed feedback loop. Crews investigated alarms and labeled them real or false. A correction model built from those labels screened the detector's alarms, and the case study credits the cut to it. It is the best-measured quantitative benefit among the Lab's industrial case files, from an anonymized study site. This example draws the feedback loop as working, because the benefit is the loop, not either model alone. The figure is the case study's own peer-reviewed result, entered as such.

  • assumed

    The feedback loop between the crews and the models is also the weak point. This example draws a check on the loop's calibration that the sources do not show running. Calibration here means the balance between the crews trusting the alarms and judging them for themselves. The loop depends on crews staying engaged, and that fails two ways. The first is alert fatigue. Too many false alarms arrive before the loop tunes them down, so crews stop labeling and the feedback stalls at the noisiest point. The second is automation bias. Crews defer and stop judging, so the labels echo the detector's own calls and the retraining learns nothing. The benefit depends on a narrow range: enough trust to respond, and enough independence to judge.

  • assumed

    What can be governed here is the loop's calibration, not the models' raw accuracy. The measured 90 percent is the reward for keeping the balance of trust and judgment in its narrow range. It is a contingent result, not a permanent property of the models. A deployment that reports the number without governing the loop reports a result it can lose. The loop runs on human engagement, and that engagement erodes in both directions if nobody manages it.

  • assumed

    This example computes no equipment failure or safety outcome. It covers the plant's own parts only: the models, the crews, and the maintenance record. The equipment and the people who depend on it are outside it. The measured cut in false alarms, the feedback loop, and the risks of fatigue and over-trust come from the case file. Nothing in this diagram computes them.

What this example does not show

Show all 2 limitations
  • This example computes no equipment failure or safety outcome. It shows how mistakes can pass between the plant's own models, crews, and maintenance record. The equipment and the people who depend on it are outside it. The measured cut in false alarms, the feedback loop, and the risks of fatigue and over-trust come from the case file. Nothing in this diagram computes them.
  • The cut of about 90 percent in false alarms is the case study's own peer-reviewed result, entered as such. This example draws the feedback loop as working, because it produced the measured benefit. It also draws a check on the loop's calibration, the balance between the crews trusting and judging the alarms. The sources do not show that check running. The diagram computes no harm from that check going unrun.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A peer-reviewed heavy-industry predictive-maintenance case study achieved a large, measured reduction in false alarms — on the order of 90 percent — through a closed operator-feedback loop: the maintenance crews investigated the alerts, labeled which were real, and the model retrained on those labels, so the false-alarm rate fell sharply over successive rounds. This is the industrial-QA domain's best-measured quantitative benefit, and it comes from an anonymized study site rather than a named-manufacturer press release, which is the pattern across this domain — the peer-reviewed magnitudes are at anonymized or smaller sites, while the named deployments report their benefit through corporate and trade channels.

    empirical
    • Peer-reviewed Hermansa, M., Kozielski, M., Michalak, M., Szczyrba, K., Wróbel, Ł., & Sikora, M. (2021). Sensor-Based Predictive Maintenance with Reduction of False Alarms — A Case Study in Heavy Industry. Sensors, 22(1), 226. https://doi.org/10.3390/s22010226 https://pmc.ncbi.nlm.nih.gov/articles/PMC8749854/
  • The lesson the case carries is that the same closed feedback loop that produced the benefit is the thing that can break it, because the loop depends on the crews continuing to engage with the alerts — investigating them, labeling them, responding — and that engagement fails in two opposite directions. Alert fatigue: if too many false alarms arrive before the loop has tuned them down, crews stop trusting the alerts and stop responding, so the feedback the model needs to improve never arrives and the loop stalls. Automation bias: if crews defer to the alerts and stop applying their own judgment, the labels the model retrains on become an echo of its own calls rather than an independent check. Either way the loop degrades, so the measured benefit is contingent on the loop staying calibrated — enough trust that crews respond, enough independence that their labels still carry real judgment.

    empirical
    • Academic Romeo, G., & Conti, D. (2025). Exploring automation bias in human-AI collaboration: a review and implications for explainable AI. AI & Society. https://doi.org/10.1007/s00146-025-02422-7 https://link.springer.com/article/10.1007/s00146-025-02422-7
    • Vendor Wittbold, K. (2026, June 18). Why Your Team Has Stopped Trusting Their Predictive Maintenance Alerts. Augury blog. https://www.augury.com/blog/machine-health/why-your-team-has-stopped-trusting-their-predictive-maintenance-alerts/

Where this connects

Institutional pressures in this domain

  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Industrial QA & operations AI domain page.

Levers available here and the patterns behind them

Documented case histories