PAN Lab example
Automated visual inspection of injectable drugs
Erring toward the scrap heap while guarding the one that gets through
Machine learning inspects filled injectable drugs for defects, tuned to err toward scrapping good vials rather than passing a bad one. Inspectors remain the backstop.
See more
Machine learning automated visual inspection examines filled injectable drug products, in vials and syringes, for particle contamination and cosmetic defects. A drug manufacturer runs it on a task once done by human inspectors and fixed-rule cameras. Its quality process is validated, meaning formally shown to work as intended.
The trade-off, chosen on purpose
An inspection can make two errors, and here they are far from equal. A false accept passes a real defect, so a compromised injectable reaches a patient. That is a patient safety failure, the thing the whole inspection exists to prevent. A false reject scraps a good vial. That is a cost: wasted product and lower yield, but no patient harmed.
So the system is deliberately tuned to over-reject. The case file calls that the right call for a safety-critical line. It is a governance decision, not a neutral default. Someone chose which error to make, and the choice is defensible because it is owned.
Qualifying an AI inside a validated process
In regulated drug manufacturing, a change to an inspection step is not simply deployed. It must be qualified: shown, within the validated process, to perform as intended and to keep performing. For a fixed-rule camera this is well-trodden. For a machine learning model whose behavior can shift with the product mix, the sensors, or a retraining, it is harder.
The US Food and Drug Administration (FDA) published a discussion paper on AI in drug manufacturing in 2023. The case file describes the regulator's framework for validating and monitoring such AI as still being developed. So qualifying the model's behavior over time is an emerging check, not a settled one.
Public feedback to the FDA, summarized in a 2025 journal article, describes how manufacturers want model updates handled. It names risk-based change control under the quality system and people verifying the results. It prefers AI that changes only in approved steps to AI that keeps learning on its own.
The human backstop
The over-reject tuning lowers the visible risk but does not remove it. It makes a missed defect rarer, not impossible. The human inspector is the layer meant to catch the defects that still slip.
Over-trust wears that layer down. If inspectors treat the AI's pass as authoritative, they scrutinize less, because the machine has already looked. The misses the tuning was meant to guard against then get past the one layer there to catch them.
How the case file reads it
The case file calls this a well-governed use of AI in a safety-critical setting. The trade-off is chosen in the safe direction, and the deployment respects a regulated qualification process. Two governance questions stay open anyway: the AI-specific qualification the regulator is still framing, and the human backstop that over-trust can hollow out. The remaining risk is small and real. It is covered only if both hold.
What the sources do not show
The case file names no manufacturer and no product. Following the caution in the research behind this example, nothing here claims that a named manufacturer's AI let a defect ship. That research, done in July 2026, found no named manufacturer publicly attributing a shipped defect or a recall to its AI inspection.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A pathway is closed when mistakes stop passing along it. Each tool has a standard setting and a stronger one.
Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. At standard strength, the cheapest combinations cost 7 of this case's 11 budget units, with three tools each. One is Gate vendor updates, Escalate checks, and Check with a second model. The other is Gate vendor updates, Gate record entries, and Escalate checks. At its stronger setting, Gate vendor updates meets the targets with Escalate checks alone, for 5 units.
More tools are not better here. The service target asks that the inspection stay useful to the work it supports. Every tool at once, ignoring the budget, misses the Service Targets Only targets. Most of the tools take something from the service. Together they take it below its target.
Under Service and Safety Targets and under All Governance Targets, this case is not fully addressable with the available tools. Both levels ask you to close every failure pathway. Four pathways stay open whatever you choose, even with every tool at once.
They are the inspectors judging the AI's calls, the batch record used to update the AI, the inspectors reading batch history, and rejected units sent to scrap. None of the tools offered for this case acts on any of them. The tools offered are the ones the research sources behind this case show the manufacturer or its AI vendor could plausibly use. That is a finding about the deployment, not a flaw in your choices.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Regulated-visual-inspection-class tuned to over-reject network: 6 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 5 assumptions
- assumed
This example draws the cost side of the trade-off as an action: a rejected unit is scrapped. It also draws a second look at rejected units before scrapping. The case file does not describe such a second look. The research sources behind this case say inspectors re-inspect rejected units while they are still physically present, and describe human verification of rejects as required. Erring toward scrap is the right direction, and over-rejection is the one error this design makes by choice. Separately, industry feedback to the FDA describes risk-based change control under the quality system. It prefers AI that changes only in approved steps to AI that keeps learning on its own. This example draws that change control as a step every model update goes through. It assumes a heavy volume of filled units against inspectors who can review only a fraction. The batch record holds information about product, not people, so this example does not treat it as sensitive.
- baseline
This example follows the inspection pattern the case file documents. It does not rebuild any one manufacturer's system. Machine learning inspects filled injectable drugs for particle contamination and cosmetic defects, in a regulated setting where patient safety is at stake. The two possible errors are not equal, and the choice between them is deliberate. A false accept, a missed defect that reaches a patient, is a patient safety failure. A false reject, a scrapped good vial, is a cost. So the system is tuned to over-reject. The case file calls this an owned governance decision in the safe direction. This example draws it as a pathway from the AI to itself.
- baseline
This example draws the qualification of the AI as a check on the AI's own calls. An AI in a validated quality process must be qualified and monitored for drift, meaning changes in how it behaves over time. The regulator's framework for AI in drug manufacturing is still being developed. So this example treats the qualification as an emerging check, not a settled one. The quality system around the AI is mature, but it has no finished way yet to validate a learning model over time.
- assumed
This example draws a second check: a guard that keeps the human inspectors an independent backstop. The over-reject tuning lowers the visible risk without removing it. A miss is rarer, not impossible, and the inspectors are the layer meant to catch the misses that remain. Over-trust wears that backstop down where it matters most: on the missed defect the design was built to prevent. So choosing the error well does not finish the governance. The qualification and the human backstop are what keep the remaining risk covered. Following the caution in the research behind this example, nothing here claims that a named manufacturer's AI let a defect ship.
- assumed
This example computes no patient safety outcome and no defect that reached a patient. The network draws only the manufacturer's own inspection: its AI, its inspectors, and its records. The patients who receive the products are not drawn in it. The trade-off between the two errors, the over-reject tuning, the developing qualification, and the risk of inspector over-trust come from the case file. Nothing in this diagram computes them.
What this example does not show
Show all 2 limitations
- This example shows no patient safety outcome and no defect that reached a patient. Patients are not part of the network. The network draws only the manufacturer's own inspection: its AI, its inspectors, and its records. The trade-off between the two errors, the over-reject tuning, the developing qualification, and the risk of inspector over-trust come from the case file. Nothing here computes them.
- Following the caution in the research behind this example, nothing here claims that a named manufacturer's AI let a defect ship. This example draws the deliberate over-reject tuning. It draws the AI's qualification and the guard on inspectors' independence as two checks, not as computed harm.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Machine-learning automated visual inspection of filled injectable drug products flags particulate and cosmetic defects that manual inspection or fixed-rule cameras would otherwise judge. In this safety-critical, regulated manufacturing setting the error trade-off is asymmetric and deliberate: a false accept — a missed defect in an injectable that reaches a patient — is a patient-safety failure, while a false reject — scrapping a good vial — is a cost, so the system is tuned to over-reject rather than risk a miss. Because the inspection sits inside a validated pharmaceutical quality process, the AI cannot simply be switched on; it must be qualified within that process, and a regulator is actively developing the framework for how AI in drug manufacturing should be validated and monitored.
empirical- Peer-reviewed Veillon, R., Shabushnig, J., Aabye-Hansen, L., et al. (2023). Applying Machine Learning to the Visual Inspection of Filled Injectable Drug Products. PDA Journal of Pharmaceutical Science and Technology, 77(5), 376-401. https://doi.org/10.5731/pdajpst.2022.012796 https://journal.pda.org/content/77/5/376
- Government U.S. FDA, CDER/OPQ (2023). Discussion Paper: Artificial Intelligence in Drug Manufacturing. Docket FDA-2023-N-0487. https://www.fda.gov/media/165743/download
Two governable surfaces follow from putting AI inside a regulated inspection. First, qualification: an AI in a validated quality process is not simply deployed but must be qualified and monitored for drift, and because the regulator's AI-specific framework is still developing, the qualification of the model's behavior over time is an emerging, not-yet-settled check rather than a solved one. Second, the human backstop: the manual inspector is what catches the false accepts the over-reject tuning is meant to avoid, so if inspectors come to defer to the AI and stop scrutinizing, that backstop erodes exactly where it matters most — the missed defect the asymmetric tuning was designed to prevent. The governable reading is that the over-reject tuning lowers the visible risk without removing it, and the qualification and the human backstop are what keep the residual risk covered.
empirical- Peer-reviewed Veillon, R., Shabushnig, J., Aabye-Hansen, L., et al. (2023). Applying Machine Learning to the Visual Inspection of Filled Injectable Drug Products. PDA Journal of Pharmaceutical Science and Technology, 77(5), 376-401. https://doi.org/10.5731/pdajpst.2022.012796 https://journal.pda.org/content/77/5/376
- Government U.S. FDA, CDER/OPQ (2023). Discussion Paper: Artificial Intelligence in Drug Manufacturing. Docket FDA-2023-N-0487. https://www.fda.gov/media/165743/download
Where this connects
Institutional pressures in this domain
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Industrial QA & operations AI domain page.
Levers available here and the patterns behind them
- Gate vendor updates — Vendor quality gate
- Gate record entries — Human-in-the-loop write gating
- Pause AI on alarms — Deployment circuit-breaker
- Review the riskiest first — Risk-tiered oversight
- Escalate checks — State-feedback vigilance
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Train the staff — AI literacy & boundary rules
- Upgrade model — Improve the model