PAN Lab example
NarxCare
The score you cannot see: an opaque prescribing-risk model
Bamboo Health's NarxCare turns state prescription records into patient risk scores. Patients cannot see or contest the scores, which can gate access to pain medication.
See more
NarxCare is proprietary software from Bamboo Health, formerly Appriss Health. It reads a patient's records in a state Prescription Drug Monitoring Program (PDMP), the state-run records of dispensed controlled-substance prescriptions. It returns three Narx Scores and an Overdose Risk Score, each from 000 to 999, to prescribers and pharmacists. Its algorithm and the inputs it derives are proprietary and undisclosed.
How the scores are built
Bamboo Health's own documentation, dated September 2023, describes the scores as follows. Each Narx Score is a three-digit number for one type of controlled substance, and a higher number means more exposure. The first two digits place the patient's exposure as a percentile against everyone else in that state's PDMP. The last digit counts active prescriptions.
Morphine milligram equivalents (MME), a standard measure of opioid dose, make up half of each Narx Score's weight. The score averages four overlapping time windows, the shortest two months and the longest a year.
The Overdose Risk Score is a logistic regression, a model that adds up weighted factors, predicting unintentional overdose death. A higher score means higher predicted risk. It uses nine ranked inputs from the PDMP. Total MME over a year weighs most, at roughly a quarter to a third.
It was first trained on more than 5,000 overdose deaths confirmed by autopsy, matched to 500,000 comparison patients. They came from a single Midwestern state, from 2013 to 2016.
Where it is used
The scores appear in the PDMP web portal, the electronic health record, or pharmacy management software. They are often placed in the patient header beside vital signs and allergies.
Counts of its reach differ by what is counted. KFF Health News and the Associated Press reported that more than 40 states and territories run their PDMPs on Bamboo Health technology. They also reported that five of the top six pharmacy chains use NarxCare.
The 2026 npj Digital Medicine study says the scoring module is switched on in statewide PDMPs in more than 20 states. A clinician perspective in the Journal of General Internal Medicine put the platform in 45-plus states. By a reach estimate its authors flag as such, it could influence over a billion patient encounters a year.
State PDMP administrators decide whether the scoring module is switched on, and which tiles, indicators, and thresholds are shown.
Advisory in name
The vendor's documentation says more than a dozen times that the scores are “intended to aid, not replace” medical decisions. It says they should “never” be the sole reason to give or refuse medication.
In practice, clinician and patient-advocacy sources document the score deciding outcomes. It is often displayed in the patient header, and its basis is invisible. Clinicians who fear Drug Enforcement Administration (DEA) scrutiny and criminal liability treat a high number as a stop sign.
The documented results are forced tapers, where a dose is cut down against the patient's wishes, patients dropped from care, and pharmacy refusals. These can themselves raise overdose risk. No public statistic says how often clinicians follow a high score or decide against it.
Documented harms
A Michigan graduate student's two dogs had controlled-substance prescriptions from the veterinarian, filled under her name. They inflated her profile and contributed to her being denied emergency pain care. Another patient was told before an MRI, “Your Narx Score is so high, I can't give you any narcotics.”
Patients cannot see, challenge, or correct their scores. NarxCare does not track the outcomes of tapering or stopping a medication. So harm to a denied patient, including a push toward the illicit market, never shows up in the records the model reads. It is never used to correct the model.
The records can also mislead in the other direction. Medications that treat opioid use disorder carry high MME, so being in recovery treatment can raise a patient's own Overdose Risk Score.
The accuracy dispute
Precision is the share of patients a score flags who truly have the outcome. On its own 2013 to 2016 data, Bamboo Health reported an Overdose Risk Score precision of about 75%. It reported recall, the share of real cases the model found, of 57%. It reported specificity, the share of non-cases correctly left unflagged, of 81%.
It also reported odds of unintentional overdose death rising with the score. Against scores of 000 to 199, the odds were 12.4 times higher at 500 to 599, and 29.3 times at 800 to 999. These are vendor self-reports, never independently reproduced.
The vendor's own external check, on a different state's data from 2017 to 2023, showed precision falling to about 52%. Bamboo attributes the fall to illicit fentanyl, which PDMPs do not track, and to wider use of treatment medication.
The independent rebuild
In a 2026 study in npj Digital Medicine, researchers rebuilt the Overdose Risk Score. They used California's CURES prescription-monitoring database, with about 17.9 million records to learn from, and commercial insurance claims data.
They could not reproduce the vendor's precision. They obtained 1% to 32% instead. They tried four kinds of statistical and machine-learning model: logistic regression, random forest, gradient boosting, and neural networks.
This is not a strict like-for-like refutation. Overdose-death records were not available to the researchers. So they trained on stand-in outcomes: starting treatment for opioid use disorder in one dataset, and opioid-related harms in the other. The result is best read as evidence that the model's secrecy keeps anyone outside the vendor from assessing its accuracy, fairness, or safety.
What other research says
A 2022 study in JAAOS Global Research and Reviews followed 1,136 knee-replacement patients. A higher Narx Narcotic Score, the Narx Score for narcotics, on admission independently predicted readmission within 30 days. The odds were 2.46 times higher for scores of 300 to 499, and 3.98 times for 500 and above, against patients new to opioids.
In the California Law Review, Jennifer Oliva argues that PDMP risk platforms were designed for law-enforcement surveillance and never validated for clinical care. She argues they likely inflate scores for women and for Black, poor, uninsured, and rural patients. The argued route is proxies such as cash payment and distance traveled. This is a documented but contested concern, not a measured rate per group.
Economist Angela Kilby rebuilt a comparable opioid-risk model. She found that predicted risk does not track whether opioids would help or harm the individual patient.
The npj paper, the Oliva article, and the clinician perspective make one shared point. NarxCare has been deployed at national scale but never independently validated for clinical care.
Who regulates it
The Food and Drug Administration (FDA) treats NarxCare as clinical decision support, software that advises clinicians, outside its premarket review of software as a medical device. The npj authors note the FDA's own guidance. It says software that gives a risk score for a disease would be a regulated device function.
Contesting the score has run through the FDA. A 2023 citizen petition, a formal public request to the agency, asked it to declare NarxCare a misbranded device and order a recall. A misbranded device is one sold with false or misleading labeling. The FDA rejected it on procedural grounds. A second citizen petition, docket FDA-2025-P-0701, has been pending since 2025 and has drawn more than 1,000 public comments.
The central dilemma
In most cases in this Lab, a correction route exists but fails. Here, in the case file's reading, no route exists for anyone. Three doors are shut at once.
The patient cannot see or contest the score, though they have the most reason to challenge an error. No outsider can audit the deployed model, because it is proprietary. No outside authority has validated it, because the FDA does not review it as a medical device.
With all three shut, the state that deploys NarxCare configures a box it cannot open. No court ruling yet binds it. The only actors left are outside the system: researchers who rebuild it and citizens who petition.
The case file's lesson is that the ability to contest a score comes before every other control. Improving the score's accuracy changes least, because none of the shut doors is an accuracy problem.
What this network is drawn from
This network follows the pattern the case file describes. It is not a reconstruction of the actual tool. It shows NarxCare, prescribers, dispensing pharmacists, the PDMP records, and the FDA with the petitions contesting it.
It never models overdose or any clinical outcome. The patients the score sorts are not in the network.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Closing a pathway means mistakes stop passing along it. The work along it goes on.
This case has a budget of 9 units and offers thirteen tools. Each costs the same at every target level.
Before any tool is used, mistakes are copied about as fast as they are corrected. The Lab calls this a tipping point. The work is strained, falling behind demand, and eight failure pathways are open.
Explore (No Targets) sets no targets. Service Targets Only asks for mistakes to be caught faster than they are copied, which the Lab calls self-correcting. It also asks for NarxCare to be helping the work.
Lingering effects is a Dynamics setting in which damage outlasts its cause, and the Lab starts with it on. With it on, no single tool meets that level's targets. Ten pairs of tools meet them for 4 units. One is Escalate checks, which raises checking when monitoring flags trouble, with Review on schedule, which reviews the deployment on a fixed rhythm. In all, more than 700 different sets of tools within the budget meet them.
With lingering effects off, Peer sharing rules at its stronger setting meets them alone, for 3 units. It sets rules for what people and systems pass to each other.
Service and Safety Targets and All Governance Targets ask you to close every failure pathway, among other targets. Both can be met, but only by spending the whole budget. Exactly two sets of tools do it, each for all 9 units.
The eight open pathways are Scores shown to prescribers, Scores shown to pharmacists, Records scored by NarxCare, and Prescribers read history. They are also Prescriptions become records, Fills become records, One score at both decisions, and Same model across states.
Escalate checks closes Scores shown to prescribers and Scores shown to pharmacists.
Mark AI-written records closes Records scored by NarxCare and Prescribers read history. In the Lab it marks machine-written content in the records so readers can weigh it. Here that content is NarxCare's scores, which this network assumes are logged with the dispensation records.
Peer sharing rules closes One score at both decisions and Same model across states. Understand the system, the tool that pays for ongoing study of the deployment, is not offered here. So with lingering effects on, Peer sharing rules works at reduced strength, but it still closes both.
Gate record entries, which requires sign-off before anything enters the records, closes Prescriptions become records and Fills become records. Store less data closes those same two instead.
So the two sets are Escalate checks, Mark AI-written records, and Peer sharing rules, with either Gate record entries or Store less data.
The case file argues for four controls around the model. In the Lab they are Gate vendor updates, Require sign-off, Review on schedule, and Escalate checks. The first three are in neither set, and none of them closes a pathway. Escalate checks is in both.
The sources describe neither of the network's two checks in the deployment. Check with a second model starts Audit of the model. Assign a challenger, which makes challenge a scheduled duty, and Peer sharing rules start Recourse to contest a score.
Using every tool on offer, each at its strongest setting, costs 48 units, more than five times the budget. It closes every failure pathway. But the added checks cut the benefit NarxCare adds below what every level with targets asks.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the NarxCare-class opaque prescribing-risk score network: 5 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 6 assumptions
- assumed
This network models the pattern of an opaque prescribing-risk score documented in the NarxCare case. It is not a reconstruction of the actual tool or its algorithm.
- baseline
The network draws two checks the sources describe as missing from the deployment. No one outside Bamboo Health can audit or validate the deployed model, because it is proprietary and undisclosed. The FDA has not regulated it, and patients cannot see or contest their scores. So no one in the deployment has a way to review, re-approve, or challenge a score.
- assumed
The vendor calls the scores advisory, but clinician and patient-advocacy sources document them deciding outcomes in practice. The network assumes a high score usually carries the decision, and prescribers seldom decide against it. The sources give three reasons. One is automation bias, the habit of trusting an automated output over one's own judgment. The others are the score's prominent placement and fear of regulatory and criminal liability. No public figure measures how often prescribers follow or overrule a score, so how strongly the network draws both is a modeling choice.
- baseline
In this network, as in the deployment, NarxCare's scores are computed from the patient's own prescription records. Medications that treat opioid use disorder carry high morphine milligram equivalents, a standard measure of opioid dose. So being in recovery treatment can itself raise a patient's Overdose Risk Score, a documented perverse effect.
- baseline
The network assumes the prescriber and the pharmacist rely on the same single opaque score. So its blind spots affect both decisions together, rather than at random. A high score behind one decision carries over to the other.
- assumed
This network models no overdose, suicide, or other clinical outcome. It traces how mistakes pass inside the institution, and the patients the score sorts are not in it. The discrimination concern documented for such scores is argued and contested. The case file records it, and nothing in this network computes it.
What this example does not show
Show all 3 limitations
- This example traces how mistakes pass inside the institution. It never models overdose, suicide, or any clinical outcome, and the patients the score sorts are not in it. A score or a denial here is an institutional signal, never a person's care or harm. The case file records the documented patient harms, the denials, and the way a denial can push a patient toward the illicit market.
- The accuracy figures are read with care. The vendor's precision of about 75% is self-reported, on its own data. The independent figure of 1% to 32% comes from a 2026 rebuild of the model. Lacking overdose-death records, those researchers trained on stand-in outcomes. So their result shows that the model's secrecy prevents outside assessment, not a like-for-like refutation of the vendor's number. The case file carries both as what they are.
- The documented discrimination concern is argued and contested, not measured. It holds that scores like this are likely inflated for women and for Black, poor, uninsured, and rural patients. The argued route is proxies such as cash payment and distance traveled. No rate per group is measured, and no disparity figure is given here. The case file carries the concern in words, as an outside observation, and nothing in this network computes it.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
NarxCare is a proprietary clinical-decision-support platform built by Bamboo Health that layers over state Prescription Drug Monitoring Programs and returns three Narx Scores plus a composite Overdose Risk Score (each 000-999) into the electronic health record, the PDMP portal, or pharmacy software, often in the patient header alongside vitals and allergies; adoption figures vary by what is counted (more than 40 states and territories run their PDMPs on Bamboo technology and five of the top six pharmacy chains use NarxCare, while the scoring module itself is switched on in more than 20 states). The vendor states the scores are intended to aid, not replace, clinical judgment and should never be sole justification for providing or refusing medication, but clinician and patient-advocacy sources document de facto determinative use — denials, forced tapers, and pharmacy refusals — driven by automation bias and fear of regulatory and criminal liability; patients cannot see, challenge, or correct their scores, the algorithm is proprietary and has not been independently validated for clinical care, and the FDA has not regulated it as a Software-as-a-Medical-Device, so contestation has instead run through FDA citizen petitions (one rejected on procedural grounds in 2023 and a second, docket FDA-2025-P-0701, pending since 2025 with more than 1,000 public comments).
empirical- Vendor Bamboo Health, Inc., NarxCare Application Overview (Version 1.0, September 2023; hosted by the Idaho Division of Occupational and Professional Licenses) https://dopl.idaho.gov/wp-content/uploads/2024/07/2023.10.04.Bamboo-Health-NarxCare-Application-Overview.pdf
- Academic Wang, Stofer, Chu, Huang, Li, Algorithmic opacity in opioid risk scoring and the need for transparent AI regulation (npj Digital Medicine, 2026; DOI 10.1038/s41746-026-02491-y) https://www.nature.com/articles/s41746-026-02491-y
- Investigative Miller and Whitehead, Artificial Intelligence May Influence Whether You Can Get Pain Medication (KFF Health News, 2023) https://kffhealthnews.org/news/artificial-intelligence-pain-medication-narx-score/
- Academic Buonora, Axson, Cohen, Becker, Paths Forward for Clinicians Amidst the Rise of Unregulated Clinical Decision Support Software: Our Perspective on NarxCare (Journal of General Internal Medicine, 2023) https://pmc.ncbi.nlm.nih.gov/articles/PMC11043299/
- Academic Oliva, Dosing Discrimination: Regulating PDMP Risk Scores (California Law Review, 2022; Vol. 110) https://www.californialawreview.org/print/dosing-discrimination-regulating-pdmp-risk-scores
- Investigative Pain News Network, Petition Asks FDA to Take NarxCare Off the Market (2023) https://www.painnewsnetwork.org/stories/2023/4/28/citizens-petition-calls-on-fda-to-take-narxcare-off-the-market-nbsp
- Trade press Medscape, Hidden Formulas, High Stakes: The Fight to Regulate Clinical Decision Support Tools (2025) https://www.medscape.com/viewarticle/hidden-formulas-high-stakes-fight-regulate-clinical-decision-2025a1000cw3
- Trade press Medscape, When an Algorithm Guides Pain Management: The Growing Backlash Against NarxCare Scores (2025) https://www.medscape.com/viewarticle/when-algorithm-guides-pain-management-growing-backlash-2025a100091n
A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.
empirical- Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
- Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Behavioral-health & crisis triage domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Peer sharing rules — Peer-edge governance
- Gate record entries — Human-in-the-loop write gating
- Gate vendor updates — Vendor quality gate
- Require sign-off — Conformity assessment gate
- Check with a second model — Cross-model verification
- Assign a challenger — Structured dissent
- Escalate checks — State-feedback vigilance
- Keep skills sharp — Deskilling-arrest mandate
- Review on schedule — Oversight cadence & retrospectives
- Upgrade model — Improve the model
- Review the riskiest first — Risk-tiered oversight
- Store less data — Data minimization
Documented case histories
- NarxCare
- REACH VET
- Vanderbilt VSAIL suicide-risk alert
- Kaiser Permanente Suicide-Risk Model
- Crisis Text Line & Loris.ai
- LyssnCrisis counselor QA at ProtoCall Services (988)
- Stratification Tool for Opioid Risk Mitigation
- ODMAP overdose spike alerts on a drug-enforcement-housed store
- The discontinuation that wasn't: a school communication scanner swapped rather than stopped
- Oxevision camera monitoring on NHS mental health wards
- Limbic Access (NHS Talking Therapies)
- Two surfaces, one program: NYC's teen teletherapy, its suicide-alert algorithm, and the ad trackers on the sign-up page
- Four retrofits and a shutdown: a companion platform's crisis screen under external pressure
- Tessa chatbot replacing the NEDA eating-disorder helpline
- Woebot (a governed app wind-down)