PAN Lab example
LyssnCrisis counselor QA at ProtoCall Services (988)
The AI watches the counselor rather than the caller: a crisis-line QA scorer
LyssnCrisis rates how counselors on 988, the US crisis line, handle calls, never the callers. No outside study has checked its scores or their benefit.
See more
LyssnCrisis is an AI quality-review tool built by Lyssn.io, a Seattle software company. Lyssn deployed it with ProtoCall Services, a crisis-call contractor in Portland, Oregon. It transcribes call audio and scores the counselor's practice, such as whether they assessed suicide risk, then returns the scores to counselors and supervisors within minutes.
How it is used
988 is the national crisis line in the United States. A ProtoCall counselor takes a 988 or backup crisis call. LyssnCrisis transcribes the call audio. It rates the counselor's conduct against a rating scheme that human reviewers first applied to calls by hand. The scheme covers whether the counselor assessed suicide risk, and which parts of the assessment they covered. It also covers empathy and active listening.
Within minutes, the counselor and their supervisor get fidelity scores, transcripts, and dashboards. A fidelity score says how closely the counselor's conduct matched the scheme. Supervisors and clinicians read the dashboards. The feedback is meant to improve the counselor's next calls.
The tool never interacts with a caller. It makes no decision about a caller and takes no automated action. The counselor keeps full discretion over every call.
The gap it addresses
Crisis centers in the 988 network are required to review 3% of their calls for quality. In June 2023, STAT News reported that ProtoCall took more than 1,000 calls a day. That was about 560,000 in the previous twelve months. ProtoCall was reviewing under 3% of its calls, mostly picked at random. So most of its crisis counseling went unmeasured.
988 guidance, as Lyssn's grant announcement cites it, recommends a suicide risk assessment in every conversation. The announcement also cites research finding that counselors assessed suicide risk in only about half of calls. The announcement dates that research to 2007, before 988 existed. It is the vendor's context, not a measure of ProtoCall.
The tool's aim is to extend measured review from that small sample toward nearly every call. The grant's abstract gave the reason to scale: a projection of up to 40 million 988 calls a year by 2026.
Who is involved
Lyssn builds and updates the scorer and curates the calls people rated by hand to train it. ProtoCall runs the counselor review process and the human supervision the tool supports. It also obtains callers' consent. The platform is described as compliant with HIPAA, a US health-privacy law, with options to remove data.
ProtoCall holds a contract with SAMHSA, a US federal agency, as a national backup provider for 988. It also runs the main 988 line for New Mexico. SAMHSA knew of the work and supported it, but it did not fund or require it.
ProtoCall's chief clinical officer, Brad Pendergraft, doubted that SAMHSA would require AI review tools, given the burden on smaller providers. Beyond the trial described below, oversight of the tool is the ordinary relationship between a vendor and its customer.
How accurate the scores are
The main published evidence is a peer-reviewed study in the journal Psychiatric Services, first published online in July 2024. Its team labeled 476 calls by hand, 193,257 statements in all. They then fine-tuned, or trained further, a transformer, a type of AI language model, on those labels.
The study compared the model's ratings with human raters' ratings. For spotting whether any risk assessment happened, the model reached 98% of the agreement human raters reach with each other.
The study also reported an F1 score, a standard accuracy measure where 1 is perfect. It averaged 0.86 for whole calls and 0.66 for single statements. The single-statement figure is lower mainly because some risk labels are rare.
These figures measure agreement with human raters. They are not error rates for the tool in real use.
Who wrote the evidence
Four of the study's authors own a stake in Lyssn, and three of them are cofounders. A ProtoCall clinician is also an author. No one outside has repeated the study.
The work was funded by a fast-track Small Business Innovation Research grant from the National Institute of Mental Health, project R44MH133517. Its principal investigator, the lead researcher, was David C. Atkins. Its three yearly awards total $2,113,800: $275,373 for fiscal 2023, $998,118 for fiscal 2024, and $840,309 for fiscal 2025. The project's end date was January 31, 2026. Lyssn's own figure of $2.1 million matches that total.
The trial of whether it helps
Whether the scores change how counselors work was put to a trial registered on ClinicalTrials.gov, the US trial registry, as NCT06299384. Lyssn was its sponsor, the party responsible for it, with ProtoCall as a collaborator. It enrolled 81 ProtoCall call-takers and assigned them at random.
The design began with four weeks of baseline measurement. Then, for twelve weeks, AI-based scoring and feedback was compared with supervision as usual. The groups then swapped. Its registered outcomes are the AI's own fidelity scores on empathy, active listening, and seven key risk-assessment questions. They also include a nine-item survey of callers after calls and measures of how the tool was put into use.
The trial ran from April 2, 2024, to October 31, 2025, and is listed as completed. Its primary completion date, when data for its main outcome were in, was September 5, 2025.
As of mid-2026 no results are posted on ClinicalTrials.gov. A search of PubMed, the index of medical research, found no results paper. Sharing of participant-level data is marked unavailable, for proprietary reasons. So every claim that counselors' skills improved remains, for now, the vendor's claim.
Lyssn's own list of research papers still showed the ProtoCall collaboration as active in July 2026, under a "Data Collection" label. That is the vendor's index, not an independent status report.
What people have said
Brad Pendergraft, ProtoCall's chief clinical officer, described the problem the tool addresses: "People can burn out in this work. They can stop doing things that are more emotionally difficult for them."
A psychologist quoted by STAT News, Virna Little, called the technology "a potential gamechanger" for identifying under-performing staff. The sources treat that as a hope voiced in commentary, not a practice anyone has evaluated.
What this case asks
Most other crisis-line networks in the Lab point the AI at the caller. One ranks callers in the queue by severity, and another runs intake through a chatbot. Here the AI scores the counselor instead, and never touches a caller. The case file calls that a real design gain. Lyssn also registered a trial of the tool's effect, with counselors assigned at random.
But moving the AI away from callers moves the governance question rather than ending it. The tool watches the counselor, and no one independent watches the tool. The reliability study was written by people who own the tool. The trial of whether the feedback helps is run by the vendor, measures counselors with the AI's own scores, and has not published.
The harm the case file describes is second-order. A wrong score, for example on a rare risk label, can be saved in the archive the tool learns from. Counselors can learn to produce whatever the dashboard rewards. The harder, more emotionally difficult techniques a burned-out counselor drops first are what a checklist can miss. And managers could judge staff by scores before the evidence shows what the scores mean.
So the question is not whether a caller gets a wrong answer. It is whether a measure no one independent has checked becomes the yardstick counselors are judged by. The case file's answer: a measure used to manage people needs independent checking, on a schedule, by someone who does not own it.
What this network is drawn from
This network follows the pattern the case file describes. It is not a reconstruction of the actual tool. It shows the scorer, the call audio, the crisis counselors, the supervisors and clinicians, and the review and training archive. Callers are outside the network, and it computes no suicide, crisis, or clinical outcome.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Closing a pathway means mistakes stop passing along it. The work along it goes on.
This case has a budget of 9 units. Each tool costs the same at every target level.
Explore (No Targets) sets no targets. Service Targets Only asks for two things. The network must be self-correcting, meaning its mistakes are corrected rather than building on each other. The scorer must also be helping the work. Before any tool is used, the scorer is helping, but the network is at a tipping point, not self-correcting. Its mistakes are not building on each other, but they are not reliably corrected either.
Three tools meet Service Targets Only on their own at their standard setting. Escalate checks and Mark AI-written records cost 2 units each, and Gate record entries costs 3. Assign a challenger does it at its stronger setting, for 3 units. Lingering effects is a Lab setting in which mistakes and reliance on the system stay after their cause is gone. The Lab starts with it on. With it on, Store less data does it at its stronger setting, for 5 units. With it off, Store less data does it at its standard setting, for 3 units, and Peer sharing rules at its stronger setting, for 3. Almost every set of tools within the budget meets this level.
Under Service and Safety Targets and All Governance Targets, the targets can be met within the budget. Both levels ask you to close every failure pathway, among other targets. Seven are open before any tool is used. Three carry the scorer's work: Calls scored for fidelity, Scores sent to counselors, and Dashboards sent to supervisors. Four involve the archive: Scores saved to the archive, Calls saved to the archive, Scored calls used for training, and Supervisors read past scores.
The cheapest way costs 7 units. Escalate checks closes Scores sent to counselors and Dashboards sent to supervisors. Mark AI-written records closes Calls scored for fidelity, Scored calls used for training, and Supervisors read past scores. Gate record entries or Store less data closes Scores saved to the archive and Calls saved to the archive.
Every way to meet either level includes Escalate checks and Mark AI-written records, with Gate record entries or Store less data. Calls scored for fidelity is the one pathway here that the Lab marks as a privacy risk, because the audio holds what callers say. Mark AI-written records is the one tool on offer that closes it. Upgrade model is in no set that meets either level.
One set differs between the two levels. Review on schedule added to the cheapest way with Gate record entries meets Service and Safety Targets. It does not meet All Governance Targets, because the scorer then adds too little to the work.
The link named Independent check of the scores stands for the check the sources say is missing. No tool on offer adds it.
More is not better here. Every tool at its strongest setting costs 36 units, four times the budget. It closes every failure pathway, but the scorer then no longer helps the work, so it meets no level's targets.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Lyssn-class counselor-QA scorer on a crisis line network: 5 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 2 assumed · 4 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 6 assumptions
- assumed
This network follows the counselor review pattern in the LyssnCrisis at ProtoCall Services case file. It is not a reconstruction of the actual tool. Its defining choice is direction: the AI scores the counselor's own practice and never acts on a caller. Most other crisis-line networks in the Lab put the AI on the caller's side instead. One ranks callers in the queue by severity, and another runs intake through a chatbot.
- baseline
The pathways that matter most return the scorer's results to people. One goes to the counselor and is meant to shape later calls. The other goes to the supervisor, for coaching and possibly for judging performance. The archive the scorer trains on matters too. Because the tool measures how staff work, the tools that matter protect the human reading of the scores and the checking of the measure. The model's raw accuracy matters less.
- baseline
The tool addresses a measured gap. 988 centers must review 3% of calls for quality, and in 2023 ProtoCall's manual review covered under 3%, mostly picked at random. 988 guidance, as Lyssn's grant announcement cites it, says a suicide risk assessment should happen in every call. The network treats the manual review as a real check over a small random sample. The tool extends measurement toward every call, rather than replacing a working full review.
- baseline
The network assumes habits pass between counselors, some spreading mistakes and some catching them. The network assumes that working to the score, and dropping harder techniques under burnout, spread among counselors. A habit of flagging odd scores to one another, and the required manual review, act as checks. The evidence that the scores are reliable is real but vendor-authored. Four of its authors own a stake in Lyssn, and three are cofounders. No one outside has repeated it.
- baseline
The network includes an independent check of the scores against human raters. As of mid-2026 the sources describe no such check, so the scores rest on the vendor's own study. Lyssn's registered trial, with 81 call-takers, was built to test whether the feedback helps. It assigned counselors at random and swapped them between AI feedback and usual supervision. Its results are unpublished as of mid-2026, and its participant-level data are proprietary. So every claim that counselors' skills improved remains the vendor's claim. Three things would stand behind the scores: an independent assessment, a standing review schedule, and vendor terms that secure independent test access. The sources describe none of them.
- assumed
Callers are outside this network, and counselors appear as a group, not as people. It computes no suicide, crisis, or clinical outcome. That boundary is required for a behavioral health case. The harm this case can show is second-order and never touches a caller directly. A wrong fidelity score can be saved in the archive the tool learns from. Or managers can judge staff by scores before the evidence of their effect is published. Any such harm is recorded outside a network like this one. The network never scores or flags a person.
What this example does not show
Show all 4 limitations
- This example does not show callers or counselors as people, and it computes no suicide, crisis, or clinical outcome. It shows how mistakes pass between the scorer, the staff, and the archive. Those outcomes are measured, if at all, outside a network like this one.
- The tool scores the counselor's practice. It never interacts with a caller or makes any decision about one. So nothing here shows how callers are triaged or cared for.
- Every claim that counselors' skills improved is the vendor's claim. The reliability study is vendor-authored, with four authors who own a stake in Lyssn, three of them cofounders. No one outside has repeated it. The registered trial of whether the feedback helps, which assigned counselors at random, completed on October 31, 2025. As of mid-2026 it had published no results, and its participant-level data are marked proprietary.
- How this network starts is not a finding about the real deployment, safe or unsafe. The reliability figures measure agreement with human raters, not how often the tool is wrong in real use. Using the scores to identify under-performing staff is a hope voiced in commentary, not a practice anyone has evaluated.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
An AI quality-assurance tool deployed on a national 988 backup line scores crisis counselors' own call practice rather than callers, expanding measured review from the under-3% of calls that had been reviewed by hand toward nearly all of them; a peer-reviewed reliability study of 476 labeled calls reported agreement with human ratings at 98 percent of human interrater agreement for detecting any risk assessment, with average F1 of about 0.86 at call level and 0.66 at statement level, and its authors include four holders of equity in the vendor.
empirical- Academic Imel, Pace, Pendergraft, Pruett, Tanana, Soma, Comtois, Atkins, Machine Learning-Based Evaluation of Suicide Risk Assessment in Crisis Counseling Calls (Psychiatric Services, 2024;75(11):1068-1074) https://pubmed.ncbi.nlm.nih.gov/39026467/
- Investigative Aguilar, A 988 operator faced with a flood of calls turns to AI to boost counselor skills (STAT News, 2023) https://www.statnews.com/2023/06/22/988-suicide-hotline-lyssn-protocall-artificial-intelligence/
- Government NIH RePORTER (National Institutes of Health), Voice-based AI to scale evaluation of crisis counseling in 988 rollout (R44MH133517) (2025) https://reporter.nih.gov/project-details/10983779
The registered randomized crossover trial of the tool's counselor feedback (81 call-takers) completed on October 31, 2025, but as of mid-2026 no results were posted to the trial registry or found in the peer-reviewed literature and participant-level data were marked unavailable for proprietary reasons, so reported counselor-skill-improvement effects remain vendor claims pending independent publication.
empirical- Government ClinicalTrials.gov (U.S. National Library of Medicine), Voice-Based AI to Scale Evaluation of Crisis Counseling in 988 Rollout (NCT06299384) (2026) https://clinicaltrials.gov/study/NCT06299384
- Vendor Lyssn.io, Academic Papers: Deployment and evaluation of Lyssn's risk and safety assessment tool at a national crisis and 988 call center (research index, 2026) https://www.lyssn.io/resources/academic-papers/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Behavioral-health & crisis triage domain page.
Levers available here and the patterns behind them
- Escalate checks — State-feedback vigilance
- Keep skills sharp — Deskilling-arrest mandate
- Mark AI-written records — Provenance labeling
- Peer sharing rules — Peer-edge governance
- Review on schedule — Oversight cadence & retrospectives
- Assign a challenger — Structured dissent
- Gate vendor updates — Vendor quality gate
- Gate record entries — Human-in-the-loop write gating
- Store less data — Data minimization
- Upgrade model — Improve the model
Documented case histories
- LyssnCrisis counselor QA at ProtoCall Services (988)
- REACH VET
- Vanderbilt VSAIL suicide-risk alert
- Kaiser Permanente Suicide-Risk Model
- Crisis Text Line & Loris.ai
- NarxCare
- Stratification Tool for Opioid Risk Mitigation
- ODMAP overdose spike alerts on a drug-enforcement-housed store
- The discontinuation that wasn't: a school communication scanner swapped rather than stopped
- Oxevision camera monitoring on NHS mental health wards
- Limbic Access (NHS Talking Therapies)
- Two surfaces, one program: NYC's teen teletherapy, its suicide-alert algorithm, and the ad trackers on the sign-up page
- Four retrofits and a shutdown: a companion platform's crisis screen under external pressure
- Tessa chatbot replacing the NEDA eating-disorder helpline
- Woebot (a governed app wind-down)