PAN Lab example
Woebot (a governed app wind-down)
The responsible wind-down: retiring a peer-reviewed CBT chatbot
In 2025 Woebot Health retired Woebot, its therapy chatbot app, on a published schedule. Users could request their chat transcripts before account data was anonymized.
See more
Woebot was a smartphone chatbot app from Woebot Health that delivered self-help cognitive behavioral therapy (CBT). Its replies followed conversation paths its authors wrote in advance, so it could not invent a new clinical claim. It stored each user's conversations as transcripts, sensitive records that had to be dealt with when it shut down.
How it worked
Woebot Health, formerly Woebot Labs, of San Francisco, was founded by Alison Darcy. The app ran from 2017 to 2025. Woebot Health offered the app directly to the public and through employers and health systems.
The conversation felt natural, but each reply followed a decision path written in advance. Roughly 1.5 million people used Woebot over its lifetime. That is a cumulative figure reported in press coverage, not an audited count of active users at one time.
What the evidence showed
The foundational study in its peer-reviewed record was a 2017 randomized controlled trial in the journal JMIR Mental Health. In such a trial, people are assigned by chance to the app or to a comparison. Its authors were Fitzpatrick, Darcy, and Vierhile, Woebot's own people. It enrolled 70 people aged 18 to 28 for two weeks. The comparison group received an information-only self-help ebook.
The Woebot group's depression symptoms fell significantly on the PHQ-9 questionnaire. The authors called it a moderate difference between the groups, an effect size of 0.44. They concluded that a chatbot can be a feasible and engaging way to deliver CBT.
That is an efficacy signal, an early sign the app might help, not regulatory validation. The trial was small and short. It was unblinded, so participants knew which group they were in. Woebot's wider published evidence includes further studies not gathered here. What the sources lack is an independent, arm's-length evaluation of the consumer app.
The prescription version, WB001
WB001 was a separate, investigational version, available only on prescription. It was an eight-week smartphone treatment for postpartum depression, to be used under a clinician's supervision. It combined CBT with elements of interpersonal psychotherapy.
In May 2021 it received a Food and Drug Administration (FDA) Breakthrough Device Designation. That status speeds up review. It is not marketing authorization, the FDA's permission to sell a device. WB001 entered a pivotal trial as medical-device software, with the first patient enrolled in January 2023. A pivotal trial is the main study meant to support an FDA decision. This one ran at several sites and was double-blind, so neither patients nor researchers knew who got which version. It stayed investigational and never received marketing authorization. The consumer app that shut down is a different product.
How it was funded
In July 2021 Woebot Health closed a $90 million funding round, called a Series B, co-led by JAZZ Venture Partners and Temasek. That brought its total funding to $114 million, per the company's press release. A secondary industry summary calls the round a Series C and puts the total at about $107.5 million. The difference is noted here, not resolved.
The funding is context for the founder's later argument. She said the cost of authorization, not a lack of investment, ended the product.
How it was wound down
In April 2025 Woebot Health announced it would shut the app down. It did so on a published schedule, set out in its own FAQ. The app was retired on June 30, 2025. Users could request a transcript of their conversations until July 15, 2025. The FAQ said all account data would be anonymized as of July 31, 2025.
Anonymizing removed personally identifying information. The service was not quietly abandoned, and identifiable data was not kept indefinitely. After the retirement, no new accounts could be created, and earlier accounts could no longer be opened.
What each side says
The founder and chief executive, Alison Darcy, attributed the shutdown largely to the cost of meeting the FDA's requirements for marketing authorization. She also pointed to a mismatch between a fast-moving AI field and regulatory frameworks that lacked a clear pathway. She framed the exit as economic and regulatory, not a clinical failure. That is her own on-record account, not an independently audited finding.
Commentators read the exit as a signal about how AI in mental health is regulated. A clinically studied, transparently governed tool that did not generate text was shut down. Meanwhile, unregulated general-purpose chatbots drew far larger audiences without clinical validation, because they are not marketed as health tools.
What to watch
The danger here is not a runaway mistake by the chatbot. It is what a running service does with the sensitive records it built up, once the business reason to keep going runs out. It is also whether an early efficacy signal gets quietly read as validation on the way out.
Watch the transcript store, and watch what the exit claims about the evidence. This network follows the pattern the case file describes. It does not reconstruct the app, and the people who used it are not in it.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Here a mistake is, for example, a flaw in a pre-written reply. Closing a pathway means mistakes stop passing along it. The work along it goes on.
This case has a budget of 10 units. Each tool costs the same at every target level.
Explore (No Targets) sets no targets. Service Targets Only asks for two things. The network must be self-correcting, meaning its mistakes are corrected rather than building on each other. The chatbot must also be helping the work. Before any tool is used, the chatbot is helping. But the network is at a tipping point, on the edge between correcting its mistakes and letting them build up.
Two tools meet Service Targets Only on their own, for 2 units each. Escalate checks raises checking when monitoring flags trouble. Mark AI-written records labels what the chatbot wrote in the transcripts. Store less data does it at its stronger setting, for 5 units.
Lingering effects is a Lab setting in which mistakes and reliance on the system stay after their cause is gone. The Lab starts with it on. While it is on, Upgrade model also does it at its stronger setting, for 5 units. With lingering effects off, Store less data does it at its standard setting, for 3 units, and so does Vet connections, for 3. Upgrade model then does not. More than 150 different sets of tools within the budget meet this level.
Under Service and Safety Targets, the targets can be met within the budget. That level asks you to close every failure pathway, among other targets, and it runs with lingering effects off. Three are open before any tool is used. They are the pathways named Users' answers shape the script, Conversations saved per user, and Staff manage and anonymize transcripts.
The cheapest way costs 7 units. Store less data closes the two transcript pathways. Mark AI-written records closes the pathway named Users' answers shape the script, and Vet connections can close it instead. This level also asks the staff to keep up with their work. The tools that close those three pathways leave the staff behind their work. Adding Escalate checks or Review the riskiest first keeps them up.
That level can be met with 14 different sets of tools, and every one includes Store less data. Review on schedule is in none of them.
Under All Governance Targets, the targets are not fully addressable with the available tools. That level keeps lingering effects on. There, closing all three pathways takes Store less data at its stronger setting and Mark AI-written records, for 7 units. That leaves the staff falling behind. Adding Escalate checks or Review the riskiest first keeps them up. But the benefit the chatbot delivers then falls below what this level asks. Every combination within the budget was checked, and none meets every target.
This case does not offer Understand the system, the tool that lets you aim other tools. Without it, Store less data and Vet connections work at reduced strength while lingering effects are on.
More is not better here. Every tool at its strongest setting costs 34 units, more than three times the budget. It closes every failure pathway, but the chatbot then no longer helps the work, so it meets no level's targets.
No tool offered here adds the Arm's-length evaluation, the independent check of the app's evidence.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Woebot-class governed wind-down of a scripted CBT chatbot network: 6 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 7 assumptions
- assumed
This network follows the pattern the Woebot case file describes. It does not reconstruct the actual app or its scripted content. It shows the retired app, offered to the public and through employers and health systems. The app was a cognitive behavioral therapy (CBT) chatbot that did not generate text. It does not show WB001, the separate investigational prescription version, and the two are never treated as one. Only the staff and clinical leadership nodes also stand for the people who worked on WB001. Two limits are the network's own choices. It assumes the chatbot uses a user's earlier sessions, which the sources do not describe. They describe storing transcripts, users requesting them, and scheduled anonymization. It also leaves out the employers and health systems that offered the app. A governed wind-down would have to notify them and end those arrangements. The network shows Woebot Health's own staff and leadership.
- assumed
The network assumes a deployment built to be safe. The chatbot followed pre-written paths instead of generating text, so the mistakes it could make in one conversation were limited by design. Staff monitored it as an organization, rather than one stretched worker acting on its output. People chose to use the app, and their answers were given for this purpose. Even so, at the start the network is at a tipping point, not self-correcting. That means it is on the edge between correcting its mistakes and letting them build up. The question becomes what a responsible exit does with the transcripts the app collected.
- baseline
What defines this network is the teardown of the transcript store. It is a large store of each user's conversations, among the most sensitive records a service like this can hold. A governed exit keeps it from leaking or being quietly abandoned. Woebot Health set a window to request transcripts, then scheduled anonymization that removed personally identifying information. The network draws a pathway where transcripts could be copied outside. A careless wind-down would open it, and a governed one keeps it shut. Store less data acts on it by keeping fewer records for less time. Vet connections acts on it by closing links out of the network that no one approved. Here, with lingering effects on, it works at reduced strength. The case file calls deleting records without reading them the pattern to avoid. A responsible teardown keeps less and removes identifying details, rather than deleting unread.
- baseline
The network also draws a check the sources do not document: an independent, arm's-length evaluation of what the app's evidence showed. The foundational peer-reviewed study was written by Woebot's own people. It was small, with 70 people, and short, at two weeks. It was unblinded, so participants knew whether they were using Woebot. It is an early efficacy signal, a first sign the app might help, not regulatory validation. Woebot's wider published evidence is not gathered here, and what is missing is an independent evaluation, not more studies. The consumer app was never a device authorized by the Food and Drug Administration (FDA). WB001, the investigational prescription version, never received FDA marketing authorization. No tool offered in this case adds this check.
- assumed
The real shutdown had an outside cause. The founder, Alison Darcy, attributed it to the cost of meeting the Food and Drug Administration's (FDA's) requirements for marketing authorization. She also pointed to regulatory frameworks that lacked a clear pathway. She framed the exit as economic and regulatory, not a clinical failure. That is her own on-record account, not an independently audited finding. So the danger in this network is not a runaway mistake by the chatbot. It is what happens to the stored transcripts when a running service is shut down. It is also whether the exit follows a published schedule or lapses under the pressure of closing.
- assumed
The network gives Woebot Health's clinical leadership a review pathway of its own, so the leadership takes part in how mistakes are passed on and caught. It stands for the internal clinical leadership above the consumer app. It also stands for the company's work with the Food and Drug Administration (FDA) on WB001, the separate investigational prescription version.
- assumed
The app's users are not in this network. Roughly 1.5 million people used it over its lifetime, a cumulative figure. The network computes no clinical, symptom, or crisis outcome. It shows the organization only, and a conversation or a transcript here stands for work inside the organization, never a person. The staff shown are Woebot Health's own clinical and operations staff. For WB001, the prescription version, they include the clinicians supervising its use.
What this example does not show
Show all 4 limitations
- This example shows the organization only. It never models symptoms, recovery, crisis, suicide, or any other clinical outcome. The people who used the app are not in the network. A conversation or a transcript here stands for work inside the organization, never a person. The app's clinical value and its effects on the people it served are documented in the case file and measured outside this example.
- The evidence is hedged as the sources hedge it. The foundational study was a 2017 randomized controlled trial, in which people were assigned by chance to the app or a comparison. Woebot's own people wrote it. It enrolled 70 people aged 18 to 28 for two weeks. It was unblinded, so participants knew which group they were in. The comparison group received an information-only self-help ebook. It reported a moderate difference in depression symptoms between the groups, an effect size of about 0.44. That is an early efficacy signal, a first sign the app might help. It is not regulatory validation or proof that the app works in general use. This example carries it that way throughout. Woebot's wider published evidence is not gathered here. What is missing is an independent, arm's-length evaluation, not more studies by the company.
- The two products are kept apart. This example shows the retired app, offered to the public and through employers and health systems, which did not generate text. WB001 was a separate, investigational, prescription-only version. It received a Food and Drug Administration (FDA) Breakthrough Device Designation, a status that speeds up review but is not marketing authorization. It entered a pivotal medical-device trial but never received FDA marketing authorization. The two are never treated as one.
- The cause of the shutdown is given as reported, not as an audited fact. The founder attributed it to the cost of Food and Drug Administration (FDA) marketing authorization and to regulatory frameworks that lacked a clear pathway. She framed the exit as economic and regulatory rather than a clinical failure. That is her own account. The lifetime figure of roughly 1.5 million users is a cumulative number from press coverage. It is not an audited count of active users at one time. Funding figures differ across sources. The company's press releases give a $90 million funding round, called a Series B, bringing total funding to $114 million. A secondary summary calls that round a Series C and puts the total at about $107.5 million.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Woebot, a rule-based (non-generative) cognitive behavioral therapy chatbot used by roughly 1.5 million people over its lifetime, was deliberately retired by its maker on a pre-announced schedule: the app was taken down on June 30, 2025, with a transcript-request window (deadline July 15, 2025) and all account data anonymized as of July 31, 2025, removing personally identifying information rather than silently abandoning the service. The founder and chief executive attributed the shutdown to the cost of meeting FDA marketing-authorization requirements and to a regulatory-pathway gap, framing the exit as economic and regulatory rather than a clinical failure - a self-reported account, not an independently audited finding. The roughly 1.5 million figure is a cumulative lifetime number reported in press coverage, not an audited point-in-time active-user count.
empirical- Investigative Aguilar, Why Woebot, a pioneering therapy chatbot, shut down (STAT News, 2025) https://www.statnews.com/2025/07/02/woebot-therapy-chatbot-shuts-down-founder-says-ai-moving-faster-than-regulators/
- Vendor Woebot Health, FAQs (Woebot app retirement) (2025) https://woebothealth.com/faq/
- Trade press HLTH, Woebot Health Is Shutting Down Its App (2025) https://hlth.com/insights/news/woebot-health-is-shutting-down-its-app-2025-04-28
The foundational study in Woebot's peer-reviewed efficacy record is an early-stage, vendor-authored trial: a 2017 randomized controlled trial in JMIR Mental Health (n=70, ages 18 to 28, two weeks, unblinded, information-only control) reported a moderate between-groups reduction in PHQ-9 depression symptoms (about d = 0.44). That is an efficacy signal, not regulatory validation; the study authors were affiliated with the tool's maker, and no independent, arm's-length evaluation of the consumer app is documented (the broader published evidence base is not assembled here). A separate, investigational, prescription-only variant (WB001) received an FDA Breakthrough Device Designation in May 2021 - an expedited-review status, not marketing authorization - and entered a pivotal Software as a Medical Device trial with the first patient enrolled in January 2023, but never received FDA marketing authorization; it must not be conflated with the consumer app.
empirical- Academic Fitzpatrick, Darcy, Vierhile, Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial (JMIR Mental Health, 2017;4(2):e19) https://mental.jmir.org/2017/2/e19/
- Vendor Woebot Health (Business Wire), Woebot Health Receives FDA Breakthrough Device Designation for Postpartum Depression Treatment (2021) https://www.businesswire.com/news/home/20210526005054/en/Woebot-Health-Receives-FDA-Breakthrough-Device-Designation-for-Postpartum-Depression-Treatment
- Vendor Woebot Health (Business Wire), Woebot Health Enrolls First Patient in Pivotal Clinical Trial of WB001 for Postpartum Depression (2023) https://www.businesswire.com/news/home/20230123005211/en/Woebot-Health-Enrolls-First-Patient-in-Pivotal-Clinical-Trial-of-WB001-for-Postpartum-Depression
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Behavioral-health & crisis triage domain page.
Levers available here and the patterns behind them
- Review the riskiest first — Risk-tiered oversight
- Escalate checks — State-feedback vigilance
- Store less data — Data minimization
- Review on schedule — Oversight cadence & retrospectives
- Vet connections — Connection authorization
- Mark AI-written records — Provenance labeling
- Require sign-off — Conformity assessment gate
- Gate vendor updates — Vendor quality gate
- Upgrade model — Improve the model
Documented case histories
- Woebot (a governed app wind-down)
- REACH VET
- Vanderbilt VSAIL suicide-risk alert
- Kaiser Permanente Suicide-Risk Model
- Crisis Text Line & Loris.ai
- LyssnCrisis counselor QA at ProtoCall Services (988)
- NarxCare
- Stratification Tool for Opioid Risk Mitigation
- ODMAP overdose spike alerts on a drug-enforcement-housed store
- The discontinuation that wasn't: a school communication scanner swapped rather than stopped
- Oxevision camera monitoring on NHS mental health wards
- Limbic Access (NHS Talking Therapies)
- Two surfaces, one program: NYC's teen teletherapy, its suicide-alert algorithm, and the ad trackers on the sign-up page
- Four retrofits and a shutdown: a companion platform's crisis screen under external pressure
- Tessa chatbot replacing the NEDA eating-disorder helpline