PAN Lab example
IRS collection chatbots
Expanded without a ruler: a federal collection chatbot with no performance measures
The Internal Revenue Service (IRS) expanded its collection chatbot and live chat, and made live chat permanent, with no performance measures or reliable statistics.
See more
The Internal Revenue Service (IRS) runs a scripted, non-generative chatbot for its Automated Collection System (ACS). ACS is the unit handling balances the IRS says are due and returns it says were not filed. The chatbot can hand taxpayers to a live chat with an assistor, an IRS employee, on a vendor's platform that also produces the program's statistics.
What was built, and when
ACS, part of the IRS Small Business/Self-Employed Division, was the first IRS unit to pilot chat. Its goal was to steer taxpayers to online self-help instead of the toll-free phone line.
The IRS piloted live chat without identity checks in November 2017. In June 2019 it added authenticated chat, where assistors can act on a verified taxpayer's account. The chatbot followed in December 2021, in English and Spanish.
Voice bots on the phone line followed in January and June 2022. Live chat was made permanent in June 2025.
The IRS spent about $7.2 million of Inflation Reduction Act funds on the chat programs in fiscal years 2024 and 2025. Congress cut that act's IRS funding from about $79.4 billion to about $26 billion.
What the audit found
The Treasury Inspector General for Tax Administration (TIGTA), a Treasury oversight office independent of the IRS, issued its audit on June 15, 2026. It is report 2026-308-029. It covered chat reports from February 2023 to December 2024.
The IRS had not analyzed its chat programs and had no performance measures for them, TIGTA found. The Taxpayer First Act requires metrics and benchmarks. The IRS expanded the program and made live chat permanent anyway.
IRS officials had claimed, in an internal 2023 review, that the chatbots cut phone demand. With no analysis of the program, TIGTA found, the IRS could not support that claim.
This outside audit is what brought the broken figures to light. The case file puts the danger this way: not a wrong answer, but an unmeasured one.
How the statistics went wrong
One report showed an assistor working 603 live chats at once, though a system control allowed three. Over 16,000 records showed possible concurrent chats, and 193 showed 4 to 603.
IRS technology officials named one possible cause, a miscalculated handle time: how long an assistor takes to serve one chat. Assistors may leave finished chats open, and the platform closes them after four hours. As of December 2025, the IRS and the vendor had not found the root cause.
In April 2025 the IRS limited assistors to one chat at a time. Reports for September to November 2025 still showed 23 assistors apparently working 2 to 11 chats at once.
The reports also held 635,684 resolution codes, the code an assistor records for how each chat ended, for 613,056 live chats. Managers knew the numbers did not reconcile, and did not investigate.
Unresolved chats
Of 635,684 chats reported, 46 percent were resolved and 54 percent were not, TIGTA counted. Of the unresolved, 136,838 were abandoned, 127,288 disconnected, and 81,377 referred to the ACS phone line. Abandoned chats rose from 33,781 in 2023 to 103,057 in 2024.
Assistors record no reason when a chat goes unresolved, so the IRS cannot say why. TIGTA says the exact number of unresolved chats cannot be determined from the data.
Assistors and the risk of disclosure
In a sample of 40 assistors, chosen by the auditors' judgment, 24 worked more than one live chat at a time. Twelve of them had an authenticated chat open while working another. TIGTA says this raises the risk of disclosing taxpayer information to the wrong taxpayer.
Managers believed authenticated chats were worked one at a time. Two assistors at two ACS locations said they could work several, authenticated or not.
The IRS cut its staff from about 103,000 to 77,000 between January and May 2025. ACS live assistors fell about 20 percent, from 159 to 127.
What the chatbot's answers showed
TIGTA tested the chatbot in March 2025. It found 29 answers across the 206 process flows it tested that could be improved, for example with broken links. Of 53 keywords typed into the search box, 44, or 83 percent, were not recognized or got an incomplete answer.
Some keywords returned computer programming code. The Spanish chatbot has no search box, so Spanish speakers can use only the preset options.
Feedback and oversight
The IRS has offered no survey after live chats since September 2022. The chatbot's survey tells taxpayers their answers help improve service, but officials said no one collects or reviews them. A new live-chat survey was still awaiting submission for clearance in December 2025.
TIGTA found the IRS out of compliance with the Office of Management and Budget's customer-experience requirements in Circular A-11.
Managers reviewed one chat per assistor each month, and the reviews could not count in appraisals. Managers said live chat was a pilot, though no policy barred evaluative reviews in a pilot. The Internal Revenue Manual had no guidance for managing live chat.
The sources describe no check of the reports against the chat records during the audit period, and no review of assistors that counted in appraisals. TIGTA made nine recommendations, and IRS management agreed with all nine. It said it had acted on four and planned action on five.
Earlier warnings, and the IRS's own claims
The IRS Advisory Council's November 2024 report found the bots were "not designed to provide the taxpayer a direct answer" but pointed to general information. It warned that several bot vendors created parallel channels that may confuse taxpayers.
It recommended single entry points, escalation to live agents, and testing with non-English speakers and users with disabilities.
A September 2022 IRS essay by Darren Guillot, a deputy commissioner, claimed over 450,000 chatbot interactions, 42 percent resolved without escalation. It claimed over 4.8 million calls to the first voice bots, 40 percent handled without a person. It claimed over 1 million calls to the authenticated voice bots since June 2022, which created, modified, or extended about 7,600 payment plans worth over $50 million.
Those are the IRS's own unaudited claims, and the page is now marked historical. The IRS's long-term AI strategy has been on hold since the office developing it was eliminated in March 2025.
Where the facts come from
The main source is TIGTA's audit, report 2026-308-029. News reports by FedScoop, CPA Practice Advisor, and Accounting Today report its findings. The IRS's own 2022 essay gives the early figures. A 2024 FedScoop report covers the Advisory Council's findings. No source names the chat vendor.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other. A closed pathway is one that few enough mistakes pass along. The work along it may go on.
This case has a budget of 10 units. Contained means the network's mistakes are corrected rather than building on each other. When the network opens, five failure pathways are open, and the pressure Monitoring goes stale is on. Two of those pathways, the assistors' account actions and their reading of accounts, drain the Privacy gauge.
Under Explore (No Targets), which sets no targets, one tool alone keeps the mistakes contained for 2 units: Escalate checks, or Mark AI-written records. Gate record entries or Store less data does it alone for 3.
Under Service Targets Only, each of those four tools also meets the targets alone. That level also asks for the automated system to be helping the work.
Under those two levels, with lingering effects on, the setting you first see, Upgrade model at its stronger setting also works alone, for 5 units. With lingering effects off, Peer sharing rules at its stronger setting works alone instead, for 3 units.
Service and Safety Targets and All Governance Targets also ask you to close every failure pathway and keep up with the work, among other targets. Both can be met. The cheapest way costs 7 of the 10 units: Mark AI-written records, Escalate checks, and either Store less data or Gate record entries.
Every combination that meets those two levels includes Mark AI-written records and Escalate checks. Escalate checks is the one tool that closes the hand-off from the chatbot to live chat. Mark AI-written records is the one tool that closes the assistors' reading of taxpayer accounts.
Counting each tool once at either setting, 15 different sets of tools meet the Service and Safety Targets within the budget. 13 meet All Governance Targets, which asks for the system to be clearly helping the work. Two sets with Gate record entries fall short there: one adds Review on schedule, the other Check copied records.
Review on schedule, Check copied records, Vet connections, Upgrade model, Keep skills sharp, and Peer sharing rules close no open failure pathway here. Each can join a combination that meets the targets, but none of them is needed.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the IRS ACS chatbot-class collection-deflection funnel network: 6 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 6 assumptions
- assumed
This example shows the missing measurement documented in the case file for the IRS ACS chat applications. It is not a copy of the real system. The ACS chatbot is scripted, does not use generative AI, and makes no eligibility or scoring decision. What matters is what the IRS did not measure. TIGTA found the IRS had no performance measures for the program while it expanded it and made live chat permanent. Reading it as a system that decides things, or as generative AI, misreads it.
- baseline
This example treats the statistical reports as the part that carries the mistakes. TIGTA found them unreliable. The concurrent-chat figures showed up to 603 chats at once against a limit of three. IRS technology officials named a miscalculated handle time, how long each chat takes, as one possible cause. Over 16,000 records showed possible concurrent chats, and the reports held 22,628 more resolution codes than chats. Managers used the reports, knew the counts did not reconcile, and did not investigate. This example assumes nothing cleans mistakes out of the reports, because ACS managers did not make sure the report data was accurate. It also assumes wrong figures that look complete are worse than a visible gap, because people read them as fact.
- baseline
This example treats TIGTA's audit as the one check that worked, and that check sits outside the IRS. It treats an evaluative review of assistors as a check the IRS did not run during the audit period. GOV.UK Chat, the UK government's public chatbot and another case in this Lab, shows the opposite pattern. The UK Government Digital Service held back a 2023 version of it that missed its accuracy bar. Here the audit alone found the problems, once. For its concurrency test it set aside most of the concurrency records.
- baseline
This example also ties assistors' workload to the risk of disclosure. In a sample of 40 assistors, chosen by the auditors' judgment, 24 worked several chats at once. Twelve of those had an authenticated chat open while working another. TIGTA says this raises the risk of disclosing taxpayer information to the wrong taxpayer. It places that risk on the assistors' account actions and on their reading of accounts. ACS live assistors fell from 159 to 127, about 20 percent, as the IRS cut about a quarter of its staff. This example assumes the cut raised each remaining assistor's load, and the unreliable reports kept any such rise out of sight. The sources give the cut, not its effect on each assistor. These are signals about the institution's records. No disclosure to any real taxpayer is claimed or calculated.
- assumed
The 603 figure and the 54 percent of chats unresolved come from data TIGTA itself calls unreliable. TIGTA says the exact number of unresolved chats cannot be determined. The 60 percent figure comes from 40 of 253 assistors, chosen by judgment, not at random, so it cannot be applied to all of them. TIGTA judged only a narrow band of figures likely accurate, 1.20 to 3.88 chats at once on average. The audit did not test the voice bots, so its error findings cover the chat programs only. The 2022 figures are the IRS's own unaudited claims, from an essay now marked historical. They are over 450,000 chatbot interactions, over 4.8 million calls to the first voice bots, and about 7,600 payment plans created, modified, or extended, worth over $50 million. No source names the chat vendor, so this example names none.
- assumed
The taxpayers who use the chat, and whether their questions are resolved, are not shown in this example. It shows how mistakes move between the chat platform, staff, and records. The Spanish chatbot has no search box, so Spanish speakers can use only preset options. That unequal access is recorded here, not calculated as an outcome for any group. Of 635,684 chats reported, 54 percent went unresolved, and abandoned chats roughly tripled from 2023 to 2024. In 81,377 chats, the assistor told the taxpayer to call the ACS phone line, the line the chat was meant to relieve. The IRS could not support its claim that the chatbots cut phone demand, because it had not analyzed their performance. A concurrency record, a resolution code, or a disclosure risk here is a signal about the institution, never a person. A safe starting point in this example is not a safety promise for any real deployment.
What this example does not show
Show all 4 limitations
- The taxpayers who use the chatbot, the live chat, and the voice bots are not shown in this example. It shows how mistakes move inside the institution. The risk that an assistor discloses one taxpayer's account in another's chat is a signal about a record pathway. It is never a calculated disclosure to any real person. The Spanish chatbot's missing search box is recorded as a concern about unequal access. It is not an outcome calculated for any group.
- The 603 figure and the 54 percent unresolved figure come from data the audit itself calls unreliable. The audit says the exact number of unresolved chats cannot be determined. The 60 percent figure comes from 40 assistors chosen by judgment, not at random, and cannot be applied to all 253. The audit judged only concurrency figures from 1.20 to 3.88 chats at once likely accurate. This example reads these figures as signs of a monitoring failure, not as exact measures.
- The audit did not test the voice bots. So its findings on the chatbot's content, 29 answers across 206 process flows and 83 percent of tested keywords, cover the chat programs only. The 2022 figures are the IRS's own unaudited claims, from an essay now marked historical. They are over 450,000 chatbot interactions and over 4.8 million and over 1 million calls to two kinds of voice bot. They include about 7,600 payment plans created, modified, or extended, worth over $50 million. The audit found the IRS could not support its claim that the chatbots cut phone demand.
- This is a scripted chatbot that does not use generative AI. It does not score people or decide eligibility. No source names the platform vendor. The case matters here because the chat surveys looked like a feedback channel, but no one read the answers. An occasional outside audit was the one working check. It also shows how assistors' workload ties to the risk of disclosure. A safe starting point in this example is not a safety promise for any real deployment.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A June 2026 Treasury Inspector General for Tax Administration performance audit (Report Number 2026-308-029) reported that the IRS expanded its Automated Collection System chatbot and live-chat program and made live chat permanent while having no performance measures for it, despite a Taxpayer First Act requirement for metrics and benchmarks, and that management's claim the bots reduced telephone demand could not be substantiated; the statistical reports the IRS did collect were deemed unreliable, in one instance showing a single assistor apparently working 603 chats at once against a systemic cap of three, attributed partly to a miscalculated handle-time metric the vendor had not resolved as of December 2025.
empirical- Government evaluation Treasury Inspector General for Tax Administration, Opportunities Exist to Improve the Quality of Chat Applications (Final Audit Report, Report Number 2026-308-029, 2026) https://www.oversight.gov/sites/default/files/documents/reports/2026-06/2026308029fr.pdf
- Trade press Bracken, IRS live chat apps have room for improvement, watchdog finds (FedScoop, 2026) https://fedscoop.com/irs-live-chat-apps-chatbots-report/
- Trade press Bramwell, Are IRS Chatbots Really Helping Taxpayers? (CPA Practice Advisor, 2026) https://www.cpapracticeadvisor.com/2026/07/08/are-irs-chatbots-really-helping-taxpayers/186261/
In the same audit, of a judgmental sample of 40 IRS ACS live assistors, 24 (60%) were found working multiple chats concurrently and 12 of those 24 had at least one authenticated chat open while working another, which TIGTA reported as raising the risk of disclosing taxpayer information to the wrong taxpayer; the audit also reported 635,684 resolution codes against 613,056 chats (a mismatch management knew of but did not investigate) and, in March 2025 hand-testing, 14% of chatbot process flows deficient and 83% of tested keywords unrecognized or insufficient, with the figures drawn from a nonprobability sample and data the audit itself characterized as unreliable and not projectable to the full assistor population.
empirical- Government evaluation Treasury Inspector General for Tax Administration, Opportunities Exist to Improve the Quality of Chat Applications (Final Audit Report, Report Number 2026-308-029, 2026) https://www.oversight.gov/sites/default/files/documents/reports/2026-06/2026308029fr.pdf
- Trade press Bramwell, Are IRS Chatbots Really Helping Taxpayers? (CPA Practice Advisor, 2026) https://www.cpapracticeadvisor.com/2026/07/08/are-irs-chatbots-really-helping-taxpayers/186261/
- Trade press Cohn, IRS chatbot results may be wrong (Accounting Today, 2026) https://www.accountingtoday.com/news/irs-chatbot-results-may-be-wrong
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Benefits navigation & public-facing chat domain page.
Levers available here and the patterns behind them
- Review on schedule — Oversight cadence & retrospectives
- Mark AI-written records — Provenance labeling
- Gate record entries — Human-in-the-loop write gating
- Store less data — Data minimization
- Escalate checks — State-feedback vigilance
- Keep skills sharp — Deskilling-arrest mandate
- Peer sharing rules — Peer-edge governance
- Check copied records — Reconcile copied records
- Vet connections — Connection authorization
- Upgrade model — Improve the model
Documented case histories
- IRS collection chatbots: expanded and made permanent with no performance measures
- Nava assistive benefits chatbot
- Caddy adviser copilot at Citizens Advice
- GOV.UK Chat
- Mass.gov Virtual Assistant
- Frida (NAV Norway)
- SSA 800-Number Conversational AI Assistant
- EDD Virtual Assistant
- Burokratt
- Singapore's chatbot fleet refresh: eighty scripted engines slated for retirement onto a shared LLM platform
- Albert France Services
- Propel in-app SNAP benefits assistant
- GetCalFresh: the nonprofit front door that carried most of California's online SNAP intake
- MyFriendBen benefits screener
- Benefits Data Trust wind-down