Skip to content

PAN Lab example

Massachusetts DTA call summaries

The record and its source: AI summaries of benefits calls

Massachusetts piloted AI summaries of food-benefit eligibility calls. Each summary is saved into the case record, and the full call transcript is not kept.

See more

The DTA call summarizer is software the Massachusetts Department of Transitional Assistance (DTA) built with Accenture. It transcribes SNAP eligibility calls on the DTA Assistance Line as they happen and drafts a structured summary from a prompt, written instructions DTA will not release. A caseworker can edit the summary before saving it into BEACON, the state's eligibility system of record.

What the pilot is

SNAP, the Supplemental Nutrition Assistance Program, is the federal food-benefits program. DTA runs it in Massachusetts. The pilot began in December 2025. The summarizer produces no score, no recommendation, and no eligibility decision.

The calls can include Social Security numbers, medical history, and immigration status. DTA says callers are notified before being connected and can opt out. That is an agency claim, and no opt-out figures are published.

What is kept and what is not

According to technical documentation reviewed by The Shoestring, an investigative news outlet, full transcripts of the calls are not saved. Only the AI summaries are kept. DTA declined to release the prompt that generates them, citing "the proprietary nature of the prompt development."

So the summary is not a note kept beside its source. It becomes the official account of what the applicant said. A mistake in it, such as something the caller said that the summary left out, or a mis-heard figure, has nothing left to be checked against. The only competing account is the client's memory.

DTA presents not keeping transcripts as a privacy protection. So one design choice both protects privacy and removes the means of checking the record.

Who relies on the summary

The summarizer never deals directly with the eligibility workers, supervisors, and hearings staff who act on its text. They read the summary in BEACON as the account of an interview they did not hear.

In December 2025, DTA decided 31,390 SNAP applications, 47% of them approvals. That month 56,454 recertifications were due. SNAP served 1,011,460 people, one in six Massachusetts residents.

How far the pilot reaches

About 400 calls had gone through the tool between the December rollout and the April 2026 reporting, DTA says. In December 2025 alone, 45,703 callers were connected to staff. A daily average of 6,454 callers could not connect at all. The pilot touched well under 1% of connected calls: a pilot with room to expand.

What the agency says

DTA says the tool is designed to cut call handling times, make case notes more consistent, and free caseworkers to focus on the conversation. No independent evaluation of these claims has been published.

Who shaped the deployment

Oversight ran through the workers' union. SEIU Local 509 reached an agreement with DTA that made using the tool voluntary and protected jobs. It did not address record provenance, meaning who or what wrote each part of the record, or safeguards for callers.

The formal privacy process was empty. The state's internal inventory lists at least 40 AI uses, and nine were disclosed after a public records request. None of the nine, this summarizer included, reported a completed privacy impact assessment. That is a review of what personal data a system handles and how it is protected.

The tool's documentation listed its plan for interaction data as still to be defined, while it was in production. After an appeal, the Supervisor of Records ordered the withheld records on the other 31 uses submitted for private review. That state official rules on public records appeals. The state's public AI page does not name the DTA summarizer.

A contested record

The records the summaries join were themselves under a federal demand. Massachusetts was one of 21 states, with the District of Columbia, suing in California v. USDA over the U.S. Department of Agriculture's demand for personal SNAP data. Preliminary injunctions in October 2025 and February 2026 blocked USDA from cutting funding over the states' refusal. The litigation is live.

Whether AI summaries held in BEACON fall within the demand is unresolved. The Massachusetts Attorney General's office declined to comment on that question.

A separate contract

In February 2026 the state also signed an enterprise contract with OpenAI, to give a general-purpose AI assistant to roughly 40,000 executive-branch employees. That is a separate track. Another state union demanded bargaining over it, and legislators criticized it.

State Representative Erika Uyterhoeven asked for "specific protections for residents who interact with MassHealth and other safety-net programs." The documented pressure targets that contract. Its bearing on the DTA pilot is inferred, not documented.

Where the facts come from

The core public record is The Shoestring's April 2026 investigation, a single outlet. The caseload and call figures come from DTA's December 2025 performance scorecard. The court case comes from the Massachusetts Attorney General's office and from JURIST, a legal news service.

As of April 2026 there is no published evaluation, no inspector-general audit, no privacy impact assessment, and no documented harm to any individual. The concern is structural and about the future. By the time any harm is judged, the evidence needed to judge it will already have been discarded, call by call, by design.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A closed pathway is one that mistakes stop passing along. The work along it may go on.

This case has a budget of 10 units. Its two pressures, Workload surges and Autonomy expands, are on at the start. Mistakes are then copied faster than anyone corrects them, which the Lab calls a cascading failure. Nine failure pathways are open.

Explore (No Targets) sets no targets. There, keeping the mistakes contained takes at least 8 units. Contained, which the Lab calls self-correcting, means mistakes are caught faster than they are passed on. One way is Mark AI-written records and Gate record entries, with Gate vendor updates at its stronger setting.

Under Service Targets Only, the same combinations also meet the service target. That target asks for the summarizer to be helping the work. Lingering effects is a Dynamics setting in which damage outlasts its cause, and it is on by default. With it on, 23 different sets of tools within the budget meet this level's targets. With it off, 31 do. Every one includes Mark AI-written records.

Under Service and Safety Targets and All Governance Targets, the targets are not fully addressable with the available tools. Both levels ask you to close every failure pathway. Four stay open whatever you choose: Draft summary to the caseworker, Entry in the AI inventory, Caseworker chooses and edits, and One prompt for every call. No tool on offer acts on them.

More money does not change that. Every tool at its strongest setting at once costs 24 units, far over the budget. It contains the mistakes but leaves those four pathways open. It also leaves the summarizer helping too little to meet the service target.

Stylized model of a documented deploymentCaseworker documentation & copilots

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Call-summary-class benefits documentation copilot network: 7 components and 14 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 6 assumed · 14 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 20 assumptions
  • assumed

    This network follows the pattern the Massachusetts DTA call summaries case file describes. It does not rebuild the actual pipeline. DTA withheld the summarization prompt as proprietary. So how the tool turns a call into a summary is not public, and the network does not guess at it.

  • baseline

    The network includes a check, Summary checked against the call, that would compare each saved summary with the call it came from. DTA's design keeps the editable summary and discards the full transcript. So the check has nothing left to read. That is the documented design, not a review someone skipped. DTA presents not keeping transcripts as a privacy protection. So a privacy control and an accountability control rest on the same choice.

  • assumed

    The network draws no link from BEACON back into the summarizer. The summarizer reads the live call. The sources describe no use of past case records in a summary and no retraining on them. The tool's documentation listed its plan for interaction data as still to be defined while it was in production. Most documentation tools in the Lab draw that link back, and a few do not. Drawing it here would claim a reuse the sources do not describe.

  • baseline

    The network shows two groups of workers because the sources describe two groups with different access to the facts. The caseworker took the call, heard it, and may edit the summary. Determination, recertification, and hearings staff later read that summary as the account of an interview they did not hear. That difference between them is the case.

  • baseline

    The network assumes the whole two-party call goes into the summarizer unredacted, Social Security numbers, medical history, and immigration status included. The sources describe no filter or redaction between the call and the summarizer. For the same reason, the network marks that link, and the saving of the summary into BEACON, as privacy-sensitive.

  • baseline

    The network draws the two links between the caseworker and the summarizer from the documented steps, not from call volume. The draft summary stands in for the note the caseworker would otherwise write, which is DTA's stated purpose for the tool. Two steps limit it: the caseworker may edit the draft, and use is voluntary under the union agreement. The caseworker's own link to the summarizer carries two choices. They choose which calls go through the tool, and they edit before saving. The caseworker does not shape how the summary is written, because the prompt is withheld.

  • baseline

    The network assumes every call that goes through the summarizer ends in BEACON. On save, the summary becomes part of the client's case record as the account of the interview. Staff also write their decisions back into BEACON. What they write is a decision reached from the summary, not the interview account itself.

  • baseline

    The network includes two further links that the documented design does not use. One is the summarizer saving into BEACON without a caseworker's save. DTA's design sends every summary through that save. It is included because the pressure Autonomy expands would act on it. The other is a later reader putting a disputed line back to the caseworker. The sources say no one downstream can compare the summary with the call after the save. They say correcting the record falls instead on the client's memory against the state's system of record.

  • baseline

    The network assumes later decisions lean on the saved summary, from documented scale and documented reliance. The summary is the lasting account of the interview. Future determinations, recertifications, discrepancy checks, and fair hearings consult it. In December 2025 there were 31,390 SNAP application decisions and 56,454 recertifications due. Monthly churn was 24%, and a new application took 12 days on average to approve.

  • baseline

    The network assumes the tool spreads slowly between caseworkers, because the pilot is small. Use is voluntary under the union agreement, so the network assumes the tool spreads by colleague example. About 400 calls had gone through it between the December 2025 rollout and the April 2026 reporting. In December 2025 alone, 45,703 callers were connected to staff. That is well under 1% of connected calls. This is a pilot with room to expand, not a deployment already in full use.

  • baseline

    One pipeline and one withheld prompt handle every call in the pilot, and the pilot is small. So a weakness in the transcription or the prompt repeats across calls instead of averaging out. The callers speak many languages. After English, the most common are Spanish, Haitian Creole, Chinese, Portuguese, and Vietnamese. There are 263,828 recipients aged 60 or over and 308,652 with a disability. So a uniform pipeline's weaknesses fall on the same groups each time. That makeup of the caseload is documented. No difference in harm between groups is computed here.

  • baseline

    The oversight part of the network is the Executive Office of Technology Services and Security (EOTSS), which owns the state's generative-AI policy and its internal AI inventory. The summarizer appears to EOTSS as an inventory entry, not as a feed of its summaries for review. It is one of nine uses disclosed from an inventory of at least 40. None of the nine, this one included, reported a completed privacy impact assessment, though the inventory had fields for one. The tool's documentation listed its plan for interaction data as still to be defined while it was in production. So the assessment link from EOTSS does nothing in the deployment as documented.

  • baseline

    The network does not show the union, SEIU Local 509, as a separate part. Its agreement with DTA made use voluntary for workers and added job protections. It reviewed no summaries and did not address record provenance or safeguards for callers. Its documented effect is the caseworker's choice of which calls go through the summarizer. So the network shows it as part of the caseworker's link to the summarizer, not as a reviewer. The sources do not establish that the agreement came before the December 2025 rollout, and the network does not assume it did.

  • assumed

    The callers are not part of the network. DTA says callers are notified before connection and can opt out, but no opt-out figures are published. That step is recorded in the case file, not drawn as part of the network.

  • baseline

    The network shows one link by which records leave the state. While the pilot ran, BEACON's records were under a demand from the U.S. Department of Agriculture (USDA) for personal data on SNAP applicants and recipients. In California v. USDA, No. 3:25-cv-06310 (N.D. Cal.), preliminary injunctions of October 15, 2025 and February 27, 2026 blocked USDA from cutting funding over the states' refusal. The court held that the proposed data protocol would likely allow sharing beyond the entities permitted under 7 U.S.C. 2020(e)(8). Both orders are preliminary, and the litigation is live. Whether AI summaries in BEACON fall within the demand is unresolved, and the state Attorney General's office declined to comment. Two other ways out are not drawn. The sources document no caseworker pasting call details into an unapproved tool. DTA says summary data is held in state-owned systems, which is itself an agency claim. The state's separate enterprise AI contract is a different matter and is kept apart from this pilot.

  • assumed

    The network draws no second model checking the summarizer. None is documented or planned, and no independent evaluation of the tool has been published. A second model would also face the same discarded transcript. So the summarizer's only link to itself is the one prompt used on every call. The network leaves out links the sources do not describe.

  • baseline

    The network assumes the Assistance Line is under strain, and that caseworkers did the work competently before the tool. In December 2025 the Assistance Line connected 45,703 callers to staff. Each day, on average, 2,709 calls were connected, 5,427 were completed in the automated phone menu, and 6,454 callers could not connect at all. The program served 1,011,460 people in 622,837 households, one in six Massachusetts residents. The work without the tool is not assumed to have been better. Nothing documents the tool replacing a clearly better process, and no evaluation exists either way. The manual process also kept no verbatim record: caseworkers wrote notes from memory. What the tool changes is who writes the account, not whether a recording survives.

  • assumed

    The agency's claimed benefits are unmeasured and are treated as claims. DTA says the tool cuts call handling times, makes case notes more consistent, and frees caseworkers to focus on the conversation. It also says data is stored in state-owned systems, follows existing access and retention policies, and callers can opt out. No independent evaluation, inspector-general audit, or completed privacy impact assessment of the tool is on record. Edit rates and edit depth at the caseworker's save are unpublished. The network computes no benefit figure from these claims.

  • baseline

    The picture rests on one outlet. The finding that transcripts are not kept, the call count, the withheld prompt, and the blank assessment fields all come from The Shoestring's April 2026 investigation. It cites technical documentation and DTA statements on the record. The caseload, call, decision, language, age, and disability figures come separately from DTA's own December 2025 performance scorecard. The scorecard notes that its churn rate was recently updated after a computational error. No court or agency has found that any individual was harmed. The concern this network shows is structural and about the future.

  • assumed

    Different effects on different groups of callers are documented, never computed. The caseload is disproportionately elderly, disabled, and non-English-speaking. Transcription quality across those groups has not been measured. Because the transcripts are discarded, the evidence that could measure it is discarded with them. The case file records this. The Lab shows how mistakes pass between the organization's parts, and estimates no harm to the people served.

What this example does not show

Show all 8 limitations
  • No outcome for any caller is modeled. The Lab follows how mistakes pass between the organization's parts. Applicants for and recipients of SNAP, the Supplemental Nutrition Assistance Program, are outside the network. The makeup of the caseload, the correction burden on a client's memory, and the caller notice and opt-out are recorded in the case file. None of them is computed on this diagram.
  • The harm this network shows is structural and about the future. No court or agency has found it. The sources name no claimant whose benefits were affected by a summary error, and nothing here claims one. The concern is that the evidence needed to show such a case is discarded by design before anyone could bring it.
  • The picture rests on one outlet. The finding that transcripts are not kept, the call count, the withheld prompt, and the blank assessment fields all come from The Shoestring. It cites technical documentation and DTA statements on the record. No inspector-general audit, agency evaluation, or court filing about this tool exists. The caseload, call, decision, language, age, and disability figures come separately from DTA's own performance scorecard. The scorecard notes its churn rate was recently updated after a computational error.
  • Every benefit and safeguard DTA describes is an agency claim, labeled as one and never measured here. They are shorter call handling times, more consistent notes, storage in state-owned systems, following existing access and retention policies, and a caller opt-out. Opt-out rates, summary edit rates, and edit depth are unpublished.
  • This is a pilot, and the network shows it as one. Roughly 400 calls had gone through the tool by the April 2026 reporting. Roughly 45,000 callers a month were connected to staff, 45,703 in December 2025 alone. That is well under 1% of connected calls. The network shows a deployment with room to expand, not one already in full use.
  • The litigation over these records is live, and its orders are preliminary. Preliminary injunctions were issued on October 15, 2025 and February 27, 2026 in California v. USDA, No. 3:25-cv-06310 (N.D. Cal.). Whether AI summaries held in BEACON fall within the demanded data is unresolved. The state Attorney General's office declined to comment on it, and the question stays open here.
  • The timing of the union agreement stays uncertain everywhere, the diagram included. The sources establish that an agreement exists that made worker use voluntary and protected jobs. They do not establish that it was concluded before the December 2025 rollout.
  • This pilot is not the state's separate February 2026 enterprise AI contract. The documented pressure from legislators and a union targets that contract. Any bearing on this pilot is inferred, and the network draws no link for it.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • In December 2025 the Massachusetts Department of Transitional Assistance piloted an Accenture-built tool that transcribes SNAP eligibility calls in real time and generates a structured, caseworker-editable summary that is saved into BEACON, the state's benefits eligibility system of record; full transcripts are not retained, only the summaries, according to technical documentation reviewed by The Shoestring, and the summarization prompt was withheld as proprietary. About 400 calls had been processed by the April 2026 reporting, against roughly 45,000 calls connected to staff per month, in a record system that fed 31,390 SNAP application dispositions in December 2025 alone. No independent evaluation, inspector-general audit, or completed privacy impact assessment of the tool is on record.

    empirical
    • Investigative The Shoestring, Massachusetts' AI program is more than meets the eye (Jonathan Gerhardson) (2026) https://theshoestring.org/2026/04/08/massachusetts-ai-program-is-more-than-meets-the-eye/
    • Government Massachusetts Department of Transitional Assistance, DTA Performance Scorecard, December 2025 (2025) https://www.mass.gov/doc/performance-scorecard-december-2025-0/download
  • Oversight of the DTA call summarizer ran through the labor channel: SEIU Local 509, representing DTA call-center workers among roughly 9,000 state employees, reached an agreement with DTA — pre- or early-deployment; the record does not establish it preceded the December 2025 rollout — that made worker use voluntary and protected jobs, without addressing record provenance or client-side safeguards. The formal privacy apparatus sat empty: none of the nine AI use cases Massachusetts disclosed from its internal inventory of at least 40, the DTA summarizer included, reported a completed privacy impact assessment; the tool's interaction-data plan was listed as still to be defined while it was in production; and details on the other 31 use cases were withheld until the Supervisor of Records ordered them submitted for in camera review.

    empirical
    • Investigative The Shoestring, Massachusetts' AI program is more than meets the eye (Jonathan Gerhardson) (2026) https://theshoestring.org/2026/04/08/massachusetts-ai-program-is-more-than-meets-the-eye/
    • Government Massachusetts Executive Office of Technology Services and Security, Artificial Intelligence at the Commonwealth (Mass.gov) (2026) https://www.mass.gov/artificial-intelligence-at-the-commonwealth
  • The record store the AI summaries enter was itself contested while the pilot ran: in California v. USDA, No. 3:25-cv-06310 (N.D. Cal.), a 21-state-plus-DC coalition including Massachusetts obtained preliminary injunctions on October 15, 2025 and February 27, 2026 blocking USDA from cutting SNAP funding over states' refusal to hand over personal SNAP applicant and recipient data, the court holding the proposed data protocol would likely permit sharing beyond the entities allowed under 7 U.S.C. 2020(e)(8). Both orders are preliminary and the litigation is live; whether AI-generated call summaries held in BEACON fall within the demanded data is unresolved, and the Massachusetts AG's office declined to comment on that question.

    empirical
    • Government Massachusetts Attorney General's Office, AG Campbell Secures Second Order Blocking Trump Administration From Cutting Off SNAP Funding Because of States' Refusal to Turn Over Personal Data of SNAP Applicants and Recipients (2026) https://www.mass.gov/news/ag-campbell-secures-second-order-blocking-trump-administration-from-cutting-off-snap-funding-because-of-states-refusal-to-turn-over-personal-data-of-snap-applicants-and-recipients
    • Investigative JURIST, US federal court blocks SNAP funding cuts over states' refusal to share recipient data (2026) https://www.jurist.org/news/2026/02/us-federal-court-blocks-snap-funding-cuts-over-states-refusal-to-share-recipient-data/
    • Investigative The Shoestring, Massachusetts' AI program is more than meets the eye (Jonathan Gerhardson) (2026) https://theshoestring.org/2026/04/08/massachusetts-ai-program-is-more-than-meets-the-eye/
  • Documented benefit-automation failures replicated determinations into downstream systems with no independent reconciliation against the source records — Michigan MiDAS actioned replicated flags and Robodebt reversed the onus onto recipients.

    empirical
    • Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
    • Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
    • Government Royal Commission into the Robodebt Scheme, Report (2023) https://robodebt.royalcommission.gov.au/publications/report
    • Investigative Law Society Journal, Crude, cruel and unlawful: Robodebt findings https://lsj.com.au/articles/crude-cruel-and-unlawful-robodebt-royal-commission-findings/
  • A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.

    empirical
    • Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
    • Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Caseworker documentation & copilots domain page.

Levers available here and the patterns behind them

Documented case histories