Skip to content

PAN Lab example

A contact centre's generative-AI agent assist

Fifteen percent on average and thirty for the newcomers

A Fortune 500 software firm gave support agents an AI copilot. On average they resolved 15 percent more issues per hour. Newer agents gained most.

See more

The copilot is a generative AI tool built on OpenAI's GPT-3 language model and further trained on past support chats. It watches each customer chat, suggests replies to the agent, and links to the firm's technical documentation. The published study does not name the tool, its maker, or the firm.

Where it was used

The firm is a Fortune 500 company that sells business-process software to small and medium-sized businesses in the United States. Its chat-support agents answer technical questions from small business owners in the United States. A support chat averages about 40 minutes, most of it spent diagnosing the problem.

Most agents in the study, 89 percent, worked outside the United States, mainly in the Philippines. Some worked for outsourcing firms. Agents handle chats alone, each in a team with a manager who gives feedback and training.

How the copilot was rolled out

The rollout took place mainly in fall 2020 and winter 2021. An agent got access after one three-hour online session run by the AI company, and no further training on the tool. Managers chose which agents went to each session.

The firm set a limited budget for the tool, which capped how many agents could use it. That budget and the limited training sessions meant agents got access at different times. A small pilot came first, in August 2020: about 50 agents, about half of them chosen at random to get the tool.

How it is used

Only the agent sees the copilot's output. The agent decides which suggestions, if any, to use, and stays responsible for what the customer is told. On average, agents adopted 38 percent of the suggestions they received.

What the study found

Erik Brynjolfsson, Danielle Li, and Lindsey Raymond studied the rollout. The Quarterly Journal of Economics published their paper, "Generative AI at Work", in 2025. A 2023 working paper from the National Bureau of Economic Research came first.

The study compares agents before and after they got the copilot with agents who did not yet have it or never got it. Apart from the small pilot, it is not a randomized trial. Its data cover about 3 million chats by 5,172 agents. Of these, 1,636 agents used the copilot during the period studied.

On average, agents with the copilot resolved 15 percent more issues per hour. The working paper reported 14 percent. Customers' messages to agents became more positive in tone, and fewer customers asked for a manager. Surveyed customer satisfaction did not change on average.

Fewer agents quit, most clearly among those with under six months on the job. The authors caution that this result may overstate the copilot's effect.

Who gained

Almost all the gain went to newer and less skilled agents. They resolved about 30 percent more issues per hour in the published paper, and 34 percent in the working paper. Agents with two months on the job and the copilot performed as well as agents with more than six months without it.

The most skilled and most experienced agents gained little or nothing. The study found some evidence that the quality of their chats dipped slightly.

Why one average misleads

The copilot mostly raised the floor. It pulled novices up toward where experienced agents already were, and did little at the top. So a 15 percent average describes almost no single agent.

The average overstates the gain for experienced agents and understates it for novices. It also hides the small quality cost at the top. An organization that reports the average and plans staffing around it is managing a benefit it has not measured.

What the gain is worth

The benefit is real and worth having. It shortens the time new agents take to get up to speed. It makes hard early months easier, which may help explain why fewer agents quit. Calling it "15 percent more productive" still misstates whom it helps. It invites two mistakes: expecting gains from experienced agents that will not come, and missing the quality cost at the top.

Open questions

The copilot's training gave extra weight to top performers' chats, and their chats are used to train it further. The authors suggest it passes the best agents' practice on to everyone else.

They raise two questions. Heavy reliance on suggestions could lower the quality of future training chats. Pay may need to reward the agents whose chats train the tool.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. The cheapest combination of tools that meets them uses two tools, Review on schedule and Escalate checks. It costs 4 of this case's 9 budget units.

Under Service and Safety Targets and under All Governance Targets, this case is not fully addressable with the available tools. Both levels ask you to close every failure pathway, among other targets. Every combination of tools that fits the budget was checked, and none meets the targets at either level.

Seven failure pathways stay open under every combination. Every tool offered here, each at its strongest setting, costs 29 units and still leaves the same seven open.

They are the deployment's everyday work. Novice and experienced agents use or edit suggestions, and both groups' resolved chats are saved. The copilot reads the chat history, novice agents read what it retrieves, and quality assurance reads resolved chats.

One tool, Check copied records, has nothing to act on here, because this network has no copied records. That is a finding about the deployment, not a flaw in your choices.

Stylized model of a documented deploymentCustomer service & contact-centre AI

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Agent-assist-copilot-class with skill compression network: 6 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 5 assumptions
  • assumed

    The published study describes two kinds of copilot output: suggested replies and links to the firm's technical documentation. This network shows the documentation links as their own part. It assumes that finding the right article matters most to an agent who does not yet know where to look. The study's authors point more to the suggested replies: agents who followed them more closely gained more. They say the links may be particularly valuable for less common problems. The network assumes a heavy workload for about 5,000 chat-support agents, and staff capacity that is limited but not strained. It assumes capacity is not strained because the study found fewer agents quit once they had the copilot. The study also records limits. Training resources and the budget for the tool were limited, and agents got no training on it after the first session. Quality assurance is assumed to read the chat records it evaluates.

  • baseline

    This network models the pattern the case describes, not a reconstruction of the actual deployment. The pattern is a gain that differs sharply by agent skill. The gain is real and carefully measured. The study compared agents who had the copilot with agents who did not yet have it, or never got it. It covered 5,172 agents. Issues resolved per hour rose about 15 percent on average. Customers' messages grew more positive in tone, and fewer agents quit. The network gives the two agent groups the same setup, because the point is that the gain is a spread, not a single number.

  • assumed

    The two agent groups have identical parts and links on purpose. The copilot and the work are the same for both. The difference the study documents is a measured result. Novice and lower-skilled agents gained about 30 percent in the published paper, and 34 percent in the 2023 working paper. The most experienced gained close to nothing, with a slight loss of quality. The network does not compute that split. It is a result recorded in the case file. It appears here only as the difference the benefit measurement exists to reveal.

  • baseline

    The reading behind this network is that an agent-assist copilot mostly raises the floor. One average then overstates the gain for the experienced agents who least need help. It also hides the slight quality cost the same study found for them. The network includes two checks that turn one number into a measured spread. One measures the copilot's benefit at each level of agent skill. The other is a scheduled quality review aimed at the most experienced agents. The published study is itself such a measurement. The sources do not say whether the firm repeats it.

  • assumed

    No customer outcome is modeled here. The network follows how mistakes move between the copilot, the agents, and their records. The customers being served are outside it. The productivity figures, the uneven gain across skill, and the slight quality loss at the top come from the case file. None of them is computed from this network.

What this example does not show

Show all 2 limitations
  • This example does not model customers. The productivity figures, the finding that the gain went mostly to newer agents, and the slight quality loss at the top come from the case file. None of them is computed here. The Lab shows how mistakes move between the copilot, the agents, and their records.
  • The figures are measured results from a peer-reviewed field study: about 15 percent on average, about 30 percent for novices, and close to nothing for the most experienced. The study compared agents who got the copilot at different times. Apart from a small pilot, it was not a randomized trial. The two agent groups have identical parts and links on purpose. The difference between them is a measured result, not a difference in how their work is set up.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The strongest field evidence for an agent-assist copilot in customer service comes from a staggered rollout of a generative-AI assistant to customer-support agents at a large software firm, studied across 5,172 agents, 1,636 of whom used it. Measured against agents who did not yet have it or never got it, the copilot raised issues resolved per hour by about 15 percent on average, and it also improved customer sentiment and agent retention. The gain, however, was sharply uneven: novice and low-skill agents improved by about 30 percent (34 percent in the 2023 working paper), agents with two months of experience performed like agents with six months and no AI, and the most experienced agents gained close to nothing, with some evidence of slight quality degradation. This is the contact-centre domain's cleanest measured benefit, and it is a distribution rather than a single number.

    empirical
    • Peer-reviewed Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2), 889-942. https://doi.org/10.1093/qje/qjae044 https://academic.oup.com/qje/article/140/2/889/7990658
    • Academic Brynjolfsson, E., Li, D., & Raymond, L. (2023). Generative AI at Work. NBER Working Paper 31161 https://www.nber.org/papers/w31161
  • The lesson the field evidence carries is skill compression: an agent-assist copilot mostly raises the floor. Because almost the entire measured gain accrues to less-experienced agents and the most experienced gain close to nothing, an average productivity number overstates the effect for the agents who least need it and hides that the tool does little for the experienced while possibly costing a small amount of quality there. The governable reading is that the benefit must be measured as a distribution across agent skill, not reported as a scalar — a copilot that helps novices a great deal and experts not at all is a real and specific benefit, and describing it with one average misstates who it helps.

    empirical
    • Peer-reviewed Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2), 889-942. https://doi.org/10.1093/qje/qjae044 https://academic.oup.com/qje/article/140/2/889/7990658

Where this connects

Institutional pressures in this domain

  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).

All of them in context on the Customer service & contact-centre AI domain page.

Levers available here and the patterns behind them

Documented case histories