PAN Lab example
Klarna AI assistant
Two-thirds of chats handled and a year later a rethink
Klarna said its AI assistant handled about two-thirds of customer-service chats in the assistant's first month. About a year later, Klarna said quality had fallen.
See more
The Klarna AI assistant is the customer-service chat assistant in the app of Klarna, a Swedish payments company. It runs on a model from OpenAI and answers customers' questions in chat. Klarna ran it to handle as many chats as possible without a human agent.
What Klarna reported
Klarna published the assistant's first-month figures in a press release on 27 February 2024. It said the assistant handled about two-thirds of its customer-service chats, some 2.3 million conversations. It said that was the equivalent work of about 700 full-time agents.
Klarna said average resolution time fell from about 11 minutes to under 2. It said customer satisfaction matched that for human agents, and repeat inquiries fell by about 25 percent. It estimated a profit improvement of about 40 million US dollars.
Every figure is Klarna's own report, and no one audited them independently. OpenAI, which made the model, published a customer story titled "Klarna's AI assistant does the work of 700 full-time agents." Neither party was independent of the deployment. The case file says the figures were reported in good faith.
What Klarna said about a year later
In May 2025, Fortune reported that Klarna planned to hire humans again. Klarna's chief executive said cost had been "a too predominant evaluation factor." The chief executive added: "what you end up having is lower quality."
Klarna committed that "there will be always a human if you want." It began piloting an "Uber-style" flexible pool of human agents.
What the two halves show together
Both halves come from the same deployment. That makes this the clearest case among the Lab's customer-service deployments of a benefit followed by a cost. Nothing in the record says the first figures were false. The Lab's case file for this deployment, its written account of the sources, says the assistant very likely did handle two-thirds of chats.
The case file's lesson is that deflection is not resolution. Deflection counts the chats the assistant handled without a human agent. Resolution asks whether the customer's problem was solved. The two can move in opposite directions: faster, cheaper, and worse.
At Klarna, the quality cost stayed hidden until Klarna's own later judgment named it. No measurement built into the deployment from the start brought it to light.
The wider picture
A Gartner survey of 5,728 customers, published in July 2024, found 64 percent would prefer that companies not use AI in customer service. Sixty percent feared AI would make it harder to reach a person.
In June 2025, Gartner predicted that 50 percent of organizations will abandon plans to reduce their customer-service workforce because of AI. In its March 2025 poll of 163 service leaders, 95 percent planned to keep human agents.
What the case file says to do
The case file says a deployment like this should measure resolution and repeat contact against deflection, instead of counting deflection as a success by itself. It says the route to a person should be protected by design, before a reversal forces it.
What this network is drawn from
This network is drawn from the public record of this deployment. It shows the structure that record describes, not Klarna's own systems.
The case starts with one pressure already applied. A pressure is a change in the deployment's conditions that you can switch on or off. This one, Autonomy expands, stands for Klarna's leadership widening what the assistant handles. The chief executive's own account, that cost had become too predominant, is the basis for it.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other.
Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met within this case's budget of 10 units. The cheapest way costs 4 units and uses two tools, Mark AI-written records and Escalate checks. Each costs 2 units at every target level.
Under Service and Safety Targets and All Governance Targets, the targets are not fully addressable with the available tools. These levels require every failure pathway to be closed. Five pathways stay open whatever you pull. Two are Agents take over escalated chats and Agents record how chats ended. The others are Assistant's chats logged, Standing agents hand chats to the pool, and Pool agents record how chats ended. None of the tools on offer closes them, even at their strongest settings.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Deflection-assistant-class with the benefit-then-cost arc network: 5 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 5 assumptions
- assumed
This example draws the on-demand pool Klarna piloted in 2025 as its own group, separate from the standing agents. An agent called in for one chat was not there for the customer's earlier chats. So the two groups know different things about a customer. The example assumes the assistant hands a customer who asks for a person straight to that pool. The case file says a design judged on how few chats reach a person makes that route harder to use. The example also assumes the agents' workload is heavy for the staff available. Klarna said the assistant did the work of about 700 full-time agents, and later began hiring agents again.
- baseline
This example follows the pattern the case file documents: a benefit first, then a cost. It is not a copy of Klarna's actual assistant. The first-month figures are Klarna's own report and were not independently audited. The assistant handled about two-thirds of chats, some 2.3 million conversations, and did the equivalent work of about 700 full-time agents. Resolution time fell from about 11 minutes to under 2. Satisfaction was said to match human agents, repeat inquiries fell by about 25 percent, and profit was projected to rise by about 40 million US dollars. The example shows these as Klarna's own published figures, a claim the deployment made about itself. Nothing in the example computes them.
- baseline
The reversal is the same deployment's own later judgment. Klarna's chief executive said cost had become too predominant a factor, and the result was lower quality. Klarna committed to always keeping a human available. Nothing here says the deflection figures were false. The case file says deflection, pursued on cost, was the wrong thing to maximize. What mattered was whether customers' problems were solved, and how well. The example shows both the figures and the reversal as claims the deployment made about itself.
- assumed
The example includes two checks from the case file. Deflection is not resolution. The first, named Measure of resolution and quality, is the measurement the case file says the deployment lacked. It would have shown the quality cost before a reversal did. The second, named Route to a person, is the safety valve customers say they want. A design built to maximize deflection tends to erode it, and Klarna's reversal restored it. The customer survey and the industry forecast are in the case file. Nothing in the example computes them.
- assumed
This example does not model what happened to any customer. It shows how mistakes move among the assistant, the agents, and the interaction record. The customers being served are outside the network. The deflection figures, the reported satisfaction, the later fall in quality, and the survey results come from the case file. Nothing in this example computes them.
What this example does not show
Show all 2 limitations
- This example does not show any customer outcome. It shows how mistakes move among the assistant, the agents, and the interaction record, and the customers are outside the network. The deflection figures, the reported satisfaction, the later fall in quality, and the survey results come from the case file. Nothing in the network computes them.
- The first-month figures are Klarna's own report, and no one audited them independently. The reversal is Klarna's own later judgment. The network shows both as claims the deployment made about itself. It includes the measure of resolution and quality and the route to a person as two checks. It does not compute any harm.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
An organization published striking first-month numbers for its customer-facing AI assistant: it handled about two-thirds of customer-service chats (some 2.3 million conversations), was described as doing the equivalent work of about 700 full-time agents, cut average resolution time from about 11 minutes to under 2, was said to match human customer satisfaction, and was projected to improve profit by tens of millions. Every one of those figures was self-reported and not independently audited. Roughly a year later the same organization reversed course on quality grounds — its chief executive said cost had become too predominant an evaluation factor and the result was lower quality — and committed to always keeping a human available to customers who want one. This is the contact-centre domain's cleanest benefit-then-cost arc: the deflection numbers and the walk-back come from the same deployment.
empirical- Vendor Klarna Bank AB (2024, February 27). Klarna AI assistant handles two-thirds of customer service chats in its first month (press release via PR Newswire). https://www.prnewswire.com/news-releases/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month-302072740.html
- Trade press Ivanova, I. (2025, May 9). Klarna plans to hire humans again, as new landmark survey reveals most AI projects fail to deliver. Fortune. https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/
The lesson the benefit-then-cost arc carries is that deflection is not resolution. A published deflection number reports how many contacts the AI handled, not whether it handled them well, and a figure that is impressive on cost can hide a quality cost that only shows up later — which is what the organization's own reversal described. The survey backdrop sharpens it: most customers say they would rather not meet AI in service and fear it makes reaching a human harder, and industry analysts expect a large share of organizations to abandon plans to reduce their customer-service workforce with AI. The governable reading is to measure resolution and repeat contact against deflection rather than counting deflection as a win by itself, and to protect the path to a human as the safety valve a deflection-maximizing design tends to erode.
empirical- Trade press Ivanova, I. (2025, May 9). Klarna plans to hire humans again, as new landmark survey reveals most AI projects fail to deliver. Fortune. https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/
- Trade press Gartner, Inc. (2025, June 10). Gartner Predicts 50% of Organizations Will Abandon Plans to Reduce Customer Service Workforce Due to AI (poll of 163 service leaders). https://www.theregister.com/software/2025/06/11/half_of_firms_set_to_abandon_plans_to_ditch_customer_service/502135
Where this connects
Institutional pressures in this domain
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
All of them in context on the Customer service & contact-centre AI domain page.
Levers available here and the patterns behind them
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Train the staff — AI literacy & boundary rules
- Mark AI-written records — Provenance labeling
- Escalate checks — State-feedback vigilance
- Upgrade model — Improve the model