Skip to content

PAN Lab example

Klarna AI assistant

Two-thirds of chats handled and a year later a rethink

A customer-facing AI assistant handled about two-thirds of chats in its first month - some 2.3 million conversations, described as ~700 agents' worth of work, resolution time cut from ~11 minutes to under 2, satisfaction said to match humans. All self-reported. About a year later the same organization reversed on quality grounds - cost had become too predominant, and the result was lower quality - and committed to always keeping a human available. The deflection numbers and the walk-back come from the same deployment.

Stylized model of a documented deploymentCustomer service & contact-centre AI

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Deflection-assistant-class with the benefit-then-cost arc network: 5 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    The reversal is drawn as the structure it actually built. The record describes piloting a flexible, on-demand agent pool rather than re-staffing the standing function, and a contingent workforce reached on demand is a different group of staff, not more of the same one - what it knows about a customer differs, because it was called in for this contact rather than present for the last. The assistant hands off to that pool directly, but only faintly: the promise the reversal made was that a customer can always reach a person, and a design measured on how few contacts reach one makes that path narrow by construction. A heavy workload against very limited capacity - the human side was thinned on cost while the assistant absorbed two-thirds of contacts, which is the arrangement the reversal was reversing.

  • baseline

    This models the benefit-then-cost pattern documented in the case file - not a reconstruction of the actual assistant. The first-month figures (handled ~2/3 of chats / ~2.3M conversations, ~700-FTE equivalent, resolution time ~11min to <2min, claimed CSAT parity, ~25% fewer repeat inquiries, ~tens of millions in projected profit) are the organization's own self-report, not independently audited - drawn as the deployer's dashboard, a claim the deployment made about itself, not a measurement the diagram computes.

  • baseline

    The reversal is the same deployment's own later judgment: the organization said cost had become too predominant an evaluation factor and the result was lower quality, and committed to always keeping a human available. Nothing here says the deflection numbers were false - only that deflection was the wrong thing to have maximized, because the metric maximized (deflection, on cost) was not the one that mattered (resolution and quality). Both the numbers and the walk-back are drawn as claims the deployment made about itself.

  • assumed

    The two governable absences are drawn as the latent checks. Deflection is not resolution: the resolution-and-quality measurement the deflection headline does not contain is the empty independent model check - the thing that would have surfaced the quality cost before a reversal did. And the escalation path to a human - the safety valve customers say they want and a deflection-maximizing design erodes, which the walk-back restored - is the empty oversight check. The survey backdrop (most customers would rather not meet AI in service; analysts expect many organizations to abandon workforce-reduction plans) is carried in the case file, not computed here.

  • assumed

    No customer outcome is modeled here. This Lab reads institutional propagation only, and the customers being served are boundary-only. The deflection figures, the self-reported satisfaction parity, the later quality reversal, and the survey backdrop live in the case file, and are never computed from anything in this diagram.

What this example does not show

  • No customer outcome is modeled. The Lab reads institutional propagation only; the customers being served are boundary-only, and the deflection figures, the self-reported satisfaction parity, the later quality reversal, and the survey backdrop live in the case file, never computed on this diagram.
  • The first-month figures are the organization's own self-report entered as such (not independently audited), and the reversal is the same organization's own later judgment; the diagram draws both as claims the deployment made about itself, with the resolution-quality measurement and the escalation path as two latent checks, not a computed harm.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • An organization published striking first-month numbers for its customer-facing AI assistant: it handled about two-thirds of customer-service chats (some 2.3 million conversations), was described as doing the equivalent work of about 700 full-time agents, cut average resolution time from about 11 minutes to under 2, was said to match human customer satisfaction, and was projected to improve profit by tens of millions. Every one of those figures was self-reported and not independently audited. Roughly a year later the same organization reversed course on quality grounds — its chief executive said cost had become too predominant an evaluation factor and the result was lower quality — and committed to always keeping a human available to customers who want one. This is the contact-centre domain's cleanest benefit-then-cost arc: the deflection numbers and the walk-back come from the same deployment.

    empirical
    • Vendor Klarna Bank AB (2024, February 27). Klarna AI assistant handles two-thirds of customer service chats in its first month (press release via PR Newswire). https://www.prnewswire.com/news-releases/klarna-ai-assistant-handles-two-thirds-of-customer-service-chats-in-its-first-month-302072740.html
    • Trade press Ivanova, I. (2025, May 9). Klarna plans to hire humans again, as new landmark survey reveals most AI projects fail to deliver. Fortune. https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/
  • The lesson the benefit-then-cost arc carries is that deflection is not resolution. A published deflection number reports how many contacts the AI handled, not whether it handled them well, and a figure that is impressive on cost can hide a quality cost that only shows up later — which is what the organization's own reversal described. The survey backdrop sharpens it: most customers say they would rather not meet AI in service and fear it makes reaching a human harder, and industry analysts expect a large share of organizations to abandon plans to reduce their customer-service workforce with AI. The governable reading is to measure resolution and repeat contact against deflection rather than counting deflection as a win by itself, and to protect the path to a human as the safety valve a deflection-maximizing design tends to erode.

    empirical
    • Trade press Ivanova, I. (2025, May 9). Klarna plans to hire humans again, as new landmark survey reveals most AI projects fail to deliver. Fortune. https://fortune.com/2025/05/09/klarna-ai-humans-return-on-investment/
    • Trade press Gartner, Inc. (2025, June 10). Gartner Predicts 50% of Organizations Will Abandon Plans to Reduce Customer Service Workforce Due to AI (poll of 163 service leaders). https://www.theregister.com/software/2025/06/11/half_of_firms_set_to_abandon_plans_to_ditch_customer_service/502135

Where this connects

Institutional pressures in this domain

  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).

All of them in context on the Customer service & contact-centre AI domain page.

Levers available here and the patterns behind them

Documented case histories