Skip to content

PAN Lab example

A contact centre's generative-AI agent assist

Fifteen percent on average and thirty for the newcomers

One agent-assist copilot drafts responses for two agent groups. Modeled on a randomized rollout that measured the distribution, not just the average: about 15% more issues resolved per hour overall, but novices gained 30-34% while the most experienced gained close to zero, with some evidence of slight quality degradation for them. The benefit is real and it is a distribution. So watch the thing a single number hides - who the copilot actually helps, and the small quality cost at the top the average flatters away.

Stylized model of a documented deploymentCustomer service & contact-centre AI

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Agent-assist-copilot-class with skill compression network: 6 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    The case describes the copilot as two things - suggested responses and knowledge retrieval - so the retrieval half is drawn as its own component. That half is where the finding lives: looking up the right article or the right prior case is most of what separates an agent who does not yet know where to look from one who does, which is why the same tool moved a novice about a third and an experienced agent close to nothing. A heavy workload against limited capacity for roughly five thousand agents on live chat, and capacity is held above the floor because the same study measured agent retention improving under this deployment - a resourced rollout, not a strained one. The quality-assurance (QA) function is wired to the record it evaluates.

  • baseline

    This models the skill-compression pattern documented in the case file - not a reconstruction of the actual deployment. The benefit is real and randomized-measured: a staggered rollout to roughly 5,000 support agents raised issues resolved per hour about 15% on average against a control group, and improved customer sentiment and agent retention. The structural skeleton is the operator-heterogeneity topology, because the point is that the benefit is a distribution, not a scalar.

  • assumed

    The two agent classes are drawn with identical edges and baselines on purpose: the copilot and the workflow are the same for both, and the documented difference (novice/low-skill agents ~+30-34%, the most experienced ~0 with some evidence of slight quality degradation) is a difference in measured outcome, not in the diagram's structure. The Lab does not compute that split; it is a recorded external observation carried in the case file, and it is drawn here only as the heterogeneity the benefit-distribution check exists to surface.

  • baseline

    The governable reading is that an agent-assist copilot mostly raises the floor, so an average productivity number overstates the effect for the experienced agents who least need it and hides the slight quality cost the same evidence flags for them. The benefit-distribution measurement is drawn as the latent independent model check, and the standing quality-assurance cadence over the quality tail as the latent oversight check - the two functions that turn a single headline number into a governed distribution.

  • assumed

    No customer outcome is modeled here. This Lab reads institutional propagation only, and the customers being served are boundary-only. The productivity figures, the skill-compression finding, and the slight quality degradation at the top live in the case file, and are never computed from anything in this diagram.

What this example does not show

  • No customer outcome is modeled. The Lab reads institutional propagation only; the customers being served are boundary-only, and the productivity figures, the skill-compression finding, and the slight quality degradation at the top live in the case file, never computed on this diagram.
  • The ~15% average / ~30-34% novice / ~0 veteran figures are a peer-reviewed randomized field study's measured outcomes entered as such; the two agent classes are drawn with identical structure on purpose, because the split is a measured OUTCOME, not a structural difference the diagram computes.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The strongest field evidence for an agent-assist copilot in customer service comes from a staggered randomized rollout of a generative-AI assistant to roughly 5,000 customer-support agents at a large software firm. Measured against a control group, the copilot raised issues resolved per hour by about 15 percent on average, and it also improved customer sentiment and agent retention. The gain, however, was sharply uneven: novice and low-skill agents improved by roughly 30 to 34 percent, agents with two months of experience performed like agents with six months and no AI, and the most experienced agents gained close to nothing, with some evidence of slight quality degradation. This is the contact-centre domain's cleanest measured benefit, and it is a distribution rather than a single number.

    empirical
    • Peer-reviewed Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2), 889-942. https://doi.org/10.1093/qje/qjae044 https://academic.oup.com/qje/article/140/2/889/7990658
    • Academic Brynjolfsson, E., Li, D., & Raymond, L. (2023). Generative AI at Work. NBER Working Paper 31161 https://www.nber.org/papers/w31161
  • The lesson the randomized evidence carries is skill compression: an agent-assist copilot mostly raises the floor. Because almost the entire measured gain accrues to less-experienced agents and the most experienced gain close to nothing, an average productivity number overstates the effect for the agents who least need it and hides that the tool does little for the experienced while possibly costing a small amount of quality there. The governable reading is that the benefit must be measured as a distribution across agent skill, not reported as a scalar — a copilot that helps novices a great deal and experts not at all is a real and specific benefit, and describing it with one average misstates who it helps.

    empirical
    • Peer-reviewed Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2), 889-942. https://doi.org/10.1093/qje/qjae044 https://academic.oup.com/qje/article/140/2/889/7990658

Where this connects

Institutional pressures in this domain

  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).

All of them in context on the Customer service & contact-centre AI domain page.

Levers available here and the patterns behind them

Documented case histories