Skip to content

PAN Lab example

Caddy adviser copilot at Citizens Advice

The gate is a job rather than a habit: a supervisor-checked adviser copilot

For once, a system built to stay safe. This copilot never talks to the public: every draft it writes goes to a separate supervisor to approve, edit, or reject before it reaches even the adviser, and only then does a human relay the answer to the client. Modeled on Caddy. Verification here is a whole job, not a habit someone under pressure might skip — which is exactly its strength. The question flips: as the pressure rises and the tool scales toward a whole national network, does the institution keep making every message pass the gate?

Stylized model of a documented deploymentBenefits navigation & public-facing chat

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Caddy-class supervisor-gated adviser copilot network: 5 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the supervisor-gated, adviser-facing verify-before-release copilot pattern documented in the Caddy case file — not a reconstruction of the actual tool. It is a deliberate sibling of the library's two US benefits copilots (the Nava-class and Benefit-Navigator-class assistive chatbots) and distinct from them: those place verify-before-use inside the same caseworker who uses the answer, while Caddy routes every draft to a separate reviewing supervisor before the adviser sees it, so the safety case is a role rather than a habit.

  • baseline

    The load-bearing control is a separate reviewing supervisor, not a personal habit: every draft is approved, edited, or rejected before it reaches the adviser, with about one draft in five changed or rejected (a developer-reported randomised-trial outcome contrast, not a per-interaction rate). The supervisor-to-adviser check is drawn strong at baseline — one step above the US siblings' collective habit — to carry that.

  • assumed

    The retrieval corpus is a closed two-source set — the government guidance site and the charity's internal knowledge base — that the copilot reads but does not write to, so record contamination is low by design at baseline; contamination enters only if a supervisor approves a subtly wrong draft or an adviser relays one, and the corpus's only write path is slow, negotiated coverage extension.

  • baseline

    Peer pathways are authored on both signs: approved answers relay worker to worker and the copilot's confidence spreads across the office (advisers were more than twice as likely to report confidence and trainees more comfortable asking the tool than a human supervisor), while a weaker collective flag-each-other habit and the per-claim citation check run as inhibiting pathways.

  • assumed

    The defining latent absence is a dormant direct model-to-adviser channel: at baseline nothing reaches the adviser unchecked, but the documented pressure to 'streamline the supervisor checks' (the 2.0 verification engine) and to route responses 'for human checks when needed' (the 2026 programme page) is exactly what would open a bypass of the every-message gate as the tool scales toward the whole network. Baseline 0; connection-auth caps it, gate-erosion pressure opens it.

  • assumed

    Served people are not in the dynamics; the harm mode on this positive pole is a wrong answer that clears the gate and is relayed to a client, recorded outside a diagram like this one, never a punitive flag. Every headline figure here is developer-reported and not independently replicated, and no differential client harm is computed.

What this example does not show

  • Served people, and the advice they do or do not ultimately receive, are not modeled here; the Lab reads institutional propagation only, and those outcomes are documented in the case file and measured outside any diagram like this one.
  • Every headline figure — the 1,000-plus randomised requests, the roughly 80% of drafts approved without revision, the about-four-minute turnaround, the doubled adviser confidence — is reported by the tool's builders in blog and conference form rather than independently replicated; the confidence and resolution numbers are adviser self-reports, not client-outcome or accuracy measures; no absolute post-supervision error rate is published; and one independent account describes the evaluation as a short pilot rather than detailing the randomisation.
  • A safe starting baseline is a property of this model, not a safety promise for any real deployment; the every-message supervisor gate that makes this shape safe is, on the case's own record, already under pressure to relax as the tool scales.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • In a developer-reported randomised controlled trial of more than 1,000 adviser support requests, an adviser-facing benefits copilot at Citizens Advice returned supervisor-checked answers in about four minutes, roughly half the previous response time, with about 80 percent of its drafts approved by supervisors without revision; these figures are reported by the tool's builders and have not been independently replicated.

    empirical
    • Government Varotsis, Transforming Civic Engagement with Caddy (Incubator for Artificial Intelligence, i.AI, UK Government, developer blog, 2025) https://ai.gov.uk/blogs/transforming-civic-engagement-with-caddy/
    • Government Department for Science, Innovation and Technology / i.AI / CASORT, Caddy (AI Knowledge Hub use case, 2025) https://ai.gov.uk/knowledge-hub/use-cases/caddy/
    • Academic Stanford Legal Design Lab / Justice Innovation, How AI is augmenting human-led legal advice at Citizens Advice (Caddy adviser copilot) https://justiceinnovation.law.stanford.edu/how-ai-is-augmenting-human-led-legal-advice-at-citizens-advice/
  • Advisers given access to the copilot were reported to be more than twice as likely to say they felt confident giving advice than a control group, a self-reported measure from post-call in-chat surveys rather than a client-outcome or accuracy measure.

    empirical
    • Government Varotsis, Transforming Civic Engagement with Caddy (Incubator for Artificial Intelligence, i.AI, UK Government, developer blog, 2025) https://ai.gov.uk/blogs/transforming-civic-engagement-with-caddy/
    • Academic Stanford Legal Design Lab, Caddy Q and A copilot (JusticeBench project page, 2025) https://www.justicebench.org/project/caddy

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Benefits navigation & public-facing chat domain page.

Levers available here and the patterns behind them

Documented case histories