Skip to content

PAN Lab example

GOV.UK Chat

The gate that said not yet: a public assistant behind a staged pilot gate

This assistant answers the public's tax, benefits and visa questions directly, in an official voice, drawing only on curated GOV.UK guidance — and the thing worth studying is not the answer but the gate in front of it. Modeled on GOV.UK Chat. An earlier 2023 version was held back for not reaching the accuracy demanded of a site like this, and the whole learning curve was published, gate by gate, alongside a rising accuracy figure. That makes it the near-opposite of the public chatbot that was told it was wrong, in public, and kept answering anyway: here the gate could say 'not yet', and did. The catch is quieter. Nearly every number the gate turns on — the accuracy, the answer rate, the blocked misuse attempts — is graded by the operator itself, with the denominators unpublished and the independent check aimed at security, not accuracy. Nothing is decided here; the answer reaches the public directly, and the operator itself named the risk that a fluent reply invites misplaced confidence — an answer taken as settled without the click back to the source. The question is not whether the gate exists — it does — but whether anyone outside reads the dial it turns on.

Stylized model of a documented deploymentBenefits navigation & public-facing chat

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the GOV.UK-Chat-class staged-gate public assistant network: 6 components and 14 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the staged-pilot-gate public-assistant pattern documented in the GOV.UK Chat case file — not a reconstruction of the actual product or any of its releases. It is deliberately the governance-inverse of the library's other public-facing generative adviser: the NYC MyCity-class chatbot is the same shape (a public asking directly, no professional intermediary, error in an official voice) but kept a documented-wrong adviser online and rejected its audit, while here the gate could say 'not yet' and did. It is also distinct from the verify-before-use benefits copilots (Nava-class, Benefit-Navigator-class), which put a caseworker between the model and the affected person; here the answer reaches the public directly.

  • baseline

    The staged pilot gate is drawn as a present oversight pathway, not an inert badge: each release was scored against a published accuracy bar before wider use, and the 2023 version was held back for not reaching the accuracy demanded for a site like GOV.UK, including a few cases of hallucination. Whether a given gate passed or held is recorded in the case file, not computed here. This is the case's payload — a public generative adviser whose gate cadence actually fired — and it is why the model-to-gate pathway runs at its normal level while the deference channel to the public is the thing the gate is meant to hold.

  • baseline

    This shape's honest defining absence is drawn as a latent check: no independent audit of the accuracy methodology gates a release. Nearly every reported figure (the 76 percent rising to 90 percent accuracy, the 88 percent answer rate, the 508 blocked jailbreak attempts) is self-reported by GDS, the system's operator; the accuracy denominators and sampling frames are unpublished, and the independent evaluation in the record is a security (jailbreak) assessment with the AI Security Institute, not an accuracy audit. The independent model check therefore starts closed — the independent look at the number the gate turns on that the record does not have — and a cross-model check opens it.

  • baseline

    The store coupling is deliberately protective, the opposite of the copilots. The model does not write generated answers into the curated official corpus, so the contamination pathway is drawn closed at baseline; the improve-the-content loop runs the benign direction, since observed answer failures pressure departmental teams to fix the underlying official guidance (Chat can only be as good as the content published on GOV.UK); and every answer routes the user back to the authoritative source page rather than substituting for it. These are the store-to-operator route-back at baseline, the evaluation-team-to-store content loop at baseline, and the model-to-store write held latent.

  • assumed

    The model's raw error rate is a modeling choice, not a measured per-interaction rate. The reported accuracy trajectory of 76 to 90 percent is a self-reported outcome statistic assessed by subject-matter experts plus automated metrics over pilot samples, and it is held below the ungated public-adviser family because the answers are grounded exclusively in a curated official corpus, the model is instructed to ignore its training data, and the whole deployment is staged behind an accuracy gate.

  • assumed

    Peer pathways are authored on both signs across the evaluation team — scoring conventions and red-team practices spread desk to desk, and colleagues cross-check each other's scores — and one public-facing model answering everyone identically is a monoculture in the Lab's qualitative vocabulary: a wrong answer repeats for every person who asks the same question rather than scattering. These are modeling assumptions, not measurements.

  • assumed

    The members of the public who use GOV.UK Chat are the operator network in this diagram, as an adoption channel only; no benefit, harm or downstream outcome to any person is computed from anything here, and there is no per-topic (for example benefits-specific) error rate or complaint data in the record. GDS states the system does not attempt to provide advice and signposts the original guidance; whatever an answer means for the person who acts on it is documented in the case file and measured outside any diagram like this one. This Lab models institutional propagation only.

What this example does not show

  • Nearly every quantitative figure in this case (the 76 to 90 percent accuracy, the 88 percent in-scope answer rate, the 508 blocked jailbreak attempts, the satisfaction rates) is self-reported by GDS, the system's operator; there is no independent audit of the accuracy methodology, the accuracy denominators and sampling frames are unpublished, and the claim that it outscores consumer AI assistants on government questions is the operator's own comparison. Treat these as operator claims, not independently verified measurements.
  • GOV.UK Chat is an information-provision system with no automated decision: it makes no eligibility, entitlement or sanction determination, so appeal and override constructs from the decision-system cases do not apply. The people who use it are modeled here only as an adoption channel; there is no per-topic (for example benefits-specific) error rate, and no published downstream-harm, complaint or user-outcome data from the roughly one in ten answers that were not accurate at the measured gates. Whatever an answer means for the person who acts on it is documented in the case file and measured outside any diagram like this one.
  • The phase labels in the record are inconsistent across documents ('experiment', 'private beta', 'pilot'), and per-phase user and question volumes should be attributed to their specific source: the transparency record describes a private beta capped at 2,000 users over 4 weeks, while the pilot-completion post reports 10,136 users in the first public pilot — these describe different phases and configurations. The all-508-jailbreaks-blocked claim sits alongside the record's own caveat that success can never be guaranteed.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The UK Government Digital Service ran what it called the government's biggest public test of generative AI to date: across two gated public pilots (a late-2024 web pilot of 10,136 users asking 23,838 questions, and an autumn-2025 GOV.UK app pilot of 641 users asking 2,670 questions in four weeks), more than 10,000 people asked GOV.UK Chat about 26,000 questions on tax, benefits, and visas. Its first 2023 version was held back in findings published January 18, 2024 because, GDS reported, answers did not reach the highest level of accuracy demanded for a site like GOV.UK, including a few cases of hallucination. GDS reports measured answer accuracy rising from 76 percent (its earliest benchmark) to 90 percent by the autumn 2025 pilot, assessed by subject-matter experts plus automated evaluation, an 88 percent answer rate for in-scope questions after a clarifying-questions feature was added, and that 508 attempts to jailbreak the system across the pilots were all prevented by its guardrails; it soft-launched to all GOV.UK app users on March 26, 2026 and officially launched on May 14, 2026. Nearly every one of these figures is self-reported by GDS, the system's operator, and the accuracy denominators and sampling frames are unpublished.

    empirical
    • Government evaluation Government Digital Service (Inside GOV.UK), 5 things we learned testing GOV.UK Chat: an AI assistant for government (2026) https://insidegovuk.blog.gov.uk/2026/03/16/5-things-we-learned-testing-gov-uk-chat-an-ai-assistant-for-government/
    • Government evaluation Government Digital Service (Inside GOV.UK), The findings of our first generative AI experiment: GOV.UK Chat (2024) https://insidegovuk.blog.gov.uk/2024/01/18/the-findings-of-our-first-generative-ai-experiment-gov-uk-chat/
    • Government Government Digital Service, Answers in seconds, 24/7: GOV.UK Chat launches in the GOV.UK app (2026) https://gds.blog.gov.uk/2026/05/14/gov-uk-chat-launches/
  • GOV.UK Chat is a retrieval-augmented assistant that, per its Algorithmic Transparency Record published October 7, 2025, answers only from roughly 700,000 vectorised chunks (36.9 GB) of curated official GOV.UK guidance, is instructed to ignore its training data, rejects questions containing phone numbers, emails, or card numbers, links every answer back to its GOV.UK source pages with a reminder to verify, and retains question data encrypted for 12 months; GDS states it does not attempt to provide advice and makes no automated decision. GDS's December 2025 vision post frames a content-dependency loop, stating that GOV.UK Chat can only be as good as the content published on GOV.UK by departmental teams. The record's independent evaluation is a jailbreak (security) assessment conducted with the AI Security Institute, alongside the record's own caveat that it is not possible to guarantee no jailbreaking attempts will succeed; there is no independent audit of the accuracy methodology, and GDS's claim that for government-related questions the tool scores higher than widely-used consumer AI assistants is the operator's own comparison.

    empirical
    • Government Department for Science, Innovation and Technology (GOV.UK Algorithmic Transparency Recording Standard), GOV.UK Chat Algorithmic Transparency Record (2025) https://www.gov.uk/algorithmic-transparency-records/dsit-gov-dot-uk-chat
    • Government Government Digital Service (Inside GOV.UK), GOV.UK has entered the Chat: our vision for GOV.UK Chat (2025) https://insidegovuk.blog.gov.uk/2025/12/16/gov-uk-has-entered-the-chat-our-vision-for-gov-uk-chat/
    • Trade press Civil Service World (Jim Dunton), GOV.UK AI chatbot achieves 90% accuracy (2026) https://www.civilserviceworld.com/professions/article/govuk-ai-chatbot-achieves-90-accuracy

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Benefits navigation & public-facing chat domain page.

Levers available here and the patterns behind them

Documented case histories