Skip to content

PAN Lab example

GDS Microsoft 365 Copilot cross-government experiment

Time saved then spent checking: a whole-office productivity copilot

This copilot drafts documents, summarises meetings, and writes email for the whole office at once, and the largest published trial of it saved a self-reported 26 minutes a day and left most users unwilling to go back. Modeled on the GDS Microsoft 365 Copilot cross-government experiment. The catch is that the trial measured adoption and time, not output quality: a companion department found no robust evidence the saved time improved productivity, and about a fifth of its users caught the tool hallucinating. The minutes saved are the minutes not spent verifying, and the tool is trusted most on grievances and evaluations — exactly where a wrong line is hardest to catch. Its drafts flow into shared documents and minutes the same tool later re-summarises, and it could not say which document an answer came from. Nothing decides anything here; the record just quietly fills with fluent text no one independently checked.

Stylized model of a documented deploymentCaseworker documentation & copilots

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Cross-government-copilot-class horizontal productivity layer network: 5 components and 16 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the horizontal-productivity-copilot pattern documented in the GDS Microsoft 365 Copilot cross-government experiment case file — not a reconstruction of the actual product or any department's deployment. It is deliberately distinct from the library's vertical copilots: the Magic-Notes-class documentation tool rests its whole safety case on one human review gate on the error-to-record pathway, and the Nava-class and Benefit-Navigator-class tools are verify-before-use, with a caseworker reading a cited answer. Here a single copilot is inserted across every group of staff at once and writes generated text broadly into shared memory it later re-summarises.

  • baseline

    The two frontline task-classes encode the evaluation's central finding: the benefit concentrated in high-verifiability tasks (drafting, transcription) while the reported accuracy concerns concentrated in low-verifiability, high-stakes tasks (grievance handling, performance evaluations) where a wrong line is least detectable by a busy reviewer. The deference channel is drawn equally strong into both, with the grievance-and-evaluation class labeled as the place a wrong line is hardest to catch; the accuracy concerns themselves are qualitative focus-group findings about perceived risk, not measured error rates.

  • baseline

    The write-back-and-re-summarise loop is this shape's distinctive feature: unlike the verify-before-use copilots with minimal write-back, this horizontal layer writes generated text broadly into shared documents, minutes, and mail, which the same tool later retrieves and re-summarises. The documented provenance failure — the tool could not identify which documents generated a response — means a contaminated line is hard to trace back once written, so the model-to-store write and the store-to-model re-summarise edges both run at baseline.

  • baseline

    This shape's defining feature is an absence: the model-side inhibiting check is drawn empty at baseline, because the experiment measured adoption and self-reported time, not per-output quality, so no output-level quality-assurance layer gated a generated draft against ground truth before it entered the shared record. The only per-output control is the individual reviewer, and per-output verification is exactly what the reported time savings consume — worst on the low-verifiability tasks where errors are least detectable.

  • assumed

    The model's raw error rate is a modeling choice, not a measured per-interaction rate. Every reported time-savings figure is self-reported and unverified, there is no audited error-incidence rate, and the 22 percent user-detected-hallucination figure from the companion evaluation is detection by users, not measured incidence — it sits alongside 43 percent who detected none, 11 percent unsure, and a further roughly one in five who did not answer.

  • assumed

    Peer pathways are authored on both signs: prompt shortcuts and copy-paste habits spread desk to desk, colleagues sanity-check each other's drafts when time allows, and one horizontal copilot drafting for the whole office homogenises the record's voice toward its phrasing — a monoculture assumption in the Lab's qualitative vocabulary, not a measurement.

  • assumed

    Served members of the public and whether the saved time improved services are not in the dynamics. The companion evaluation found no robust evidence that task-level time savings produced department-level productivity gains, and whatever the civil servants' outputs mean for the people they serve is documented in the case file and measured outside any diagram like this one; this Lab models institutional propagation only.

What this example does not show

  • Every reported time-savings figure is self-reported and unverified; the findings report says they should be considered alongside additional literature on the true value, and independent coverage recomputed the 13-day-per-year extrapolation to about 4.6 days on a 253-working-day basis. A companion departmental evaluation found no robust evidence that the task-level time savings produced department-level productivity gains.
  • No audited output-quality or error-incidence measurement exists. The 22% user-detected-hallucination figure is detection by users, not a measured error rate, and sits alongside 43% who detected none, 11% unsure, and a further roughly one in five who did not answer; the grievance-handling and performance-evaluation accuracy concerns are qualitative focus-group findings about perceived risk. In one observed-task exercise the copilot's users completed spreadsheet analysis more slowly and to worse quality, and produced presentation slides over 7 minutes faster but to worse quality and accuracy that then needed correction.
  • The members of the public these civil servants serve, and whether the saved time improved the services they receive, are not modeled here; the Lab models institutional propagation only, and those outcomes are documented in the case file and measured outside any diagram like this one.
  • The HMRC 50,000-licence and roughly 50-million-pound-a-year figures are analyst projections in the evaluation (up to 50,000 licences until March 2028), not a funded deployment, and the roughly 28,000-licence rollout rests on independent press reporting rather than the evaluations; per-licence cost figures cited publicly are commercial reference prices, not the government's negotiated licence cost.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The Government Digital Service ran a cross-government experiment with Microsoft 365 Copilot from September 30 to December 31, 2024, with about 20,000 employees across 12 organisations, and published the findings report on June 2, 2025. Participants self-reported saving an average of about 26 minutes per working day (the report extrapolates this to roughly 13 days a year from the median values of six reported time-savings ranges; independent coverage recomputed it to about 4.6 days on a 253-working-day basis), 17% reported no clear savings, adoption held near 80% after peaking at about 83%, and 82% said they would not want to return to working without it. The experiment measured adoption and self-reported time rather than output quality: the report recorded no audited error rate, flagged significant accuracy concern for low-verifiability tasks such as grievance handling and performance evaluations, noted external web data was used without built-in verification, and documented a provenance failure in which the tool struggled to identify which documents generated a response.

    empirical
    • Government evaluation Government Digital Service (DSIT), Microsoft 365 Copilot Experiment Cross-Government Findings Report (HTML) (2025) https://www.gov.uk/government/publications/microsoft-365-copilot-experiment-cross-government-findings-report/microsoft-365-copilot-experiment-cross-government-findings-report-html
    • Government Government Digital Service, Microsoft 365 Copilot Experiment Cross-Government Findings Report publication page (2025) https://www.gov.uk/government/publications/microsoft-365-copilot-experiment-cross-government-findings-report
    • Trade press The Register (Thomas Claburn), UK govt study Copilot AI saved workers 26 minutes a day (2025) https://www.theregister.com/2025/06/03/uk_government_study_ai_time_savings/
  • A companion Department for Business and Trade evaluation of Microsoft 365 Copilot (1,000 licences, October to December 2024; published August 28, 2025) reported 72% user satisfaction but concluded it did not find robust evidence that time savings were leading to improved productivity; in observed tasks its users completed spreadsheet data analysis more slowly and to worse quality and accuracy than non-users, and produced presentation slides over 7 minutes faster on average but to worse quality and accuracy that then needed correction. In its diary study, 22% of respondents said they had identified hallucinations, 43% detected none, and 11% were unsure, with a further roughly one in five not answering, so the figure reflects user-detected hallucination rather than audited incidence. A Department for Work and Pensions evaluation (3,549 licences; published January 29, 2026) measured 19 minutes a day saved across eight routine tasks against a comparison group (95% confidence interval 17 to 22 minutes), found 85% rating meeting-note accuracy good or very good, reported that users consistently reviewed outputs before use, and concluded the tool is complementary to human expertise and requires consistent human oversight.

    empirical
    • Government evaluation Department for Business and Trade, Microsoft 365 Copilot pilot DBT evaluation report (2025) https://www.gov.uk/government/publications/microsoft-365-copilot-pilot-dbt-evaluation-report
    • Trade press The Register (Paul Kunert), M365 Copilot fails to up productivity in UK government trial (2025) https://www.theregister.com/2025/09/04/m365_copilot_uk_government/
    • Government evaluation Department for Work and Pensions, An Evaluation of DWP Microsoft 365 Copilot Trial (2026) https://www.gov.uk/government/publications/an-evaluation-of-dwps-microsoft-copilot-365-trial/an-evaluation-of-dwps-microsoft-365-copilot-trial

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Caseworker documentation & copilots domain page.

Levers available here and the patterns behind them

Documented case histories