Skip to content

PAN Lab example

Ambient scribe RCT + monitoring playbook

The trial and the playbook: an evidenced and monitored ambient scribe

An ambient scribe drafts the note; the clinician edits and signs. Modeled on the family's strongest-evidenced deployment: a randomized trial found real benefit - exhaustion down, ~22 minutes a day returned - and its monitoring exists as a published playbook, not a promise. So watch two rarer things than the benefit: what the trial did and did not measure, and whether the monitoring stays pointed at the coding-arms-race risk the better documentation creates.

Stylized model of a documented deploymentClinical documentation copilots (ambient scribes)

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Ambient-scribe-class with a published monitoring playbook network: 4 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    The oversight here watches the model in production, not the finished note: this deployment published an operations playbook for monitoring ambient AI in production, so the monitoring function takes the draft stream directly and its check runs faintly. That is the structural difference from the quality-assurance-over-the-record sibling, and it is documented rather than drawn for variety. Capacity is drawn above demand because the evidence of record is a 24-week stepped-wedge randomized trial across 66 practitioners - a small, operationally supported population, the opposite of a scaled rollout. The independent check runs faintly for the same reason: a pragmatic randomized controlled trial (RCT) is the strongest causal evidence in this domain.

  • baseline

    This models the ambient-scribe pattern on the same generation-into-the-record topology as the well-governed pole - not a reconstruction of the actual system. This org's distinctives are evidence and monitoring, not shape: a stepped-wedge individually randomized trial (66 practitioners, 71,487 notes) and a published production-monitoring playbook.

  • baseline

    The independence check is drawn present, at a low level, because the trial is a causal, randomized estimate of the well-being and time benefit - exhaustion significantly reduced, professional fulfillment a recorded null, ~22 minutes/day returned - the family's strongest causal evidence, and an honest one for recording the null. The trial measured well-being and time, not per-note error rates; the class-level ~31% hallucination profile is carried by the domain's other cases, not re-claimed here.

  • assumed

    The monitoring check is drawn empty for the player to strengthen, but the assumption records its real-world distinctive: here the org-side monitoring function exists as a published operations playbook rather than an assumed practice - the rare case where oversight is a citable, inspectable subsystem. It must stay pointed at the documented coding-arms-race risk: better AI documentation raises coding intensity, payers recalibrate, and clinician attestation liability grows.

  • assumed

    No care outcome is modeled here. This Lab reads institutional propagation only, and the patients whose visits are transcribed are boundary-only. The randomized-trial estimates, the monitoring playbook, and the coding-arms-race risk live in the case file, and are never computed from anything in this diagram.

What this example does not show

  • No care outcome is modeled. The Lab reads institutional propagation only; the patients whose visits are transcribed are boundary-only, and the trial estimates, the monitoring playbook, and the coding-arms-race risk live in the case file, never computed on this diagram.
  • The randomized trial measured practitioner well-being and time, not per-note error rates; the ~31% ambient-note hallucination profile is a separate product-class evaluation carried by the domain's other cases, not a finding of this trial.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • The strongest causal evidence in the ambient-scribe family comes from a 24-week stepped-wedge, individually randomized trial of an ambient scribe across 66 practitioners and 71,487 notes (38 percent AI-generated), which found work exhaustion significantly reduced, professional fulfillment unchanged (a recorded null), roughly 22 minutes per day of documentation time returned, and diagnostic coding accuracy improved. Unlike the larger first-party deployment reports, this is a randomized estimate of the well-being and time effects — though it measured practitioner well-being and time, not per-note error rates.

    empirical
    • Peer-reviewed Afshar, M., Baumann, M.R., Resnik, F., et al. (2025). A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being. NEJM AI. https://doi.org/10.1056/AIoa2500945 https://pubmed.ncbi.nlm.nih.gov/41625485/
  • The same team that ran the ambient-scribe trial released an open operations playbook for safety and effectiveness monitoring of ambient AI in production — the rare case where the organization-side monitoring function exists as a citable, designed subsystem rather than an assumed practice. That monitoring is the org's stated answer to a documented system-level risk of the technology: a coding arms race, in which better AI documentation raises coding intensity, payers recalibrate in response, and clinician attestation liability grows — so the improved coding accuracy the trial measured sits next door to an upcoding pressure the monitoring is meant to watch.

    empirical
    • Academic Afshar, M., et al. (2025). A Novel Playbook for Pragmatic Trial Operations to Monitor and Evaluate Ambient Artificial Intelligence in Clinical Practice. NEJM AI. https://doi.org/10.1056/AIdbp2401267 https://ai.nejm.org/doi/full/10.1056/AIdbp2401267
    • Peer-reviewed Dai, T., Kvedar, J.C., & Polsky, D. (2025). Policy brief: ambient AI scribes and the coding arms race. npj Digital Medicine. https://doi.org/10.1038/s41746-025-02272-z https://www.nature.com/articles/s41746-025-02272-z

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.

All of them in context on the Clinical documentation copilots (ambient scribes) domain page.

Levers available here and the patterns behind them

Documented case histories