PAN Lab example
Kaiser Permanente ambient AI scribe
The draft becomes the record: a well-governed ambient scribe
An ambient AI drafts the clinical note; the clinician edits and signs. Modeled on the largest documented scribe deployment. The benefit is real - thousands of hours of documentation returned. But the draft does not stay a draft: it becomes a permanent record later clinicians and later tools read as fact. So watch the write into the record - because ambient notes are documented to hallucinate about a third of the time, and what survives the review is copied forward as truth.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Ambient-scribe-class generation into the record network: 4 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This is, in the published record, the only ambient-scribe deployment running a standing internal quality-assurance (QA) function over AI output, and that is drawn as a real inbound pathway: the QA program samples signed notes out of the record, so its check runs faintly rather than empty. Demand is drawn heavy against limited capacity because the scale is a pilot that became 7,260 physicians and roughly 2.58 million encounters - every one of which is a draft somebody has to read - while the QA function samples rather than reviews exhaustively. The independent benefit check stays empty: the hours-saved figures here are the deployer's own first-party measurements, not an arm's-length evaluation.
- baseline
This models the generation-into-the-record pattern documented in the ambient-scribe case file - not a reconstruction of the actual system. Its defining feature is that the AI drafts a note written into a permanent record and read forward as fact, so the governable object is the write into the store, not the alert - the structural reason this domain is separate from alerting clinical decision support.
- baseline
The copied-forward-as-fact dynamic is drawn on the note-to-model feature edge and the note-to-clinician read edge: a signed note re-enters the scribe's next draft and is read by later clinicians (and downstream tools) as ground truth, so an uncaught AI error becomes inherited truth with a long half-life. Against a validated per-note hallucination rate near 31 percent, that is why the clinician review and the quality assurance (QA) reconciliation are contamination controls, not politeness - a review degraded into a rubber stamp is a contamination source.
- assumed
This is the well-governed pole: in the real deployment the write is gated twice, by the clinician's edit-and-sign review (drawn active, at a substantial level) and by a standing quality-assurance (QA) program that samples the AI output across the deployment - a real subsystem with a real cost. The QA reconciliation is drawn as the latent oversight check, empty at baseline, so the player strengthens it with the reconciliation and cadence levers; the assumption records that the governed deployment runs it as a standing function.
- baseline
The independence bound is drawn on the independent model check, empty at baseline: the deployment's benefit numbers are the deployer's own first-party measurements, and an independent multisite study of 8,581 clinicians found a more modest effect with no meaningful after-hours relief. The scale figures are this system's dashboard, not the product class's guarantee.
- assumed
No care outcome is modeled here. This Lab reads institutional propagation only, and the patients whose visits are transcribed are boundary-only. The time-saved figures, the hallucination rates, and the coding-intensity risk live in the case file, and are never computed from anything in this diagram.
What this example does not show
- No care outcome is modeled. The Lab reads institutional propagation only; the patients whose visits are transcribed are boundary-only, and the time-saved figures, hallucination rates, and coding-intensity risk live in the case file, never computed on this diagram.
- The deployment's benefit numbers are the deployer's own first-party, peer-reviewed measurements; independent multisite data shows a more modest effect, and the ~31% ambient-note hallucination rate is a product-class evaluation, not a measurement of this specific system.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
The largest documented ambient-scribe deployment ran a 10-week pilot at an integrated medical group and then scaled to 7,260 physicians and 2,576,627 patient encounters over fourteen months, with roughly 16,000 hours of documentation time saved and sustained physician support measured along the way. The system records the visit and drafts the clinical note; the clinician edits and signs, and the model-to-record write is gated both by that clinician review and by a standing internal quality-assurance program over the AI output — a real subsystem with a real cost, because the drafted note becomes a permanent record that later clinicians and later tools read as fact.
empirical- Academic Tierney, A.A., Gayre, G., Hoberman, B., et al. (2024). Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catalyst Innovations in Care Delivery. https://doi.org/10.1056/CAT.23.0404 https://catalyst.nejm.org/doi/full/10.1056/CAT.23.0404
- Academic Tierney, A.A., et al. (2025). Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses. NEJM Catalyst Innovations in Care Delivery. https://doi.org/10.1056/CAT.25.0040 https://divisionofresearch.kaiserpermanente.org/ai-assisted-notetaking-gains-steady-support-from-kaiser-permanente-physicians/
The scale numbers from a single ambient-scribe deployment are the deployer's own first-party measurements and should be read as that system's dashboard rather than a guarantee of the product class: a multisite study of 8,581 clinicians across five health systems found more modest effects — on the order of 13 to 16 fewer minutes per day with no meaningful after-hours relief — and a validated per-note evaluation found hallucinations in about 31 percent of ambient-generated notes under structured review, versus about 20 percent of physician-written gold-standard notes, making ambient notes more thorough but less accurate. The clinician review and quality-assurance program are the controls that stand between that error rate and a contaminated permanent record.
empirical- Peer-reviewed Rotenstein, L.S., et al. (2026). Changes in Clinician Time Expenditure and Visit Quantity With Adoption of Artificial Intelligence-Powered Scribes: A Multisite Study. JAMA. https://doi.org/10.1001/jama.2026.2253 https://pubmed.ncbi.nlm.nih.gov/41920565/
- Peer-reviewed Palm, K.H., Manikantan, K., Mahal, N., Belwadi, S.K., & Pepin, R.J. (2025). Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe. Frontiers in Artificial Intelligence, 8. https://doi.org/10.3389/frai.2025.1691499 https://pmc.ncbi.nlm.nih.gov/articles/PMC12586549/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
All of them in context on the Clinical documentation copilots (ambient scribes) domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Keep prompts neutral — Framing and mirroring reduction
- Peer sharing rules — Peer-edge governance
- Gate record entries — Human-in-the-loop write gating
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Escalate checks — State-feedback vigilance
- Review the riskiest first — Risk-tiered oversight
- Store less data — Data minimization
- Upgrade model — Improve the model