PAN Lab example
An ambient AI scribe at a multi-specialty health system
One number and two outcomes: an ambient scribe that split by clinician
One ambient scribe drafts for two clinician groups. Modeled on an evaluation that measured the distribution, not just the average: it helped 85.8% of primary-care physicians but only 36.4% of specialists, and the burnout change was not significant. So watch the thing a single benefit number hides - who the tool actually helps, and who carries the same review burden without the relief it promised.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Ambient-scribe-class with operator heterogeneity network: 5 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
The specialist correction pathway is drawn at full strength against the primary-care pathway at a substantial level because the evaluation found the benefit splitting by specialty - 85.8 percent of primary-care physicians reported improved satisfaction against 36.4 percent of medical specialists - so the specialist carries the same review burden on a fluent draft without the same relief. A heavy workload against limited capacity reflects that split rather than a uniform load. The evaluation function reads the note record, which is how the split was found at all; drawn as a real inbound pathway, its check stays empty because tracking the distribution on a schedule, rather than reporting one average, is the thing this deployment did not institutionalize.
- baseline
This models the operator-heterogeneity pattern documented in the Sutter case file - not a reconstruction of the actual system. The structural skeleton is the standard generation-into-the-record topology; what differs is that two groups of staff are fed by the same scribe (n_U:2) and the measured benefit split sharply between them, so a uniform service term overstates the effect for the group it helps least.
- assumed
The two operator classes are drawn with identical edges and baselines on purpose: the tool and the workflow are the same for both, and the documented difference (85.8% of primary-care physicians reported improved satisfaction against 36.4% of medical specialists) is a difference in measured outcome, not in the diagram's structure. The Lab does not compute that satisfaction split; it is a recorded external observation carried in the case file, and it is drawn here only as the heterogeneity the benefit-distribution check exists to surface.
- baseline
The null bound is the honest counterweight this org supplies to the family: note time and cognitive load were reduced, but the burnout change (42.1% to 35.1%) was not statistically significant. The strongest deployments' larger numbers should be read against this null, not in place of it. The benefit is a distribution, not a scalar.
- assumed
No care outcome is modeled here. This Lab reads institutional propagation only, and the patients whose visits are transcribed are boundary-only. The satisfaction split, the note-time figures, and the non-significant burnout change live in the case file, and are never computed from anything in this diagram.
What this example does not show
- No care outcome is modeled. The Lab reads institutional propagation only; the patients whose visits are transcribed are boundary-only, and the satisfaction split, note-time figures, and non-significant burnout change live in the case file, never computed on this diagram.
- The two user groups are drawn with identical structure on purpose: the documented 85.8%-vs-36.4% satisfaction split is a difference in measured outcome, a recorded external observation, not a value derived from the diagram, which never computes the benefit or its distribution.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A peer-reviewed evaluation of an ambient documentation platform at a large multi-specialty system found note time per appointment reduced (6.2 to 5.3 minutes) and NASA-TLX cognitive load reduced — but the burnout change (42.1 to 35.1 percent) was not statistically significant, the domain's honest null bound of cognitive-load relief without a demonstrated burnout effect.
empirical- Peer-reviewed Stults, C.D., Deng, S., Martinez, M.C., et al. (2025). Evaluation of an Ambient Artificial Intelligence Documentation Platform for Clinicians. JAMA Network Open, 8(5), e258614. https://doi.org/10.1001/jamanetworkopen.2025.8614 https://pubmed.ncbi.nlm.nih.gov/40314951/
In the same evaluation, benefit varied sharply by clinician group: 85.8 percent of primary-care physicians reported improved satisfaction against 36.4 percent of medical specialists — the same tool, in the same system, under the same workflow, helping one operator class and largely failing another, so any uniform service term overstates the effect for the group it helps least.
empirical- Peer-reviewed Stults, C.D., Deng, S., Martinez, M.C., et al. (2025). Evaluation of an Ambient Artificial Intelligence Documentation Platform for Clinicians. JAMA Network Open, 8(5), e258614. https://doi.org/10.1001/jamanetworkopen.2025.8614 https://pubmed.ncbi.nlm.nih.gov/40314951/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
All of them in context on the Clinical documentation copilots (ambient scribes) domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Keep prompts neutral — Framing and mirroring reduction
- Peer sharing rules — Peer-edge governance
- Gate record entries — Human-in-the-loop write gating
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Escalate checks — State-feedback vigilance
- Review the riskiest first — Risk-tiered oversight
- Store less data — Data minimization
- Upgrade model — Improve the model