PAN Lab example
A commercial code assistant across three enterprises
Big lift for novices but slower for experts: a coding assistant
One coding assistant suggests code to an org's whole engineering team. Modeled on a pre-registered randomized rollout: a big lift for less-experienced developers, and - measured independently - about 19% slower for experts who believed themselves 20% faster. Two traps to watch: the engineer who accepts a suggestion is also its only reviewer, and the individual speed-up degrades org delivery unless the review and testing gates absorb the churn it adds.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Coding-assistant-class with the operator-verifier collapse network: 5 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
A heavy workload against limited capacity: the evidence of record is company-run randomized trials across thousands of employed developers, and the review burden scales with the suggestion stream rather than with the trial. The seniority split already drawn is what the evidence is actually about - a large lift for less-experienced developers against an independent finding that experienced developers on familiar code were slowed while feeling faster - so the two correction pathways stay asymmetric.
- baseline
This models the coding-assistant pattern documented in the case file - not a reconstruction of the actual deployment. Two structural traps define it: the engineer who accepts a suggestion is also its reviewer of record (operator-verifier collapse, drawn on the junior developer's baseline-1 correction edge), and the shared repository is written into and read as the corpus later engineers and assistants use (drawn on the repo-to-model and repo-to-operator edges).
- assumed
The two developer classes are drawn with the same edges but the case documents sharply different measured outcomes: a pooled +26% completed tasks concentrated among less-experienced developers, and an independent randomized study measuring experienced developers ~19% slower on familiar code while believing themselves ~20% faster. The Lab does not compute that distribution; it is a recorded external observation carried in the case file, drawn here only as the heterogeneity a uniform service term overstates.
- assumed
The independence check is drawn present, at a low level, because the anchor evidence is a pre-registered, peer-reviewed randomized experiment - stronger than the self-reports elsewhere in this domain - though several authors are vendor-affiliated and the firms ran the experiments themselves; the pre-registration and peer-reviewed venue are the checks that make it credible. No security evaluation is in this org's own record; the class-level security risk is carried by the domain's gated-adoption case.
- baseline
The gates check is drawn empty because the org-composition bound is the domain's defining risk: a cross-industry program measured a ~7.2% decrease in delivery stability per 25% increase in adoption, so an individual speed-up degrades org delivery unless code-review and testing gates absorb the churn. Resourcing that gate is what makes the individual gain compose to the organization's outcome.
- assumed
No product or software outcome is modeled here. This Lab reads institutional propagation only, and the downstream users of the software are boundary-only. The productivity numbers, the expert slowdown and perception gap, and the delivery-stability bound live in the case file, and are never computed from anything in this diagram.
What this example does not show
- No product or software outcome is modeled. The Lab reads institutional propagation only; the downstream users of the software are boundary-only, and the productivity numbers, the expert slowdown, and the delivery-stability bound live in the case file, never computed on this diagram.
- The +26% and the ~19%-slower figures are recorded external measurements from separate studies (a pre-registered vendor-affiliated field experiment and an independent expert randomized controlled trial (RCT)); the diagram draws the two developer groups with identical structure and never computes the benefit or its distribution.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Company-run, pre-registered, peer-reviewed randomized rollouts of a commercial code-completion assistant across 4,867 developers at three enterprises found a pooled 26.08 percent increase in completed tasks, with gains concentrated among less-experienced developers. An independent randomized study of 16 experienced open-source maintainers on 246 tasks in familiar repositories bounded the expert tail from the other direction: those developers were about 19 percent slower with the AI while believing themselves about 20 percent faster — a measured perception-reality gap that means a uniform productivity number overstates the effect for senior engineers.
empirical- Academic Cui, Z.K., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2025). The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science. https://doi.org/10.1287/mnsc.2025.00535 https://pubsonline.informs.org/doi/10.1287/mnsc.2025.00535
- Industry Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR. https://doi.org/10.48550/arXiv.2507.09089 https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
Individual coding-assistant gains do not automatically compose to organization-level delivery outcomes: a cross-industry research program measured a roughly 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability for every 25 percent increase in AI adoption, evidence that the churn the assistant adds must be absorbed by code-review and testing gates or the individual speed-up degrades the organization's delivery performance.
empirical- Reference Google Cloud DORA (2024). Accelerate State of DevOps Report 2024. https://dora.dev/research/2024/dora-report/
Where this connects
Institutional pressures in this domain
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Software engineering AI (coding assistants) domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Keep prompts neutral — Framing and mirroring reduction
- Peer sharing rules — Peer-edge governance
- Gate record entries — Human-in-the-loop write gating
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Escalate checks — State-feedback vigilance
- Review the riskiest first — Risk-tiered oversight
- Store less data — Data minimization
- Upgrade model — Improve the model