PAN Lab example
GitHub Copilot at ZoomInfo
Measured carefully but measuring the wrong thing: an ordinary rollout
A well-run four-phase rollout to 400+ developers, carefully documented: acceptance rates, satisfaction, per-language deltas, stated limitations. Modeled on an ordinary competent adoption. It does most things right - but its headline metric, acceptance rate, measures how the tool feels, not what it produced, and it reported no security check at all. Watch the two gaps a careful rollout can still leave: measuring feel instead of output, and an unknown no one wrote down.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Ordinary competent rollout: telemetry measures feel network: 4 components and 9 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 2 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- baseline
This models the ordinary-competent-adoption pattern documented in the case file - not a reconstruction of the actual rollout. Its value is the documentation quality of an average case (four phase gates, telemetry definitions, per-language deltas, stated limitations, published by the deployer); its two gaps are what make it instructive, not damning.
- baseline
The first gap is drawn on the independent model check, empty at baseline: the evaluation instrument is acceptance-rate telemetry, which the productivity literature ties to perceived rather than real productivity, on a population whose perception is measured to be miscalibrated - so a carefully measured acceptance number is still the wrong number. No measurement of real delivered output exists in the record. A number can be rigorously measured and still measure feel rather than outcome.
- assumed
The second gap is drawn on the latent oversight check, empty at baseline: no security evaluation was reported at all - an unrecorded unknown, one honesty step below the bank's recorded-inconclusive finding. Both end without a resolved security answer, but a recorded inconclusive finding is a governed unknown (named, carried forward, resolvable) while an unrecorded absence is invisible on every list of open questions. The honesty ladder has three rungs - resolve it, record it unresolved, or leave it unrecorded - and this org sits one below the bank.
- assumed
No product outcome is modeled here. This Lab reads institutional propagation only, and the business customers are boundary-only. The acceptance and satisfaction figures, and the missing security evaluation, live in the case file, and are never computed from anything in this diagram.
What this example does not show
- No product outcome is modeled. The Lab reads institutional propagation only; the business customers are boundary-only, and the acceptance and satisfaction figures and the missing security evaluation live in the case file, never computed on this diagram.
- The rollout report is a self-reported, non-peer-reviewed account; its acceptance telemetry measures adoption feel rather than delivered output, and the absence of a security evaluation is an unrecorded unknown drawn as a latent check, not a computed finding.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A mid-size enterprise ran a systematic four-phase evaluation-to-rollout of a commercial coding assistant across more than 400 developers, publishing acceptance telemetry (a 33 percent suggestion-acceptance rate, with 20 percent of suggested lines accepted), a 72 percent satisfaction figure, documented per-language variation, and stated limitations. Its evaluation instrument is acceptance-rate telemetry — which the productivity literature identifies as the measure most correlated with perceived productivity rather than outcome, and perception is measured to be miscalibrated for experienced developers, so acceptance telemetry captures adoption feel, not delivered output.
empirical- Industry Bakal, G., Dasdan, A., Katz, Y., Kaufman, M., & Levin, G. (2025). Experience with GitHub Copilot for Developer Productivity at Zoominfo [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2501.13282 https://arxiv.org/abs/2501.13282
- Academic Ziegler, A., Kalliamvakou, E., Li, X.A., et al. (2024). Measuring GitHub Copilot's Impact on Productivity. Communications of the ACM, 67(3). https://doi.org/10.1145/3633453 https://dl.acm.org/doi/10.1145/3633453
The deployment report stated its limitations but reported no security evaluation at all — an unrecorded unknown, one step less honest than a deployment that runs a security check and records the result as inconclusive, because an absence no one has written down is not a governed object and cannot be carried forward or resolved. The value of the case is the documentation quality of an ordinary, competent adoption — phase gates, telemetry definitions, per-language deltas, and stated limitations by the deployer itself — with the missing security question priced as the one thing even that documentation did not name.
empirical- Industry Bakal, G., Dasdan, A., Katz, Y., Kaufman, M., & Levin, G. (2025). Experience with GitHub Copilot for Developer Productivity at Zoominfo [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2501.13282 https://arxiv.org/abs/2501.13282
Where this connects
Institutional pressures in this domain
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Software engineering AI (coding assistants) domain page.
Levers available here and the patterns behind them
- Check copied records — Reconcile copied records
- Keep skills sharp — Deskilling-arrest mandate
- Gate vendor updates — Vendor quality gate
- Gate record entries — Human-in-the-loop write gating
- Check with a second model — Cross-model verification
- Review on schedule — Oversight cadence & retrospectives
- Escalate checks — State-feedback vigilance