PAN Lab example
Gated coding-assistant rollout at a regulated bank
The gate that recorded what it couldn't resolve: a bank's rollout
A regulated bank piloted a coding assistant with ~100 engineers, evaluated, then scaled to ~1,000. Modeled on its own published rollout. It reported gains - and recorded the security impact as explicitly inconclusive: a real gate that named what it could not resolve rather than asserting it away. So watch the honest move and its limit: a recorded unknown is a governed object, but the class-level evidence says insecure generation is common, so the unknown still has to be closed.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Gated coding-assistant rollout with a recorded unknown network: 4 components and 9 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 2 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- baseline
This models the trial-then-gate pattern documented in the case file - not a reconstruction of the actual rollout. Its governance topology is a bounded pilot (~100 of 5,000 engineers, six weeks), an explicit evaluation, and a scale decision (to ~1,000) - drawn as the present gate check, at a low level, because the bank actually exercised it.
- baseline
The case's defining datum is an honestly recorded unknown, drawn as the latent security check, empty at baseline: the bank ran a security evaluation and explicitly recorded the impact as inconclusive, carrying the unknown forward into the scaled deployment rather than resolving it by assertion or omitting it. A recorded unknown is a governed object - named and available to resolve - one step ahead of an unrecorded absence; the honest form of a governance finding.
- assumed
The unknown is not hypothetical: an independent assessment found ~40% of generated programs vulnerable across CWE-top-25 scenarios, with overconfident acceptance documented, so the carried-forward security question sits against a class-level literature in which insecure generation is common. The governable move is to keep resolving it - a security check that actually runs against generated code before it merges, the gate the pilot could not yet close.
- assumed
No banking or product outcome is modeled here. This Lab reads institutional propagation only, and the bank's customers are boundary-only. The productivity claims, the security-inconclusive finding, and the class-level vulnerability rate live in the case file, and are never computed from anything in this diagram. The report is a non-peer-reviewed preprint by the bank's own engineers; the self-recorded inconclusive finding is the honest counterweight inside its own account.
What this example does not show
- No banking or product outcome is modeled. The Lab reads institutional propagation only; the bank's customers are boundary-only, and the productivity claims, the security-inconclusive finding, and the class-level vulnerability rate live in the case file, never computed on this diagram.
- The rollout report is a non-peer-reviewed preprint by the bank's own engineers reporting its own success; the self-recorded inconclusive security finding is the honest counterweight inside that account, and the ~40% vulnerable-generation rate is a separate class-level assessment, not a measurement of this deployment.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before scaling a commercial coding assistant to roughly 1,000 engineers, publishing its own measurement of the rollout. The bank's engineers reported productivity and code-quality improvements — and recorded the security impact as explicitly inconclusive, a real gating decision taken and documented under uncertainty rather than resolved by assertion, with the honestly recorded unknown carried forward into the scaled deployment.
empirical- Industry Chatterjee, S., Liu, C.L., Rowland, G., & Hogarth, T. (2024). The Impact of AI Tool on Engineering at ANZ Bank: An Empirical Study on GitHub Copilot within Corporate Environment [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2402.05636 https://www.theregister.com/2024/02/10/anz_bank_github_copilot/
- Trade press The Register (2024, February 10). ANZ Bank test drives GitHub Copilot, decides it's worth the effort https://www.theregister.com/2024/02/10/anz_bank_github_copilot/
What the bank's inconclusive security finding leaves open is not hypothetical: an independent security assessment of code generated by a widely used assistant found that about 40 percent of generated programs contained vulnerabilities across scenarios spanning the CWE top-25 weaknesses, and separate research documents developers accepting insecure suggestions with overconfidence — so the security unknown a deployment carries forward unresolved sits against a class-level literature in which insecure generation is common.
empirical- Peer-reviewed Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., & Karri, R. (2022). Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions. In 43rd IEEE Symposium on Security and Privacy (SP 2022). https://doi.org/10.48550/arXiv.2108.09293 https://arxiv.org/abs/2108.09293
Where this connects
Institutional pressures in this domain
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Software engineering AI (coding assistants) domain page.
Levers available here and the patterns behind them
- Gate vendor updates — Vendor quality gate
- Pause AI on alarms — Deployment circuit-breaker
- Gate record entries — Human-in-the-loop write gating
- Check with a second model — Cross-model verification
- Review on schedule — Oversight cadence & retrospectives
- Escalate checks — State-feedback vigilance