PAN Lab example
Google ML code completion
Owning every node: the strength and the missing check
One organization builds the model, owns the repository, runs the review gates, and defines the telemetry. Modeled on a large in-house code-completion deployment. That unification is the strength - every lever is inside one boundary - and the risk: the party that builds, deploys, and measures the tool is the same, so the numbers are a self-report and no external check exists. Watch the check that owning everything designs out.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the In-house-completion-class: one org owns every node network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This deployment is the one in its domain that built its own assistant: an internal platform team wrote the completion model and ran the control-group measurement, so that team is drawn as its own group of staff with a tuning pathway into the model - a pathway a deployment running a vendor product simply does not have, and the reason improve-model is genuinely this organization's lever while vendor-gate is not. The independent check runs faintly rather than empty because iteration time and the fraction of code the model wrote were measured against a control group; it is held faint rather than higher because the building, deploying, and measuring party are the same organization. The monorepo pathway is drawn at full strength: the repository is both the corpus the model is trained on and the place its output lands, which is the tightest such loop in this catalogue. A heavy workload against limited capacity for a population above ten thousand developers.
- baseline
This models the one-organization-owns-every-node pattern documented in the case file - not a reconstruction of the actual system. Its defining feature is unification: the model (built in-house), the store (the monorepo that trains and contextualizes it), the gates (a code-review culture predating the assistant), and the telemetry all sit inside one boundary, so the organization can in principle tune the whole loop.
- baseline
The defining absence is independence, drawn on the independent model check, empty at baseline: when the building, deploying, and measuring party are the same organization, no external check exists at all, and the numbers are an engineering-blog self-report rather than a peer-reviewed or independently reproduced evaluation. The internal review culture is strong and owned; the missing check is the external one the organization cannot supply to itself.
- assumed
The composition bound applies in the organization's own research idiom (drawn as the latent external-oversight check): AI amplifies an organization's existing strengths and weaknesses rather than substituting for them, so the gates and platform quality it owns are what decide whether the acceptance rate becomes better software or just more of it - the individual gain does not compose to the org outcome on its own.
- assumed
No product or software outcome is modeled here. This Lab reads institutional propagation only, and the downstream users of the software are boundary-only. The acceptance rates and iteration-time figures live in the case file, and are never computed from anything in this diagram.
What this example does not show
- No product or software outcome is modeled. The Lab reads institutional propagation only; the downstream users of the software are boundary-only, and the acceptance rates and iteration-time figures live in the case file, never computed on this diagram.
- The reported numbers are a first-party engineering-blog self-report with a control group, not a peer-reviewed or independently reproduced evaluation; the diagram draws the missing external check as latent, it does not compute the benefit.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
An in-house machine-learning code-completion system built, deployed, and measured by a company's own platform organization for more than 10,000 internal developers reported, against a control group, a 25 to 34 percent suggestion-acceptance rate, a 6 percent reduction in coding iteration time versus control, and 3 percent of new code characters coming from the model at the time of measurement. The measuring party, the building party, and the deploying party were the same organization, and the numbers were published as an engineering-blog self-report rather than a peer-reviewed or independent evaluation.
empirical- Vendor Tabachnyk, M., & Nikolov, S. (2022). ML-Enhanced Code Completion Improves Developer Productivity. Google Research Blog. https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/
The same company's cross-industry research program reported that AI-assisted software development amplifies an organization's existing strengths and weaknesses rather than substituting for them, with policy clarity and platform investment identified as the levers that determine whether AI adoption improves or degrades delivery — evidence that the individual coding gains do not compose to organization-level outcomes on their own, and that the deploying organization's existing gates and platform quality are what decide the result.
empirical- Reference Google Cloud DORA (2025). State of AI-assisted Software Development (2025 DORA Report). https://dora.dev/dora-report-2025/
Where this connects
Institutional pressures in this domain
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Software engineering AI (coding assistants) domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Keep prompts neutral — Framing and mirroring reduction
- Peer sharing rules — Peer-edge governance
- Gate record entries — Human-in-the-loop write gating
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Escalate checks — State-feedback vigilance
- Review the riskiest first — Risk-tiered oversight
- Store less data — Data minimization
- Upgrade model — Improve the model