PAN Lab example
Google ML code completion
Owning every node: the strength and the missing check
Google's platform team built a code-completion system for more than 10,000 Google developers, then measured it themselves. No outside party checked the results.
See more
Google's ML code completion is a machine-learning system that suggests code to Google developers in their internal code editor. Google's platform team built it and trained it on code from Google's monorepo, its single shared code repository. Google built it for more than 10,000 of its developers.
What Google reported
Google published the measurement on the Google Research blog in July 2022. It used a control group, a set of developers compared with those using the system. Developers accepted 25 to 34 percent of the system's suggestions. Coding iteration time fell 6 percent compared with the control group. At the time, 3 percent of the characters in new code came from the system.
The sources for this network do not define coding iteration time further.
The numbers are an engineering-blog self-report. They were not peer reviewed or independently reproduced.
One company owns every part
One organization builds the system, owns the repository, runs the review gates, and defines what is measured. Google built the system in-house. The monorepo's code trains the system and gives it context. Google's code-review and testing culture predates the assistant.
That unification is the strength. Every lever sits inside one company, so Google can, in principle, tune the whole loop from end to end. None of the deployments running a vendor's tool can do that.
The risk in owning everything
The same organization built the system, deployed it, and measured it. So the published numbers are a self-report, and no outside party checked them.
That is not a charge of bad faith. The control group and the published method are more than most deployments offer. It is a fact about who can see the result.
Owning every part is the best position from which to tune the whole loop. It is also the position where a missing outside check is hardest to notice, because everything is already inside.
What Google's own research found
DORA, Google Cloud's cross-industry research program on software development, published its 2025 report on AI-assisted software development. It drew on about 5,000 technology professionals and more than 100 hours of qualitative data. It is industry research with a published method, not a peer-reviewed study.
The report found that AI amplifies an organization's existing strengths and weaknesses rather than replacing them. It named clear policy and investment in the platform as what decides whether adoption helps or hurts software delivery.
Applied to Google's own deployment, the individual gains do not become an organization-wide result on their own. The review gates and platform quality Google owns decide whether the acceptance rate becomes better software or just more of it. DORA does not audit this system.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other.
Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met within this case's budget of 11 with a single lever costing 2. One is Mark AI-written records, which marks machine-written code where developers and the system read the monorepo. Another is Keep prompts neutral, which keeps the requests developers and the platform team give the system free of leading framing.
Under Service and Safety Targets and All Governance Targets, the targets can also be met within the budget, and every failure pathway must be closed. Every combination that meets them includes three levers: Mark AI-written records, Keep prompts neutral, and Escalate checks. Escalate checks raises the checking done by developers and the platform team when monitoring flags trouble.
Each combination also closes the pathway by which developers commit code, with Gate record entries or with Store less data. The cheapest combinations cost 9 of the 11. Two combinations also use Peer sharing rules, which adds to the Outside review and re-approval pathway, and they use the whole budget. At All Governance Targets, one combination closes every pathway but slows the work too much to meet the service target.
No combination that meets these targets uses Check with a second model, which has a second model check each suggestion. It closes none of the failure pathways, and adding it to the cheapest combination would exceed the budget.
So the targets can be met without the outside check this case is about. Meeting them shows the loop inside Google contained. It does not supply the independent confirmation of Google's numbers that the record lacks.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the In-house-completion-class: one org owns every node network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 5 assumptions
- assumed
Google's platform team built the completion system and ran the control-group measurement. So this example shows that team as its own group, able to change the system. A company running a vendor's product has no such group. That is why improving the model is a lever Google holds here, and gating a vendor is not offered. The control-group measurement is a real internal check, but the party that built the system ran it. The monorepo is both the code the system is trained on and the place its accepted suggestions are committed. Among this Lab's software examples, that is the tightest loop between a system and its own record. The example assumes a heavy review workload against limited capacity, for more than ten thousand developers.
- baseline
This example follows the pattern the case file documents, in which one organization owns every part. It is not a reconstruction of Google's actual system. Google built the system in-house. Its monorepo trains the system and gives it context. Its code-review culture predates the assistant, and its platform team defines what is measured. All of these sit inside one company, so Google can, in principle, tune the whole loop.
- baseline
What this deployment lacks is independence. When one organization builds, deploys, and measures a system, no outside check exists at all. Google's numbers are an engineering-blog self-report, not a peer-reviewed or independently reproduced evaluation. Its internal review culture is strong and its own. The missing check is the outside one, which no organization can supply to itself.
- assumed
Google's own DORA research program found that AI amplifies an organization's existing strengths and weaknesses rather than replacing them. This example applies that finding to Google's own deployment. The review gates and platform quality Google owns decide whether the acceptance rate becomes better software or just more of it. The individual gain does not become an organization-wide result on its own.
- assumed
This example does not model any product or software outcome. It follows only how mistakes are passed between the parts of the deployment. The people who use Google's software are outside it. The acceptance rates and iteration-time figures come from the case file, and nothing in this network computes them.
What this example does not show
Show all 2 limitations
- This example does not show what happens to the software or to the people who use it. It follows how mistakes are passed between the system, the developers, and the monorepo. Google's acceptance rates and iteration-time figures come from its published measurement, and nothing in this network computes them.
- Google's numbers are its own report on an engineering blog, measured with a control group. They were not peer reviewed or independently reproduced. The network includes the outside check the record lacks, so you can see what adding one changes. It does not compute the system's benefit.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
An in-house machine-learning code-completion system built, deployed, and measured by a company's own platform organization for more than 10,000 internal developers reported, against a control group, a 25 to 34 percent suggestion-acceptance rate, a 6 percent reduction in coding iteration time versus control, and 3 percent of new code characters coming from the model at the time of measurement. The measuring party, the building party, and the deploying party were the same organization, and the numbers were published as an engineering-blog self-report rather than a peer-reviewed or independent evaluation.
empirical- Vendor Tabachnyk, M., & Nikolov, S. (2022). ML-Enhanced Code Completion Improves Developer Productivity. Google Research Blog. https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/
The same company's cross-industry research program reported that AI-assisted software development amplifies an organization's existing strengths and weaknesses rather than substituting for them, with policy clarity and platform investment identified as the levers that determine whether AI adoption improves or degrades delivery — evidence that the individual coding gains do not compose to organization-level outcomes on their own, and that the deploying organization's existing gates and platform quality are what decide the result.
empirical- Reference Google Cloud DORA (2025). State of AI-assisted Software Development (2025 DORA Report). https://dora.dev/dora-report-2025/
Where this connects
Institutional pressures in this domain
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Software engineering AI (coding assistants) domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Keep prompts neutral — Framing and mirroring reduction
- Peer sharing rules — Peer-edge governance
- Gate record entries — Human-in-the-loop write gating
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Escalate checks — State-feedback vigilance
- Review the riskiest first — Risk-tiered oversight
- Store less data — Data minimization
- Upgrade model — Improve the model