Skip to content

PAN Lab example

A commercial code assistant across three enterprises

Big lift for novices but slower for experts: a coding assistant

GitHub Copilot suggests code to developers. Trials at three companies found large gains for less-experienced developers. An independent study found experienced developers slower using AI.

See more

GitHub Copilot is a commercial AI assistant that suggests code as a developer types in a code editor. The developer accepts or rejects each suggestion and saves accepted code into the company's shared code repository. The same assistant makes suggestions to every developer who uses it.

The trials at three companies

Microsoft, Accenture, and a Fortune 100 company that the study does not name gave GitHub Copilot to randomly chosen developers. In all, 4,867 developers took part. The trials compared developers who had the assistant with developers who did not.

Cui, Demirer, Jaffe, Musolff, Peng, and Salz wrote up the trials. The study was registered publicly in advance. The peer-reviewed journal Management Science published it in 2025.

Pooled across the three companies, the trials found a 26.08 percent increase in completed tasks. The gains were concentrated among less-experienced developers, who were also more likely to take up the assistant. For experts, the gains were near zero.

Who ran the trials

The companies ran the experiments themselves. Several of the study's authors are affiliated with the assistant's vendor. The case file names the advance registration and the peer-reviewed journal as the checks that make the numbers credible despite that.

The independent study of experienced developers

METR, an independent nonprofit research institute, ran a separate randomized study in 2025. It measured 16 experienced open-source developers working 246 tasks in mature code repositories they already knew well.

When allowed to use AI tools, they were about 19 percent slower than without them. They believed the AI had made them about 20 percent faster.

The study is a preprint and has not been peer reviewed. In February 2026, METR announced a redesigned follow-up experiment. Its 16 developers were not part of the trials at the three companies. The study shows what AI tools did for these experts on code they knew well.

What the two studies show together

The benefit changes with experience. One number cannot show it. It is large for less-experienced developers, and near zero or negative for experts on familiar code. A single productivity figure, quoted without that curve, reports the result for less-experienced developers and lets it stand for everyone.

Most deployments track how the tool feels, through the share of suggestions developers accept. GitHub's own researchers, writing in Communications of the ACM in 2024, found that share is the usage measure most closely linked to how productive developers feel. The case file says it is the measure that is miscalibrated for the engineers the tool helps least.

Two features of how the tool is used

First, the developer who accepts a suggestion is also its reviewer of record. So no independent check stands between an accepted suggestion and the shared repository, unless the company builds one.

Second, the shared repository the assistant writes into is the code later developers copy from. Later assistants train on it and read it as context. So an accepted error does not stay in one place. It becomes inherited code that stays in use for a long time.

These two traps come from the case file's reading of this kind of deployment. The sources do not describe code review, testing, or security checks at the three companies.

What happens across a whole organization

Google Cloud's DORA research program runs a cross-industry survey on software delivery. Its 2024 report estimated that each 25 percent increase in AI adoption went with a 1.5 percent decrease in delivery throughput. It also estimated a 7.2 percent decrease in delivery stability.

The case file reads this as evidence that one developer's speed-up does not add up to better delivery for the organization on its own. It adds up only if the code review and testing gates keep up with the extra changes the assistant adds.

The survey is industry data with a published method, not peer reviewed. It does not study these three companies.

What this network is drawn from

This network is drawn from the public record of these studies. It shows the structure that record describes, not any company's own system. The sources for this case report no security evaluation of the assistant's code at the three companies.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met within this case's budget of 11 units. The cheapest way costs 2 units and uses one tool, Escalate checks.

Under Service and Safety Targets and under All Governance Targets, the targets can also be met within the budget. The cheapest way costs 9 units and uses four tools together: Mark AI-written records, Keep prompts neutral, Escalate checks, and Store less data.

More checking is not always better here. Using every tool at its strongest setting closes every failure pathway, but it meets the targets at none of the three levels that set them. The added checks and waits leave the assistant adding too little to the developers' work.

Stylized model of a documented deploymentSoftware engineering AI (coding assistants)

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Coding-assistant-class with the operator-verifier collapse network: 5 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 6 assumptions
  • assumed

    This example assumes the developers' review work is heavy for the time they have. The evidence is company-run randomized trials across thousands of employed developers. The amount of review grows with the number of suggestions a developer receives. The example also assumes experienced developers review suggestions more closely. The evidence is about the split between the two groups. The trials found a large gain for less-experienced developers. An independent study found experienced developers on familiar code slowed while feeling faster.

  • baseline

    This example follows the coding-assistant pattern documented in the case file. It is not a copy of any company's actual deployment. Two structural traps define it. First, the developer who accepts a suggestion is also its reviewer of record, on the pathway named Acceptance and review in one step. Second, the shared repository is written into and also read. The assistant and later developers read it, on the pathways named Repository as the assistant's context and Developers read the repository.

  • assumed

    Both developer groups receive the same suggestions, accept or reject them, and save code into the repository. The case documents very different measured outcomes for them. The trials found a pooled 26.08 percent increase in completed tasks, concentrated among less-experienced developers. An independent randomized study measured experienced developers about 19 percent slower on familiar code, while they believed they were about 20 percent faster. The example does not compute that difference. It is an outside measurement carried in the case file. The example shows it only as the difference that a single benefit figure hides.

  • assumed

    This example includes a check on the assistant, because the main evidence is a randomized experiment. It was registered in advance and peer reviewed, which is stronger than the self-reports elsewhere in this domain. Several authors are affiliated with the assistant's vendor, and the firms ran the experiments themselves. The advance registration and the peer-reviewed journal are the checks that make it credible. No security evaluation is in this case's own record. The security risk for this kind of tool is shown in the Lab's example of ANZ Bank's gated rollout.

  • baseline

    This example treats the organization-wide effect as the domain's defining risk. Google Cloud's DORA research program estimated a 7.2 percent decrease in delivery stability for each 25 percent increase in AI adoption, across industries. So one developer's speed-up can lower the organization's delivery unless the code review and testing gates keep up with the extra changes. The sources do not show these gates doing that at the three companies. Giving that check people and time is what lets the individual gain add up for the organization.

  • assumed

    This example does not model any product or software outcome. It shows how errors move among the assistant, the developers, and the repository. The people who use the finished software are outside the network. The productivity figures, the experts' slowdown and mistaken sense of speed, and the delivery-stability estimate come from the case file. Nothing in this example computes them.

What this example does not show

Show all 2 limitations
  • This example does not show any product or software outcome. It shows how errors move among the assistant, the developers, and the repository. The people who use the finished software are outside the network. The productivity figures, the experts' slowdown, and the delivery-stability estimate come from the case file. Nothing in the network computes them.
  • The 26 percent gain and the roughly 19 percent slowdown are outside measurements from two separate studies. One is a field experiment registered in advance, with several authors affiliated with the vendor. The other is an independent randomized study of experienced developers. The network shows both developer groups receiving the same suggestions, and it never computes the benefit or how it differs between them.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • Company-run, pre-registered, peer-reviewed randomized rollouts of a commercial code-completion assistant across 4,867 developers at three enterprises found a pooled 26.08 percent increase in completed tasks, with gains concentrated among less-experienced developers. An independent randomized study of 16 experienced open-source maintainers on 246 tasks in familiar repositories bounded the expert tail from the other direction: those developers were about 19 percent slower with AI tools while believing themselves about 20 percent faster — a measured perception-reality gap that means a uniform productivity number overstates the effect for senior engineers.

    empirical
    • Academic Cui, Z.K., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2025). The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science. https://doi.org/10.1287/mnsc.2025.00535 https://pubsonline.informs.org/doi/10.1287/mnsc.2025.00535
    • Industry Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR. https://doi.org/10.48550/arXiv.2507.09089 https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
  • Individual coding-assistant gains do not automatically compose to organization-level delivery outcomes: a cross-industry research program measured a roughly 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability for every 25 percent increase in AI adoption, evidence that the churn the assistant adds must be absorbed by code-review and testing gates or the individual speed-up degrades the organization's delivery performance.

    empirical
    • Reference Google Cloud DORA (2024). Accelerate State of DevOps Report 2024. https://dora.dev/research/2024/dora-report/

Where this connects

Institutional pressures in this domain

  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.

All of them in context on the Software engineering AI (coding assistants) domain page.

Levers available here and the patterns behind them

Documented case histories