Skip to content

PAN Lab example

GitHub Copilot at ZoomInfo

Measured carefully but measuring the wrong thing: an ordinary rollout

ZoomInfo rolled out GitHub Copilot to over 400 developers in four phases. It measured acceptance and satisfaction, not delivered output, and reported no security evaluation.

See more

GitHub Copilot is a commercial code-completion assistant. It suggests code to developers as they write, and each developer accepts or rejects each suggestion. ZoomInfo adopted it under enterprise terms for its own engineers.

What ZoomInfo did

ZoomInfo is a mid-size United States company. It evaluated Copilot and rolled it out in four phases, across more than 400 developers. It defined the usage measures it collected, ran satisfaction surveys, broke results down by programming language, and stated the limitations of its study.

ZoomInfo published the account itself, as a preprint and on its engineering blog. The sources behind this case do not describe what each of the four phases involved.

What it reported

Developers accepted 33 percent of Copilot's suggestions, and 20 percent of the lines it suggested. The satisfaction figure from ZoomInfo's surveys was 72 percent. Results varied by programming language.

What acceptance rate does and does not show

Acceptance rate is the case's headline number, and ZoomInfo defined and tracked it with care. A GitHub-authored study in Communications of the ACM found it is the usage measure most strongly correlated with how productive developers feel.

How developers feel is not what they deliver. A separate randomized study, a 2025 preprint that is not peer-reviewed, followed 16 experienced open-source developers on code they knew well. With an AI assistant they were about 19 percent slower. They believed it had made them about 20 percent faster.

So acceptance rate measures how the tool feels, however carefully it is counted. No measure of the code ZoomInfo actually delivered appears in the record.

What the report leaves out

The report states its limitations, but it reports no security evaluation at all. That leaves an unknown nobody wrote down. The case file does not put this down to bad faith.

ANZ Bank's rollout of the same tool, a separate case, ran a security evaluation and recorded the result as inconclusive. Neither ends with a security answer. But a recorded unknown is named, carried forward, and can be resolved. An unrecorded one appears on no list of open questions, so nobody can be assigned to close it.

There are three steps: resolve the unknown, record it unresolved, or leave it unrecorded. Leaving it unrecorded is the ordinary default. On security, this rollout sits on that lowest step.

What is known about the tool's security elsewhere

A 2022 peer-reviewed study, presented at the IEEE Symposium on Security and Privacy, had Copilot write 1,689 programs for scenarios built around high-risk kinds of software weakness. About 40 percent of them contained vulnerabilities. That is a finding about the tool, in scenarios chosen to test security, not a measurement of ZoomInfo's code.

Why this case is here

This case is an ordinary, competent adoption, and it does most things right. It is not a randomized experiment, and it is not a regulated bank's gated rollout. Its value is how well an average rollout was documented.

Because the rollout is careful, its two gaps are instructive rather than damning. It measures feel instead of output, and it leaves its security unknown unwritten.

Where these facts come from

ZoomInfo published the rollout report itself. Its authors are Bakal, Dasdan, Katz, Kaufman, and Levin (2025). It is a self-reported preprint, not peer-reviewed. The acceptance and satisfaction figures come from it, and so does the absence of a security evaluation.

The studies of acceptance rate, of developers' speed, and of Copilot's security are separate research. None of them studied ZoomInfo.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other.

This case has a budget of 10 units. Explore (No Targets) sets no targets. There, and under Service Targets Only, two tools costing 4 of the 10 units are enough to keep the mistakes on this network contained. Under Service Targets Only, the same two also meet the service target, a floor on how much Copilot helps the work. One such pair is Review on schedule with Escalate checks.

Under Service Targets Only, more tools are not better. With every offered tool at its highest setting, ignoring the budget, Copilot's net help to the work falls below the service target.

Under Service and Safety Targets and All Governance Targets, the targets are not fully addressable with the available tools. Those levels require every failure pathway closed. Three stay open in every combination, even with every tool at its highest setting.

The three are developers accepting or rejecting Copilot's suggestions, Copilot drawing on the repositories, and developers reading the repositories. No tool offered in this case acts on those three pathways. So what stands in the way is the shape of this deployment, not a shortfall of budget.

Stylized model of a documented deploymentSoftware engineering AI (coding assistants)

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Ordinary competent rollout: telemetry measures feel network: 4 components and 9 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 2 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 4 assumptions
  • baseline

    This network follows the ordinary, competent adoption the case file documents. It draws the parts the case file describes, not every step of the rollout. The rollout's value is how well an average rollout was documented. ZoomInfo itself published its four-phase rollout with a gate at each phase, definitions of its usage figures, results by programming language, and stated limitations. Its two gaps make it instructive, not damning.

  • baseline

    The first gap is the measure of delivered output. ZoomInfo judged the rollout by acceptance rate. A GitHub-authored study found acceptance rate is the usage measure most strongly correlated with how productive developers feel. A separate 2025 preprint, not peer-reviewed, found experienced developers felt faster with an AI assistant on code they knew well, while working slower. So a carefully counted acceptance figure still measures feel, not output. The record contains no measure of delivered output.

  • assumed

    The second gap is the security evaluation, and ZoomInfo's report describes none at all. ANZ Bank's rollout, a separate case, ran one and recorded the result as inconclusive. Neither ends with a security answer. A recorded unknown is named, carried forward, and can be resolved. An unrecorded one appears on no list of open questions. There are three steps: resolve an unknown, record it unresolved, or leave it unrecorded. On security, this rollout sits one step below the bank.

  • assumed

    No product outcome is modeled here. The network shows only how mistakes can be passed on between Copilot, ZoomInfo's developers, and its repositories. ZoomInfo's business customers are outside it. The acceptance and satisfaction figures, and the missing security evaluation, come from the case file. Nothing in this network computes them.

What this example does not show

Show all 2 limitations
  • No product outcome is modeled. This example shows only how mistakes can be passed on inside ZoomInfo's engineering work. ZoomInfo's business customers are outside it. The acceptance and satisfaction figures, and the missing security evaluation, come from the case file. Nothing in this network computes them.
  • The rollout report is ZoomInfo's own account, and it is not peer-reviewed. Its usage figures measure how the tool feels to use, not the output delivered. The network shows the missing security evaluation as a check between people. That is a reading of the record, not a computed finding.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A mid-size enterprise ran a systematic four-phase evaluation-to-rollout of a commercial coding assistant across more than 400 developers, publishing acceptance telemetry (a 33 percent suggestion-acceptance rate, with 20 percent of suggested lines accepted), a 72 percent satisfaction figure, documented per-language variation, and stated limitations. Its evaluation instrument is acceptance-rate telemetry — which the productivity literature identifies as the measure most correlated with perceived productivity rather than outcome, and perception is measured to be miscalibrated for experienced developers, so acceptance telemetry captures adoption feel, not delivered output.

    empirical
    • Industry Bakal, G., Dasdan, A., Katz, Y., Kaufman, M., & Levin, G. (2025). Experience with GitHub Copilot for Developer Productivity at Zoominfo [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2501.13282 https://arxiv.org/abs/2501.13282
    • Academic Ziegler, A., Kalliamvakou, E., Li, X.A., et al. (2024). Measuring GitHub Copilot's Impact on Productivity. Communications of the ACM, 67(3). https://doi.org/10.1145/3633453 https://dl.acm.org/doi/10.1145/3633453
  • The deployment report stated its limitations but reported no security evaluation at all — an unrecorded unknown, one step less honest than a deployment that runs a security check and records the result as inconclusive, because an absence no one has written down is not a governed object and cannot be carried forward or resolved. The value of the case is the documentation quality of an ordinary, competent adoption — phase gates, telemetry definitions, per-language deltas, and stated limitations by the deployer itself — with the missing security question priced as the one thing even that documentation did not name.

    empirical
    • Industry Bakal, G., Dasdan, A., Katz, Y., Kaufman, M., & Levin, G. (2025). Experience with GitHub Copilot for Developer Productivity at Zoominfo [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2501.13282 https://arxiv.org/abs/2501.13282

Where this connects

Institutional pressures in this domain

  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.

All of them in context on the Software engineering AI (coding assistants) domain page.

Levers available here and the patterns behind them

Documented case histories