PAN Lab example
Gated coding-assistant rollout at a regulated bank
The gate that recorded what it couldn't resolve: a bank's rollout
ANZ Bank tried GitHub Copilot with about 100 engineers, then extended it to about 1,000. It recorded the tool's effect on security as inconclusive.
See more
GitHub Copilot is a commercial coding assistant that suggests code to engineers as they write. ANZ Bank adopted it under enterprise terms, and each engineer decides whether to accept a suggestion. The bank's trial recorded its effect on code security as inconclusive.
How the rollout ran
ANZ Bank is a regulated bank with about 5,000 engineers. The case file lists its jurisdiction as Australia.
The bank ran a structured six-week experiment with about 100 of those engineers. It evaluated the results, then decided to extend the assistant to about 1,000 engineers. It carried what the trial found into the larger rollout.
That sequence is a bounded trial, an explicit evaluation, and a scale decision. The case file notes that most deployments in this field skip it.
What the bank found
The bank's engineers reported gains in productivity and code quality. They recorded the assistant's effect on security as explicitly inconclusive.
The case file treats that recorded unknown as the case's most valuable finding. The bank named the question and carried it forward. It did not settle it by assertion or leave it out.
A recorded unknown can be answered later. A deployment that says nothing about security leaves no question on record for anyone to answer.
What security research says about tools like this
A 2022 independent, peer-reviewed study by Pearce and colleagues tested code written by GitHub Copilot. About 40 percent of 1,689 generated programs contained security vulnerabilities. Its test scenarios were drawn from the top 25 weakness types in the Common Weakness Enumeration, a public list of kinds of software flaw.
A separate peer-reviewed study, by Perry and colleagues in 2023, found that people using an AI coding assistant wrote less secure code in four of five tasks. They were also more likely to believe their code was secure.
Neither study measured ANZ's deployment. Together they show that the bank's open security question is a live risk, not a formality, in a rollout to about 1,000 engineers.
Who wrote the report
The rollout report, by Chatterjee, Liu, Rowland, and Hogarth, is a 2024 preprint written by the bank's own engineers. It was not peer-reviewed, and it reports the bank's own success. The Register, a technology news site, covered the rollout in February 2024.
The bank measured its own rollout, rather than a vendor or an academic team. That is both the case's strength and its limit. The inconclusive security finding is the bank recording a limit on its own good news.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other.
This case offers seven tools, the ones within the bank's control, and a budget of 9 units. Each tool's price is the same at every target level.
Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met within the budget by more than one combination of tools. Escalate checks meets them on its own for 2 units. So does Gate record entries, for 3.
Under Service and Safety Targets and All Governance Targets, this case is not fully addressable with the available tools. Using all seven at their highest settings costs 27 units, three times the budget. Even then, three failure pathways stay open:
Engineers accepting or rejecting GitHub Copilot's suggestions.
The assistant reading the repository to shape its suggestions.
Engineers reading the repository as their codebase.
No tool this case offers closes them. They are how engineers use the assistant and their own code. This is a finding about the deployment, not a gap in your approach.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Gated coding-assistant rollout with a recorded unknown network: 4 components and 9 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 2 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 4 assumptions
- baseline
This network follows the trial-then-scale pattern the case file documents. It is not a reconstruction of the actual rollout. The bank ran a six-week trial with about 100 of its 5,000 engineers. It then evaluated the trial and decided to extend the assistant to about 1,000. The network draws that sequence as one of its checks, because the bank actually carried it out.
- baseline
The case turns on an unknown the bank wrote down. The bank recorded the assistant's effect on security as inconclusive. It carried that unknown into the larger rollout, rather than settling it by assertion or leaving it out. The network draws it as its own check: the security evaluation of generated code. A recorded unknown has a name and can be settled later. That puts it one step ahead of a gap nobody recorded.
- assumed
The unknown is not hypothetical. A 2022 study found security vulnerabilities in about 40 percent of the 1,689 programs GitHub Copilot generated in its tests. Its test scenarios were drawn from the top 25 weakness types in the Common Weakness Enumeration, a public list of kinds of software flaw. Separate research found developers accepting insecure suggestions with too much confidence. The step that would settle the bank's question is a security check that runs on generated code before it is saved to the repository. The trial did not get that far.
- assumed
No banking or product outcome is modeled here. This network shows only how mistakes move between the assistant, the engineers, and the repository. The bank's customers are outside it. The productivity claims, the inconclusive security finding, and the vulnerability rate come from the case file. Nothing in this network computes them. The report is a preprint by the bank's own engineers, not peer-reviewed. Its self-recorded inconclusive finding is the report's own caution against its good news.
What this example does not show
Show all 2 limitations
- This example models no banking or product outcome. It shows how mistakes can pass between the assistant, the engineers, and the repository, and nothing more. The bank's customers are outside it. The productivity claims, the inconclusive security finding, and the vulnerability rate come from the case file, and nothing here computes them.
- The rollout report is a preprint by the bank's own engineers, not peer-reviewed, and it reports the bank's own success. Its inconclusive security finding is the report's own caution against its good news. The 40 percent vulnerability rate comes from a separate study of GitHub Copilot's code. It is not a measurement of this deployment.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A regulated bank ran a structured six-week internal experiment with about 100 of its 5,000 engineers before scaling a commercial coding assistant to roughly 1,000 engineers, publishing its own measurement of the rollout. The bank's engineers reported productivity and code-quality improvements — and recorded the security impact as explicitly inconclusive, a real gating decision taken and documented under uncertainty rather than resolved by assertion, with the honestly recorded unknown carried forward into the scaled deployment.
empirical- Industry Chatterjee, S., Liu, C.L., Rowland, G., & Hogarth, T. (2024). The Impact of AI Tool on Engineering at ANZ Bank: An Empirical Study on GitHub Copilot within Corporate Environment [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2402.05636 https://www.theregister.com/2024/02/10/anz_bank_github_copilot/
- Trade press The Register (2024, February 10). ANZ Bank test drives GitHub Copilot, decides it's worth the effort https://www.theregister.com/2024/02/10/anz_bank_github_copilot/
What the bank's inconclusive security finding leaves open is not hypothetical: an independent security assessment of code generated by a widely used assistant found that about 40 percent of generated programs contained vulnerabilities across scenarios spanning the CWE top-25 weaknesses, and separate research documents developers accepting insecure suggestions with overconfidence — so the security unknown a deployment carries forward unresolved sits against a class-level literature in which insecure generation is common.
empirical- Peer-reviewed Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., & Karri, R. (2022). Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions. In 43rd IEEE Symposium on Security and Privacy (SP 2022). https://doi.org/10.48550/arXiv.2108.09293 https://arxiv.org/abs/2108.09293
Where this connects
Institutional pressures in this domain
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Software engineering AI (coding assistants) domain page.
Levers available here and the patterns behind them
- Gate vendor updates — Vendor quality gate
- Pause AI on alarms — Deployment circuit-breaker
- Gate record entries — Human-in-the-loop write gating
- Check with a second model — Cross-model verification
- Review on schedule — Oversight cadence & retrospectives
- Train the staff — AI literacy & boundary rules
- Escalate checks — State-feedback vigilance