PAN Lab example
The same AI running hands off: the agentic office
In this made-up example office, AI agents act on cases themselves, and a stretched staff waves most of their work through.
See more
The AI here is a generic assistant with autonomous agents, run by the office. It is not a real product. Its agents are AI programs that act without a person approving each step: they draft decisions, act on cases, and file entries in the case records.
One AI, three offices
This office is invented to illustrate a point. It is not a documented deployment. It is one of three made-up offices in the Lab. All three use the same AI, which makes the same mistakes, as often, in each. Only the way each office is run differs.
In the supervised office, a person checks every output before it enters the record. The professional office runs the AI under full guardrails: checks on its output, sign-off before record entries, and time set aside for staff to verify.
How this office is run
Filing is automatic: the agents write into the case records themselves. Most of the agents' output is used by other agents, which cannot easily check facts. Mistakes pass quickly from the agents to staff and to the records. The records are audited little.
A small team oversees the agents. Their time for review stays the same while the agents' output grows. So they check a shrinking share of it, and wave most of it through.
What the comparison found
A published simulation compared the three made-up offices, all running the same AI. The supervised and professional offices contained mistakes. This office amplified them.
These results come from the simulation, not from measuring any real office. The comparison's point is that the culture around the tool moved the outcome more than the tool did.
Why this office is dangerous
The danger is not the agents' speed alone. It is that review stays the same while the work speeds up. Suppose checking kept pace, through software that reviews each agent output as fast as it is made. Then checking would keep up with the work.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A pathway is closed when mistakes stop passing along it. The work along it may go on.
This case has a budget of 12 units. Each tool costs the same at every target level. Explore (No Targets) sets no targets.
Under Service Targets Only, the targets can be met in many ways, but no single tool is enough. Lingering effects is a Lab setting, on when a case opens, in which damage outlasts its cause.
With lingering effects on, the cheapest ways cost 5 units. One is Escalate checks with Peer sharing rules at its stronger setting. The other is Gate record entries with Escalate checks. Gate record entries requires sign-off before anything enters the case records. With lingering effects off, Escalate checks with Peer sharing rules costs 4.
Service and Safety Targets and All Governance Targets both ask you to close every failure pathway, among other targets. Both can be met, but only just. The cheapest ways cost 11 of the 12 units and use five tools.
Under Service and Safety Targets, two sets of five work. Both hold Mark AI-written records, Escalate checks, Peer sharing rules, and Keep prompts neutral. The fifth is Gate record entries or Store less data.
Under All Governance Targets, only the set with Gate record entries works. At that level lingering effects are on, and Store less data works at reduced strength. Understand the system, a tool not offered here, funds study of what the deployment is really doing, so other tools can be aimed. It then leaves Staff enter outputs in records open.
At either level, the spare unit can buy the stronger setting of Mark AI-written records, Peer sharing rules, or Keep prompts neutral.
More is not better here. Every tool at its strongest setting, ignoring the budget, closes every pathway. It still misses the targets at all three levels that set them, because the AI then adds too little to the work.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Agentic low-oversight office network: 4 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 4 assumptions
- baseline
How far mistakes build up in this office comes from a simulation comparing three made-up offices. No real office was measured.
- assumed
The AI makes the same mistakes, as often, in all three modeled offices. Only the office around it differs: who checks, what is filed automatically, and how much records are audited.
- assumed
Everyone in this office is assumed to check the AI's work in the same way. The example does not show differences between staff members.
- assumed
From the start, agents pass outputs to other agents and staff pass shortcuts to each other. Coworkers seldom give second opinions, and no second model checks the agents. That lack defines this office.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
In the sociotechnical simulation, the same AI in three modeled office cultures, stylized and not real workplaces, led to very different outcomes. Mistakes built on one another far more under low-oversight autonomy than under human supervision or high-governance professional controls.
scenarioillustrative PAN-run resultNo published source is attached to this claim yet.
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Verify output — Put a verifier on the agent
- Upgrade model — Improve the model
- Gate record entries — Human-in-the-loop write gating
- Mark AI-written records — Provenance labeling
- Escalate checks — State-feedback vigilance
- Pause AI on alarms — Deployment circuit-breaker
- Vet connections — Connection authorization
- Peer sharing rules — Peer-edge governance
- Keep prompts neutral — Framing and mirroring reduction
- Store less data — Data minimization
- Assign a challenger — Structured dissent
- Check with a second model — Cross-model verification
Documented case histories
- Magic Notes (Beam)
- Minute / Local Transcribe
- Massachusetts DTA call summaries
- Justice Transcribe
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- DWP Whitemail Insights and Vulnerability Scanner
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- CDTFA Axyom Assist
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot
- Amsterdam Smart Check