PAN Lab example
The same AI under full guardrails: the professional office
A modeled office, not a real one: an AI assistant under strong checks, tested by staff pasting client data out and a silent vendor update.
See more
The AI assistant in this case is not a real product. It is a generic assistant that drafts work for an office's professional staff, within a scope set by professional standards. It runs on a vendor's host, and the vendor can update it.
What kind of office this is
This office is modeled, not real. The simulation behind this case models it on a social work agency. It has strong oversight, rules on how long records are kept, active audit and re-checking, and investment in records that can be verified.
Four guardrails define it. Staff have time budgeted for checking the assistant's drafts. Staff write to the case records through a write gate, a check each entry passes before saving. Every record carries a label showing where it came from.
The fourth guardrail is that someone with authority reviews the deployment on a schedule. The network does not draw that reviewer as a separate part. Here, Require sign-off adds that authority.
Output checks also test each draft for scope and consistency before staff see it.
Three offices, one assistant
The Lab draws the same assistant in three modeled offices, to compare them. In the agentic office it acts with little oversight. In the supervised office, workers check its work. This is the third office, and the most governed.
The assistant is the same in all three. What differs is the oversight around it, so any difference between the offices comes from how each one is governed.
The two pressures in this case
Two pressures that even strong offices face are on from the start. First, staff quietly paste case details into consumer AI tools that nobody vetted. Second, the vendor ships a model update that shifts the assistant's behavior under controls tuned to the old version.
Because the assistant runs on the vendor's host, that update is also a second way client data can leave the office's boundary.
What this case teaches
The lesson is diminishing returns: each added control buys less in an office already this well governed. The case also shows what it costs to keep controls from drifting.
Where the facts come from
No real deployment backs this office, so it has no case file. Its shape comes from the simulation's published comparison of the same AI in three office cultures.
The survey findings in the readings come from a 2025 to 2026 national survey of social workers in the United States. The University of Texas at Austin ran it with NASW, the National Association of Social Workers.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. On the three pathways that leave the office, what passes on is client data. A tool closes a pathway when mistakes, or client data, stop passing along it. The work along it goes on.
This case has a budget of 6 units. Before any tool is used, three failure pathways are open: Staff entries through write gate, Client data pasted into outside tools, and Assistant output to vendor host. The last two also drain the Privacy gauge, which falls while client data can leave the office.
Require sign-off costs 1 unit, or 3 at its stronger setting. Review on schedule costs 2, or 4 at its stronger setting. Keep skills sharp, Gate vendor updates, Keep prompts neutral, and Assign a challenger cost 2 each, or 3 each at their stronger settings. Route more work through the assistant and Let it keep working records cost 2 each, or 4 each at their stronger settings. Upgrade model and Store less data cost 3 each, or 5 each at their stronger settings. Vet connections costs 3 and has no stronger setting.
Understand the system costs 3 under Explore (No Targets) and Service Targets Only, and 4 under the two higher levels. Its stronger setting costs 6. While it is on, Upgrade model, Assign a challenger, Require sign-off, and Gate vendor updates each cost 1 unit less, or 2 less at its stronger setting. No tool ever costs less than 1 unit.
Lingering effects is a Dynamics setting in which damage outlasts its cause. It is on by default, and always under All Governance Targets. With it on, the network starts at a tipping point, not self-correcting.
Explore (No Targets) sets no targets. There, and under Service Targets Only, any one of five tools costing 2 units makes the network self-correcting. They are Keep skills sharp, Review on schedule, Keep prompts neutral, Gate vendor updates, and Assign a challenger. Under Service Targets Only, each also meets the service target, which asks that the assistant be helping the work. Many other combinations within the budget do too. With lingering effects off, the targets are met before any tool is used.
Under Service and Safety Targets, the targets can be met. That level also asks you to close every failure pathway, refill the Privacy gauge, and have staff keeping up with the work. Lingering effects is off there. The cheapest combination costs 5 of the 6 units: Store less data with Keep skills sharp. Store less data closes all three failure pathways, and Keep skills sharp lets staff keep up with the work. Every combination that meets these targets includes both. The others add Require sign-off, or use Keep skills sharp at its stronger setting.
Under All Governance Targets, the targets are not fully addressable with the available tools. That level also asks that the assistant be clearly helping the work. Store less data is the one tool offered that closes Staff entries through write gate, and it trims the help the assistant gives. No combination that keeps the network self-correcting and closes every failure pathway leaves the assistant clearly helping. The best leaves it helping, but not clearly.
More money does not change that. The Lab tried every combination of the twelve tools and their settings, at any cost, and none meets every target.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the High-governance professional office network: 5 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 4 assumptions
- baseline
How widely mistakes take hold among staff in this office comes from the simulation behind this case. No real office was measured.
- assumed
Staff consult a colleague before high-stakes calls, following the professional norm of seeking advice on difficult ones. How much this holds mistakes back is assumed.
- assumed
The assistant is assumed to make the same kinds of mistakes, just as often, in all three modeled offices. Only the oversight around it differs.
- assumed
The office is assumed to follow its professional-standards controls as designed, with no drift at the start.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
In the sociotechnical simulation, the same AI in three modeled office cultures, stylized and not real workplaces, led to very different outcomes. Mistakes built on one another far more under low-oversight autonomy than under human supervision or high-governance professional controls.
scenarioillustrative PAN-run resultNo published source is attached to this claim yet.
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Caseworker documentation & copilots domain page.
Levers available here and the patterns behind them
- Upgrade model — Improve the model
- Review on schedule — Oversight cadence & retrospectives
- Keep skills sharp — Deskilling-arrest mandate
- Understand the system — Understand the system
- Require sign-off — Conformity assessment gate
- Gate vendor updates — Vendor quality gate
- Keep prompts neutral — Framing and mirroring reduction
- Vet connections — Connection authorization
- Store less data — Data minimization
- Assign a challenger — Structured dissent
Documented case histories
- Magic Notes (Beam)
- Minute / Local Transcribe
- Massachusetts DTA call summaries
- Justice Transcribe
- Illinois DCFS Augintel
- GDS Microsoft 365 Copilot cross-government experiment
- NJ AI Assistant
- DWP Whitemail Insights and Vulnerability Scanner
- UK Home Office asylum AI copilots: interview summarisation and policy search
- Learned Hand AI clerk pilot (LA and Riverside courts)
- SSA Insight
- CDTFA Axyom Assist
- VA claims automation (automated survivor-benefit decisions)
- Trelleborg's Welfare Robot
- Amsterdam Smart Check