Skip to content

PAN Lab example

The same AI with a human checking: the supervised office

In this made-up office, staff review an AI assistant's drafts before anything enters the case records. The risk is that review slowly becomes a formality.

See more

The AI here is an assistant in a made-up example office, not a real product. Staff ask it questions, and it drafts answers and suggestions for them to review. One assistant serves the whole team, and when staff ask, it quotes existing case records into new drafts.

What this example is

This office is made up for illustration. It is one of three in the Lab built around the same AI. The Lab is built on the sociotechnical simulation, a computer simulation of how people and AI tools work together. It runs the same assistant in three made-up office cultures. Only the oversight around the AI differs between them.

The simulation's guidance calls this office the everyday baseline, where workers can and do check the AI.

The other two offices

In oversight, this office falls between the Lab's other two office examples. In the low-oversight office, the same AI also acts on cases as an agent, and a stretched staff waves most of its work through. In the professional office, the same AI works under full guardrails. Checking time is budgeted, and entries to the record pass a gate. Records are labeled with their source, and someone with authority reviews the deployment on a schedule.

Who decides

Staff decide what enters the case records. The assistant drafts and suggests, and a person reviews its output before acting on it. The office is set up so that the assistant makes no entries of its own.

Where the risk lies

Review is this office's main safeguard, and it is only as good as the judgment behind it. The simulation's guidance describes how that judgment wears down. When an AI is usually right, checking slowly becomes a formality, then gets skipped on easy cases. Nothing breaks, and people check less over time.

So the failure here is slow. Review keeps happening while the judgment behind it fades.

Mistakes also have routes besides the review. Staff share adopted answers with colleagues in team chat. How staff frame a question can nudge the assistant toward the answer they expect. Because one assistant serves the whole team, a mistake it makes once can repeat for everyone who asks.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. When a tool closes one, mistakes stop passing along it, though the work along it can go on.

Two failure pathways are open when the network opens: Drafts reviewed and adopted by staff, and Staff enter reviewed content.

The Lab has four target levels. Explore (No Targets) sets no targets. Under Service Targets Only, the network meets the targets before you use any tool, and more than 400 combinations of tools also meet them. Two tools put more of the work on the assistant: Route more work through the assistant and Let it keep working records. Used alone, either one breaks the targets, because mistakes are then harder to contain.

Under Service and Safety Targets and All Governance Targets, the targets can be met. Both levels ask you to close every failure pathway, among other targets. Every combination that meets them uses Escalate checks and Store less data together. Escalate checks closes Drafts reviewed and adopted by staff, and Store less data closes Staff enter reviewed content.

That pair is the cheapest way, at 5 of this case's 8 budget units. Ten different sets of tools meet the targets at each of these two levels. Counting each tool's stronger setting as a separate choice, there are sixteen. Understand the system costs 4 units at these two levels, and no set that meets the targets includes it.

No pathway here sends client data out of the office, so nothing drains the Privacy gauge, the Lab's reading of how well client data is kept safe.

Comparative teaching networkCaseworker documentation & copilots

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Human-supervised office network: 3 components and 9 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 4 assumptions
  • baseline

    This office is made up. It follows one of three office cultures in a computer simulation of people working with AI. Nothing real was measured.

  • assumed

    Staff pass answers to each other and compare notes on odd ones. The shared assistant's mistakes are assumed common to the whole team, not measured.

  • assumed

    The assistant is equally prone to mistakes in all three office examples. Only the oversight around it differs.

  • assumed

    Review quality is assumed steady over time. The slow loss of reviewing skill is left out when the network opens.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • In the sociotechnical simulation, the same AI in three modeled office cultures, stylized and not real workplaces, led to very different outcomes. Mistakes built on one another far more under low-oversight autonomy than under human supervision or high-governance professional controls.

    scenarioillustrative PAN-run result

    No published source is attached to this claim yet.

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Staff turnover — Experienced skepticism leaves; new staff calibrate their trust on the tool itself.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Caseworker documentation & copilots domain page.

Levers available here and the patterns behind them

Documented case histories