PAN Lab example
Albert France Services
Killed without a number: a sovereign adviser assistant that no metric ever measured
Albert, a French state AI assistant, drafted answers for France Services advisers in a pilot. No error rate was published, and it was not extended.
See more
Albert France services was a generative AI assistant piloted at France Services, France's one-stop public service counters. DINUM, the French state's digital directorate, built it with ANCT, the national agency for territorial cohesion. Advisers typed a question about a citizen's procedure or benefits, and it drafted an answer from official documents, citing its sources, for them to check.
How it was used
France Services is France's network of one-stop counters. There, advisers help citizens with administrative procedures and benefits. The network has more than 2,750 counters, assisting nearly 800,000 procedures a month.
Albert searched a curated base of documents. They were practical sheets from service-public.fr, the government's public information site, and documentation from the national operators France Services represents. These cover taxes, pensions, family benefits, health insurance, and other services. Practical sheets are the site's information pages on each procedure. Each answer came with related questions, practical links, and simulators, which are online calculators.
Albert was advisory and decided nothing. The adviser was expected to verify, modify, and validate every answer before relaying it. The adviser could also ignore the tool entirely, and many reportedly did, preferring an ordinary search.
The pilot
The pilot began with a panel of about sixty volunteer advisers. It grew to roughly eighty advisers at more than forty counters in six departments, France's administrative areas: Vienne, Deux-Sèvres, Rhône, Allier, Meurthe-et-Moselle, and Var. AFP, the French news agency, counted forty-eight counters in its January 2026 reporting.
A joint DINUM and ANCT experiment team produced three successive versions from advisers' feedback journals and qualitative interviews.
The launch and the error
DINUM presented Albert as the State's free and sovereign generative AI, created by and for public agents. Sovereign here means under French state control. It won the innovation prize at the Victoires des Acteurs publics, a public sector awards event, on 7 February 2024.
On 23 April 2024, Prime Minister Gabriel Attal launched it at the France Services counter in Sceaux. He called it a 100 percent sovereign French AI that would revolutionize public services. He promised that AI would free agents for higher-value work, not replace them. The sovereignty descriptions are the government's own.
In a demonstration before the Prime Minister, Albert answered that renewing an identity card is free. The cost that applied in the case shown was 25 euros. The error became emblematic of the tool's reliability problems. The sources do not date it to the launch itself.
What was never measured
No error rate, usage volume, adoption frequency, or override count was ever published for Albert France services. An override is an adviser setting aside the tool's answer. The sources describe no instrument that measured whether its answers were right.
How its failures came to light
Its failures came to light through the people who used it. Several unions documented recurring technical malfunctions and plainly wrong answers. Advisers reported that it often answered less well than an ordinary search. According to the union Solidaires Finances Publiques, an investigative television broadcast in April 2025 featured unenthusiastic testimony from France Services agents.
The case file reads this as the case's lesson. When nothing measures a tool, the people using it become the error detector of last resort. That detector works slowly, in public, and late.
How it ended
According to Solidaires Finances Publiques, the project had stopped by September 2025, with no announcement. The union says it learned this in a ministerial AI strategy working group, where Albert no longer appeared among the projects presented. It puts the project's cost at about 1.3 million euros. It calls the project a top-down failure, built without consulting advisers. These are the union's claims.
On 9 January 2026, DINUM announced that Albert, as tested at 48 counters, would not be generalized in its current form. Generalized means extended across the whole network. DINUM cited the pilot's record. It also disputed the failure framing. It said most of the experimental projects grouped under the Albert name are sustained and fully operational. Albert is also the name of the state's wider AI program, described below.
AFP reported DINUM's annual AI budget at about 1.2 million euros since 2024, with Albert France services a minimal share. That figure is not directly comparable to the union's cost claim.
What came next
The wider Albert program continues as shared government infrastructure. It includes the Albert API, an access point to AI models for ministries, and Albert Data, which gathers public reference bases. The state's successor tool is Assistant IA, French for AI assistant. It grew out of a conversation product in the wider Albert program. It is a newer adviser tool for agents across ministries, tested with about 10,000 of them through June 2026. It uses models from the French vendor Mistral AI.
A full evaluation of that test is due in summer 2026. It must notably establish what extending the tool across government would cost. As of mid-2026 it had not been published.
Alongside the January 2026 decision, DINUM removed the web-search function from the Albert API. It retired the old Albert-branded model names by 15 February 2026.
What this case asks
The usual question is which control failed. Here the question is what to build when nothing was watching. The case file points away from the assistant itself. It names a review rhythm, an outside bar a version must clear before launch, and a standing check inside the system.
What this network is drawn from
This network follows the pattern the case file describes. It does not reconstruct the actual tool or any of its versions. It shows the assistant, its search, the knowledge base, the advisers, the experiment team, and the union and press reports. Citizens are outside the network.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A closed pathway is one that mistakes stop passing along. The work along it may go on.
This case has a budget of 8 units. Explore (No Targets) sets no targets. Under Service Targets Only, the targets are met before any tool is used or pressure added. Every combination of tools within the budget keeps them met.
Service and Safety Targets and All Governance Targets also ask you to close every failure pathway. Before any tool is used, one is open: Draft answers to advisers. The targets can be met at both levels within the budget.
Escalate checks is the only tool on offer that closes that pathway. Once trouble is flagged, it has advisers check the assistant's answers more closely against the sheets they cite. The sources describe no monitoring on this deployment to flag it. On its own it meets the targets at both levels for 2 units. Every combination that meets them includes it.
All Governance Targets also asks for the assistant to be clearly helping the work. A few combinations with Escalate checks fall short of that. One is Review on schedule, Understand the system, and Escalate checks, which costs all 8 units.
Understand the system costs 3 units under Explore (No Targets) and Service Targets Only, and 4 under the two higher levels. Its stronger setting costs 6 units. While it is on, Vet connections, Keep prompts neutral, Keep skills sharp, and Upgrade model each cost 1 unit less. At its stronger setting they cost 2 units less, never below 1.
More is not better here. Every tool at its strongest setting costs 31 units, nearly four times the budget. That meets the targets at none of the three levels that set them, because the assistant then adds too little to the work.
No tool on offer adds the check this deployment lacked: an error measurement or independent accuracy audit before each version.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Albert-class sovereign adviser assistant (deployment-and-withdrawal) network: 7 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 7 assumptions
- assumed
This network follows the pattern the Albert France Services case file describes. It does not reconstruct the actual tool or any of its three versions. It is a companion to the Lab's two US benefits chatbots for caseworkers. Nava Labs built both: its earlier chatbot and the Benefit Navigator piloted with Imagine LA. In those, a caseworker stands between the tool and the person served, and the risk is trusting the tool too much. Albert was a French state pilot whose advisers trusted it too little and abandoned it. DINUM decided not to extend it in its current form. It is also close to the opposite of GOV.UK Chat, the UK government's public chatbot. That rollout turned on a published, if self-reported, accuracy figure. This deployment published no error figure at all.
- baseline
The case's defining gap is the independent check it never had. The sources record no error measurement or independent accuracy audit that a version had to pass before release. No error rate, override count, or usage figure was ever published. So its failures came to light through the people using it. Unions documented malfunctions, advisers grew disaffected, and it gave one wrong answer on identity card cost in a demonstration before the Prime Minister. The network draws that check to show what was missing. No tool offered in this case adds it.
- baseline
The network assumes the workforce was the last line of error detection. It draws the DINUM and ANCT experiment team as a way to correct the tool. The team had access to it and produced three versions from qualitative feedback, and the sources describe no error measurement by it. The network draws the union and press reports as the detection that actually happened. Those reports gathered advisers' disaffection and the demonstration error into the reputation that framed the tool. The case's lesson is the contrast. The body placed to measure error is not recorded doing so, and people not placed to measure it did.
- assumed
The network assumes the knowledge base protects against mistakes rather than spreading them, as in the Lab's caseworker chatbots. The assistant did not write its answers into the curated official base, which the national operators maintain. Each answer was built from that base. Each pointed the adviser back to its cited official source to check before relaying. Ignoring the tool cost advisers almost nothing, which cut both ways. It kept a wrong answer from spreading when advisers ignored it. It also let use of the tool quietly collapse.
- assumed
How often the assistant makes mistakes is a modeling choice, not a measured rate, because no error rate was ever published. The network gives it somewhat more mistakes than the Lab's caseworker chatbots built on well-curated documents. That follows the sources. Advisers found it worse than an ordinary search, unions reported recurring malfunctions, and one demonstration error became emblematic. The network does not give it many mistakes, because its answers were built from a curated official base. The case file records three versions. Whether any of them improved is not worked out here.
- assumed
The network draws three assumed pathways. First, advisers pass on to each other the view that Albert answered worse than an ordinary search, and the habit of ignoring it. That spreading wears away use of the tool rather than spreading its errors, because the tool costs nothing to ignore. Second, advisers share the habit of checking answers before relaying them, which holds mistakes back. Third, one state assistant answering every counter means a wrong answer repeats rather than scatters. These are assumptions, not measurements.
- assumed
The citizens served are not part of the network. The network shows the France Services advisers as the people who use the tool. It works out no benefit, harm, or other outcome for any citizen. The sources hold no error rate, override count, or complaint data by topic, such as benefits alone. The case's whole point is that no such measurement was published. The harm in this kind of case is a wrong answer relayed to a citizen, or an abandoned tool. The case file documents wrong answers and advisers ignoring the tool, but no outcome for any citizen. The cost figures appear only in words, never in the network. The union Solidaires Finances Publiques claims the project cost about 1.3 million euros. The French news agency AFP reports DINUM's annual AI budget at about 1.2 million euros since 2024, with Albert France services a minimal share.
What this example does not show
Show all 3 limitations
- This example does not model the citizens served, or whether they completed their benefits claims or procedures. It follows how mistakes pass between parts of the deployment. The case file does not record those outcomes either. There is no error rate, override count, or complaint data by topic, such as benefits alone. The case's whole point is that no such measurement was ever published.
- The sources hold no figures on this deployment's performance. No error rate, usage volume, adoption frequency, or override count was ever published, so nothing here implies a measured error rate existed. The disputed figures are carried as claims, not facts. The cost of about 1.3 million euros and the September 2025 stop are claims by the union Solidaires Finances Publiques. The union inferred the stop because the project no longer appeared in a ministerial working group. AFP reports DINUM's annual AI budget at about 1.2 million euros since 2024, with Albert France services a minimal share. DINUM disputes the failure framing, saying most of the experimental projects grouped under the Albert name are sustained and fully operational. A full evaluation of the successor, including its cost, was still pending as of mid-2026.
- The emblematic wrong answer on identity card cost was given in a demonstration before the Prime Minister. The sources do not date it to the 23 April 2024 launch. French identity card fees vary by case. Renewing an expired card is free, while replacing a lost or stolen card costs 25 euros. The sources do not say which kind of card the case shown involved. So this example calls it a wrong answer in the case shown, not a ruling on the fee schedule. Mistral AI is named as the vendor whose models the successor tool uses. The case file names no specific model.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Albert France Services was a sovereign, in-house generative AI assistant built by DINUM with ANCT to help France Services counter advisers answer citizens' benefits and procedure questions from a curated base of official documents, presented by the Prime Minister as a sovereign French AI in April 2024 and, in a demonstration before him, giving a wrong answer on identity-card cost. Piloted from an initial panel of about sixty volunteer advisers to roughly eighty advisers across more than forty counters (forty-eight at final count per AFP) in six departments over three iterated versions, it was, per a January 12, 2026 AFP dispatch, formally not going to be generalized 'in its current form,' a decision DINUM announced on January 9, 2026 while stating that the majority of Albert-brand projects are sustained and fully operational. No error rate, usage volume, or override count for the tool was ever published; AFP reports DINUM's annual AI budget at about 1.2 million euros since 2024 with Albert France Services a minimal share, a figure distinct from and not directly comparable to the union Solidaires Finances Publiques' separate claim of a roughly 1.3 million euro project cost.
empirical- Trade press Weka.fr (AFP dispatch), Albert, l'outil d'IA generative, experimente a France Services ne sera pas generalise (2026) https://www.weka.fr/actualite/administration/article/albert-l-outil-d-ia-generative-experimente-a-france-services-ne-sera-pas-generalise-209194/
- Trade press Acteurs Publics, Derriere l'echec mediatique d'Albert, un projet d'IA plus global qui s'ancre dans l'Etat (2026) https://acteurspublics.fr/articles/de-chatbot-experimental-a-socle-interministeriel-pour-lia-de-letat-le-parcours-dalbert-ia/
- Advocacy Solidaires Finances Publiques, Entre ici Albert, au pantheon des IA souveraines (2026) https://solidairesfinancespubliques.org/le-syndicat/dossiers/ia-a-la-dgfip/7192-albert-france-service.html
- Government France services / ANCT, Experimentation d'un modele d'assistance aux conseillers France services base sur l'intelligence artificielle (2024) https://www.france-services.gouv.fr/actualites/experimentation-dun-modele-dassistance-france-services-IA
For Albert France Services no instrumented error-detection channel existed: no error rate, override count, or usage figure was published during the pilot, and the failures that framed the tool surfaced through the operator side, with several unions documenting recurring malfunctions and wrong answers, an investigative-television broadcast in April 2025 (per Solidaires Finances Publiques) featuring unenthusiastic agent testimony, and advisers reporting answers worse than an ordinary search. According to Solidaires Finances Publiques the project had in fact stopped by September 2025 with no announcement, inferred from Albert no longer appearing among projects presented in a ministerial working group (a union claim). Alongside the January 2026 non-generalization decision DINUM migrated the Albert API model aliases off the 'albert-' branding and removed the web-search functionality, retiring legacy aliases by February 15, 2026, while a successor adviser tool that integrates models from the vendor Mistral AI was in test with about 10,000 public agents through June 2026, gated by a summer-2026 evaluation that must notably establish the cost of a generalization.
empirical- Advocacy Solidaires Finances Publiques, Entre ici Albert, au pantheon des IA souveraines (2026) https://solidairesfinancespubliques.org/le-syndicat/dossiers/ia-a-la-dgfip/7192-albert-france-service.html
- Trade press Next (next.ink), Albert: l'IA souveraine de la Dinum ne sera pas generalisee dans sa forme actuelle (2026) https://next.ink/brief-article/albert-lia-souveraine-de-la-dinum-ne-sera-pas-generalisee-dans-sa-forme-actuelle/
- Trade press Weka.fr (AFP dispatch), Albert, l'outil d'IA generative, experimente a France Services ne sera pas generalise (2026) https://www.weka.fr/actualite/administration/article/albert-l-outil-d-ia-generative-experimente-a-france-services-ne-sera-pas-generalise-209194/
- Trade press Acteurs Publics, Derriere l'echec mediatique d'Albert, un projet d'IA plus global qui s'ancre dans l'Etat (2026) https://acteurspublics.fr/articles/de-chatbot-experimental-a-socle-interministeriel-pour-lia-de-letat-le-parcours-dalbert-ia/
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Benefits navigation & public-facing chat domain page.
Levers available here and the patterns behind them
- Review on schedule — Oversight cadence & retrospectives
- Understand the system — Understand the system
- Escalate checks — State-feedback vigilance
- Keep skills sharp — Deskilling-arrest mandate
- Keep prompts neutral — Framing and mirroring reduction
- Peer sharing rules — Peer-edge governance
- Vet connections — Connection authorization
- Store less data — Data minimization
- Upgrade model — Improve the model
- Gate vendor updates — Vendor quality gate
Documented case histories
- Albert France Services
- Nava assistive benefits chatbot
- Caddy adviser copilot at Citizens Advice
- GOV.UK Chat
- Mass.gov Virtual Assistant
- Frida (NAV Norway)
- SSA 800-Number Conversational AI Assistant
- EDD Virtual Assistant
- Burokratt
- Singapore's chatbot fleet refresh: eighty scripted engines slated for retirement onto a shared LLM platform
- IRS collection chatbots: expanded and made permanent with no performance measures
- Propel in-app SNAP benefits assistant
- GetCalFresh: the nonprofit front door that carried most of California's online SNAP intake
- MyFriendBen benefits screener
- Benefits Data Trust wind-down