Skip to content

PAN Lab example

Burokratt

The network of networks: a federated public-service chatbot

Estonia's Burokratt puts many public institutions' chatbots behind one window. A mistake in its shared routing or knowledge can appear in every institution's answers.

See more

Burokratt is Estonia's national network of public-sector chatbots, run by the Information System Authority (RIA). Each institution's chatbot answers from its own knowledge base and, from 2025, from one shared module built from the state portal eesti.ee. A central classifier sends each citizen's question to the right institution's chatbot.

How it is used

A citizen types a question into one chat window. The central classifier, a program that sorts each question by the institution that should answer it, sends it to that institution's chatbot. It also oversees the handover between chatbots, which exchange messages behind the window.

When a chatbot cannot answer, it hands the live chat to a customer-service representative of that institution. Burokratt answers questions only. It makes no decision on anyone's eligibility or benefits.

How the chatbots answer

The first generation, from 2021 to 2024, worked from rules and intents: it recognised what kind of question was asked and gave the answer written for it. It was built on RASA, an open-source framework for understanding written language. It ran in the national State Cloud with the state's login service.

A 2025 programme added retrieval-augmented generation. The chatbot looks up passages in the institution's own knowledge base and in the shared module, and an AI language model writes the answer from them.

The shared module is built from public information on eesti.ee about healthcare, social benefits, pensions, transport, employment, education, and taxation. It lets the central assistant answer questions that cross institutions from one window.

How it began

Burokratt grew out of Estonia's 2019 national AI strategy, known as the kratt strategy. A vision and concept paper followed in 2020. A 2022 European Commission case study says an early version was tested in 2021 in three agencies. They were the Police and Border Guard Board, the Consumer Protection and Technical Regulatory Authority, and the National Library. The first full rollout came in 2022, beginning with the Consumer Protection agency.

How many institutions take part

How big the network is depends on how you count. RIA's own page lists 20 participating organisations. Among them are the Tax and Customs Board, the Police and Border Guard Board, Statistics Estonia, and the Health Insurance Fund. Others are the state portal eesti.ee, the National Library, and the municipality of Rae Parish.

Some of those entries are portals or sites rather than separate institutions. That may explain why trade press quoting officials reports 18 organisations integrated.

An independent study by Kaun and Manniste drew on twelve interviews with insiders from September to November 2023. It records Burokratt as first piloted in 2021 in two institutions. Eight more, including one municipality, joined near the end of 2023, for ten in total at that time.

Who runs it and what it costs

RIA runs Burokratt under the Ministry of Justice and Digital Affairs. The software is free and open source under the MIT licence. Its public code organisation holds 89 repositories, with active development through mid-2026. Institutions pay roughly 150 euros a month for State Cloud hosting, plus usage fees.

Major parts were built by the vendors Net Group, Texta, Stacc, Solita, and Microsoft. The independent study found that procurement drove frequent changes of vendor.

The European Commission's Public Sector Tech Watch entry, last updated in June 2025, still lists the project as in development. It records RIA's statement that partnerships are hard to set up. RIA says the expertise needed to develop this kind of AI together is very high.

The funding

The European Commission's Recovery and Resilience Facility project record lists a 53 million euro EU contribution. It describes that sum as being for two projects altogether, bundling Burokratt with core digital infrastructure and a move to cloud computing. So it is not spending on the chatbots alone.

The nearer figures are about 1.5 million euros spent through early 2022. Roughly 13 million euros more was budgeted over the following four years.

Interest from abroad

Estonia's chief data officer, Ott Velsberg, has described the network as interoperable in the manner of X-Road, Estonia's system for exchanging data between agencies. He has offered the open-source stack for free reuse. Belgium and Finland are named as negotiation partners, and Luxembourg has shown softer interest at conferences. Links with Finland's AuroraAI have been explored.

What each side says

Two kinds of source describe Burokratt, and they do not agree. Official pages, an Enterprise Estonia briefing, and trade press call it a Siri of public services. They report a place on a top-100 list from UNESCO and the International Research Centre on Artificial Intelligence. They frame Burokratt as rewriting how people deal with the state.

The independent academic study is more sober. It found Burokratt marketed as advanced AI while working much like a list of frequently asked questions. It found that use differs considerably by institution: minimal in one, about one-third of daily requests in another. It also found conflicts with the welfare state's value of serving everyone alike.

The study found chat conversations turned into data that could be analysed. That potential went largely unused beyond basic usage statistics. A comparison of phone and email volumes sometimes read into Burokratt belongs to a separate Swedish municipal chatbot in the same paper.

A survey experiment by Alishani and Homburg, published online in December 2025, studied citizens. Intended use rose with how useful people found the chatbot and how much they trusted the technology. Privacy concerns mattered for using it to get a service, not for getting information. Trust in government, explanations, and the amount of information given made no difference to intended use.

What the sources do not show

Session volumes, rates of handing chats to a person, and answer-accuracy figures were located in no public source. The sources found no dedicated body for evaluating the algorithms, no published evaluation framework, and no national audit report on Burokratt. No harm, error, or discrimination incident is documented. That reflects the lack of published evaluation, not evidence that nothing went wrong.

The independent academic studies and the fully public code are the main outside check. So this case shows how a network of this kind is built and governed. It is not an account of measured performance.

What this case asks

In a network like this, the links every chatbot shares carry the most risk. A stale or wrong entry in the shared module is not one institution's mistake: it can appear in every institution's answers. A classifier that sends a question to the wrong chatbot misdirects it across the whole network. Each institution's own knowledge base keeps its mistakes to its own answers.

The independent study shows Estonia built the links between institutions, the classifier, and the shared module faster than anything to evaluate them.

What this network is drawn from

This network follows the pattern the Burokratt case file describes. It is not a reconstruction of the real platform. It draws one institution to stand for all of them: its chatbot, its knowledge base, and its staff. It also draws the global classifier, the shared module, RIA's central team, and RIA's delivery oversight. Citizens who ask questions are outside the network.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. Closing a pathway means mistakes stop passing along it. The work along it goes on.

This case has a budget of 11 units. Each tool costs the same at every target level.

Explore (No Targets) sets no targets. Every other level asks that mistakes stop building on one another across the network, and that the chatbots stay useful to the work. Before any tool is used, mistakes build on one another, so no level's targets are met.

Under Service Targets Only, the targets can be met. The cheapest ways cost 2 units and use one tool: Mark AI-written records, Peer sharing rules, or Escalate checks.

Under Service and Safety Targets, you must also close every failure pathway. Seven start open. Four lead into the chatbot: Question routed to a chatbot, Trainers write chatbot content, Chatbot reads its knowledge base, and Chatbot reads the shared module. Two lead into the stores: Trainers update the knowledge base and Portal content added to shared module. One leads to staff: Chat handed to a representative.

The targets can be met, but only just. Six combinations of tools and settings within the budget meet them. The cheapest cost 10 units: Escalate checks, Keep prompts neutral, Vet connections, and either Store less data or Gate record entries. Every one includes Escalate checks and Keep prompts neutral. They are the only tools on offer that close Chat handed to a representative and Trainers write chatbot content.

Understand the system is a tool that funds ongoing study of what the deployment is really doing. This case does not offer it. Under All Governance Targets, four tools work at reduced strength without it. They are Vet connections, Store less data, Peer sharing rules, and Check copied records. Vet connections then closes no pathway. Two combinations meet the targets, each costing the full 11 units. Both use Escalate checks, Keep prompts neutral, Peer sharing rules, and Mark AI-written records, with either Store less data or Gate record entries.

Three tools close no pathway on this network: Upgrade model, Review on schedule, and Check copied records. The network draws no copying between stores for Check copied records to act on.

More is not better here. Every tool at its strongest setting costs 42 units, nearly four times the budget. It closes every pathway but misses the targets at all three levels that set them, because the chatbots then add too little to the work.

Stylized model of a documented deploymentBenefits navigation & public-facing chat

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Burokratt-Estonia-class federated public-service chatbot network network: 7 components and 14 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 7 assumed. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 7 assumptions
  • assumed

    This network follows the pattern the Burokratt case file describes: many institutions' chatbots sharing one routing classifier and one knowledge module. It is not a reconstruction of the real platform. One institution is drawn to stand for all of them. The institutions differ in how much their chatbots are used, not in how they are built. So the network describes that range instead of drawing a second institution.

  • assumed

    No published session volumes, rates of handing chats to staff, or answer-accuracy figures for Burokratt were found. So how often mistakes are made and passed on here is estimated from how the system is built and governed. It is not measured from how it runs. No harm or error incident is documented. That reflects the lack of published evaluation, not evidence that none occurred.

  • assumed

    Every institution's chatbot reads the shared eesti.ee module. So a stale or wrong entry there can appear in every institution's answers, while a mistake in one institution's own store stays in its own answers. RIA's intake of eesti.ee content is the most consequential write in the network. The sources describe shared content used in answers beside an institution's own content, not copied into its store. So the network draws no direct link between the two stores.

  • assumed

    Practices spread between institutions through RIA's central onboarding, the channel the sources name. The network draws this once, from RIA's team to the institution's staff. What institutions give back is described as promoted reuse and is not drawn separately. The one gap the sources document is drawn once, as the Cross-institution check. The sources found no dedicated body for evaluating the algorithms, no published evaluation framework, and no national audit report. They also found no check that the shared module agrees with institution content, and no peer check of answers across institutions. These belong to the same gap, so they are not drawn as further checks.

  • assumed

    The network's governance part stands for the oversight the sources show. That is RIA's oversight of the platform from demonstration to production, and the EU Recovery and Resilience Facility's milestone monitoring. It is drawn as two links: reports on delivery to governance, and a review of the central team's platform work. This is oversight of delivery, kept apart from the missing evaluation of the chatbots' answers. How much both links carry is estimated, not measured.

  • assumed

    Citizens who ask the chatbots questions are outside the network. Burokratt makes no eligibility or benefit decision. What spreads here is the quality of institutions' answers, not any outcome for a citizen. Any such outcome would be documented in the case file and measured outside this network. Use differs across institutions: about one-third of daily requests in one, minimal in another. The network draws an institution at the busier end of that range. At a little-used institution, fewer questions are routed and handed over, and trainers write and update less. Taking over chats and retraining the chatbot then get little practice.

  • assumed

    Citizens' questions pass through the classifier RIA's central team runs, and the network assumes the team reads the shared module it maintains. Neither is drawn as a link into the team. No source describes the team taking up or changing an answer as it is routed. The team shapes the network through its intake of eesti.ee content into the shared module, which is drawn.

What this example does not show

Show all 6 limitations
  • Citizens who ask the chatbots questions are not modelled here. The network shows only how mistakes spread among the institutions and their systems. Burokratt makes no eligibility or benefit decision, and nothing here computes an outcome for any citizen.
  • No published session volumes, rates of handing chats to staff, or answer-accuracy figures for Burokratt were found. No harm or error incident and no audit office report were found either. So how much each link carries here is estimated from how the system is built and governed. The lack of documented incidents reflects the lack of published evaluation, not evidence that none occurred.
  • Use varies widely across institutions and is low in some. The independent study found use differing considerably by institution: minimal in one, about one-third of daily requests in another. A comparison of phone and email volumes sometimes read into this case belongs to a separate Swedish municipal chatbot in the same paper. It is not used here.
  • The 2026 plan for a cooperative network of AI agents, one per institution, is a stated plan. So is the large language model adapted to Estonian. Neither is a working capability, and this example treats both as future plans throughout.
  • The 53 million euro EU figure is a Recovery and Resilience Facility contribution for two projects altogether. It bundles Burokratt with digital infrastructure and a move to cloud computing, so it is not spending on the chatbots alone. The nearer figures are about 1.5 million euros spent through early 2022, and roughly 13 million euros budgeted over the next four years.
  • Two kinds of source describe Burokratt, and this example keeps them apart. The promotional kind calls it a Siri of public services and reports a UNESCO top-100 listing. The independent academic kind found it working like a list of frequently asked questions, with uneven and in places low use. Nothing promotional is stated here as proven performance.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • Burokratt is Estonia's national network of public-sector chatbots operated by the Information System Authority: each participating institution runs its own assistant, a central classifier routes a citizen's query between them and oversees the handover, and from 2025 a shared knowledge module built from the eesti.ee state portal feeds cross-domain answers. RIA's page lists 20 participating organisations and trade press reports 18 integrated; an independent 2025 ethnography drawing on twelve insider interviews (conducted in late 2023, when the system spanned ten institutions) found it marketed as advanced AI while functioning much like an FAQ list, with use differing considerably by institution and low in some. No published session volumes, escalation-to-human rates, or answer-accuracy figures, and no dedicated algorithmic-oversight body, published evaluation framework, or national-audit report on the network, were located in the public record.

    empirical
    • Government Information System Authority (RIA), Republic of Estonia, Burokratt (2025) https://www.ria.ee/en/state-information-system/personal-services/burokratt
    • Trade press GovInsider, Estonia eyes cross-border interoperability for Burokratt, its Siri of public services (2025) https://govinsider.asia/intl-en/article/estonia-eyes-cross-border-interoperability-for-burokratt-its-siri-of-public-services
    • Academic Kaun, Manniste, Public sector chatbots: AI frictions and data infrastructures at the interface of the digital welfare state (New Media and Society, 2025) https://journals.sagepub.com/doi/10.1177/14614448251314394

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Benefits navigation & public-facing chat domain page.

Levers available here and the patterns behind them

Documented case histories