Skip to content

PAN Lab example

Frida (NAV Norway)

Ask for a human: the handover boundary as a governed surface

Frida is the chatbot of NAV, Norway's welfare agency. The case asks how easily citizens reach a person and who reads chats ending without one.

See more

Frida is the chatbot at the front of the anonymous chat on nav.no, run by the Norwegian Labour and Welfare Administration (NAV). It matches each question to a scripted answer about welfare rules. It decides no case, and no per-answer error rate has been published.

How it is used

A citizen types a question in the chat on nav.no. Frida answers first, 24 hours a day. On weekdays from 9:00 to 15:00, the citizen can ask Frida for a human advisor.

NAV says the chat is anonymous and it can see no personal information. Frida gives general guidance, not decisions. It writes to no individual case record.

Frida runs on a commercial no-code platform from the vendor boost.ai. A NAV team of about six non-technical staff, whom the vendor calls AI trainers, keeps the scripted answers current. NAV states Frida launched in summer 2018, and it still runs as of 2026.

The surge that made it a research subject

During Norway's COVID-19 lockdown, NAV faced a surge in citizen inquiries of about 250 percent. The vendor's case study says Frida answered more than 270,000 coronavirus-related inquiries. The NAV-funded Frida@work report records nearly 11,000 inquiries in Frida on some days between March and May 2020.

At the peak, in week 13 of 2020, the volume equalled the capacity of about 220 human advisors. That is the figure in the vendor's case study and in a peer-reviewed journal study. The Frida@work report says 230. The sources leave this difference unresolved.

The headline figures come from the vendor's marketing case study and are its claims. They are the 270,000 inquiries, the 80 percent resolved without a person, the 220 advisors, and the 250 percent surge. The Frida@work report independently supports the surge and the one-in-five transfer rate.

Who decides how easy it is to reach a person

About four in five conversations end without a person. That is a completion rate, not an accuracy rate.

The Frida@work project found that about one in five conversations moved to a person when citizens chose freely between Frida and human chat. When NAV removed that explicit choice, about 30 percent moved. The two figures come from different setups of the chat, so they are not a clean before-and-after comparison.

The citizen asks for a person. NAV decides how visible that option is.

What goes wrong inside the conversations

Independent chat-log studies found irrelevant answers and omitted important information. One study found three kinds of failure: gaps in vocabulary, uncertainty about which rules apply, and citizens misreading the rules. The most critical failures happened when a misunderstanding went undetected.

The most critical of them happen inside conversations that end without a person. Nothing in those conversations prompts the citizen to ask for one.

What the advisor receives at the handover

When a conversation moves to an advisor, Frida's transcript is handed over with it. The Frida@work project found that context survives the handover imperfectly. Citizens are often unsure whether they are now chatting with a person, because the chat window looks nearly the same.

A University of Agder thesis found advisors developed their own ways of using the chatbot, beyond its intended use. Frida@work recommended a separate internal chatbot to help chat staff at the handover. In autumn 2024 NAV piloted an internal AI assistant for its advisors. In effect, it follows that recommendation.

Who oversees it

Frida is unusually well studied. Its oversight includes a study NAV commissioned, a three-university project NAV funded, independent chat-log studies and theses, and NAV's own analysis unit.

These studies produce knowledge about Frida. None of them is a per-answer accuracy audit, and none is a step Frida's answers must pass.

Does it reduce the workload

NAV's analysis unit published a channel-use study in 2025. The chatbot icon was prominent on nav.no from autumn 2020 until it was hidden about two years later. Contact-centre inquiries have fallen steadily since 2019, apart from the early pandemic.

The study found the chatbot's visibility appears not to change contact-centre inquiry volumes. It credits the decline to self-service improvements, new application systems, SMS notifications, and changed contact-centre practices. So the sources do not show Frida reducing human workload outside the crisis peak.

The study also found that chatbot use goes together with use of the nav.no search engine. Average call length rose.

What this network is drawn from

This network follows the handover pattern the case file documents. It is not a reconstruction of Frida itself. It shows Frida, its library and chat logs, the citizens, the chat advisors, the AI trainers, and the research around it. What an answer means for the person who acts on it lies outside the network.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A pathway is closed when mistakes stop passing along it. The work along it goes on.

This case has a budget of 8 units. Each tool costs the same at every target level.

Explore (No Targets) sets no targets. Service Targets Only asks that mistakes die out rather than build on one another, and that Frida helps the work. Before any tool is used, Frida helps the work, but mistakes are copied about as fast as they are corrected.

Under Service Targets Only, five tools meet the targets on their own. Escalate checks and Mark AI-written records each do it for 2 units, and Gate record entries for 3. Store less data and Upgrade model do it at their stronger settings, for 5 each. In all, 203 distinct sets of tools within the budget meet these targets.

Service and Safety Targets and All Governance Targets both ask you to close every failure pathway, among other targets. Before any tool is used, five are open. Three are Citizen takes Frida's answer, Scripted answer looked up, and Trainers read the chat logs. The other two are Trainers update the answers and Chats logged for the trainers.

The targets can be met under both levels, with three tools. Escalate checks closes Citizen takes Frida's answer. When monitoring flags trouble, it has people check Frida's answers more closely. The sources name no such monitoring at NAV.

Mark AI-written records closes Scripted answer looked up and Trainers read the chat logs. Here it marks Frida's side of the logged chats as machine-written, so people reading the chats can weigh it.

Gate record entries or Store less data closes Trainers update the answers and Chats logged for the trainers. The first requires sign-off before anything enters the library. The second writes and keeps less there.

So two sets of tools meet these targets. Each costs 7 units, or 8 with Mark AI-written records at its stronger setting. No other combination within the budget meets them. Escalate checks is the one tool on offer that closes Citizen takes Frida's answer.

More is not better here. Every tool at its strongest setting at once costs 41 units, about five times the budget. It closes every failure pathway, but Frida then no longer helps the work enough. It misses the targets at every level that sets them.

Stylized model of a documented deploymentBenefits navigation & public-facing chat

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Frida-class chatbot handover boundary network: 6 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 7 assumptions
  • assumed

    This network follows the handover pattern the Frida case file documents. It does not reconstruct the chatbot itself. In the period the sources describe, Frida matches questions to scripted answers on a commercial platform and does not generate text. It gives anonymous general guidance and makes no benefit decision. So do not read it alongside the cases where a system decides eligibility or scores fraud. Some chatbots in this Lab answer a caseworker, who checks each answer before passing it on. Frida has no one in that seat. A person is reached only when the citizen asks for one.

  • baseline

    The network treats the handover to an advisor as something NAV governs. The share of conversations moved to a person is one of the few figures here that is both measured and set by governance. About one in five moved when citizens chose freely, and about 30 percent when NAV removed the explicit choice. Each figure belongs to its own setup of the chat. The network follows the free-choice setup. The case file records that NAV, not the citizen, sets how visible the human option is. The peak-load figures differ across the vendor, a journal study, and the NAV-funded report. They put week 13 of 2020 at the capacity of about 220 or 230 advisors. The network leaves that difference unresolved.

  • baseline

    The network treats a conversation that ends without a person as finished, not as answered correctly. About 80 percent end that way, but that is a completion rate. No per-answer error rate has been published. Chat-log studies found irrelevant answers, omitted information, and three kinds of failure in knowledge of the rules. The most critical failures happened when a misunderstanding went undetected inside a conversation Frida completed. So the network includes a per-answer accuracy audit, a check the sources do not describe in use, even among the many studies of Frida. Check with a second model is the tool here that adds it.

  • baseline

    The network assumes the transcript handed to the advisor carries the conversation's context imperfectly. That comes from the NAV-funded three-university project. It found citizens moved to a person are often unsure whether they are chatting with a person or a machine. A University of Agder thesis found advisors developed their own workarounds at the handover. NAV in effect adopted the project's recommendation of an internal chatbot for chat staff. It piloted an assistant for advisors in autumn 2024, a separate system. Frida writes no new answers and no case record. The sources describe the trainers updating the scripted answers from live conversations. The network still counts that loop, and the logging, as places a mistake could be passed on.

  • assumed

    How often Frida's answers are wrong is a modeling choice here, not a measured rate. No per-answer accuracy rate has been published, and the roughly 80 percent figure is a completion rate. The vendor's figures come from its marketing case study and are labelled as its claims. They are more than 270,000 inquiries, 80 percent resolved without a person, work equal to 220 full-time staff, and a 250 percent surge. The choice reflects two facts. No professional stands between Frida and the citizen, and the signature failure is the undetected misunderstanding. Also, in the period described, Frida gives scripted answers and does not generate text.

  • assumed

    The network assumes mistakes pass between peers in two ways. One public chatbot answers everyone from the same library, so a wrong answer repeats rather than scatters. The advisors' own handover workarounds spread across the NKS chat team. The network also draws the research around Frida as a part that studies the chatbot but approves nothing. These are modeling assumptions, not measurements.

  • assumed

    The anonymous citizens who use Frida appear here only as people who take up its answers. The network computes no benefit, harm, or other outcome for any person. The sources measure no difference in how groups of citizens are served, and no error rate by topic, such as benefits questions. So the network models none, and says so in words instead. NAV's own 2025 analysis found the chatbot's visibility appears not to change contact-centre inquiry volumes. It credits the steady decline since 2019 to self-service improvements, new application systems, SMS notifications, and changed contact-centre practices. So this example does not present Frida as reducing human workload outside the crisis peak. What an answer means for the person who acts on it is documented in the case file, outside this network.

What this example does not show

Show all 3 limitations
  • Frida gives anonymous general guidance and makes no benefit decision. It does not decide eligibility, entitlement, or a sanction, so the appeals and overrides of the decision cases do not apply here. Do not read it alongside the eligibility or fraud-scoring cases. This example shows citizens only as people who take up Frida's answers. It computes no benefit, harm, or other outcome for any person. What an answer means for the person who acts on it is documented in the case file, outside this network.
  • The headline pandemic figures come from the platform vendor's marketing case study and are its claims. They are more than 270,000 inquiries, 80 percent of conversations resolved without a person, a peak equal to about 220 full-time staff, and a 250 percent surge. The NAV-funded research report independently supports the surge and the one-in-five transfer rate. It gives 230 advisors for week 13 of 2020, where the vendor and a journal study say 220. This example leaves that difference unresolved. The 80 percent figure is a completion rate, not an accuracy rate, and no per-answer error rate has been published. The transfer figures depend on the setup of the chat. About one in five moved under free choice, and about 30 percent when NAV removed the explicit choice. Any handover figure has to say which setup it describes.
  • The sources contain no independent audit of Frida's answer accuracy and no measured difference between groups of citizens. So this example does not show effects on fairness. The research around Frida is unusually deep, but it studies the tool rather than approving its answers. NAV's own 2025 channel-use analysis found the chatbot's visibility appears not to change contact-centre inquiry volumes. It credits the steady decline since 2019 to self-service improvements, new application systems, SMS notifications, and changed contact-centre practices. So this example does not present Frida as reducing human workload outside the crisis peak. Two master's theses behind this case were checked through Norway's national research archive record, not read in full, because the universities' repositories moved platforms. One chat-log study calls the chatbot by a made-up name, Anna, while clearly identifying NAV's chatbot.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • Frida is the chatbot at the front line of the Norwegian Labour and Welfare Administration's (NAV) anonymous contact-center chat channel; NAV states it launched in summer 2018 and, as of 2026, that citizens first meet Frida (open 24 hours a day) and can ask it for a human advisor on weekdays between 9:00 and 15:00, with the channel anonymous and no personal information visible to NAV. During the COVID-19 lockdown NAV reported a roughly 250 percent surge in inquiries; the platform vendor's case study reports the chatbot answered more than 270,000 coronavirus-related inquiries and that about 80 percent of enquiries were resolved without escalating to a human, and NAV's own funded research report records nearly 11,000 inquiries in Frida on some days between March and May 2020 with a week-13-2020 peak equal to the capacity of about 230 human advisors, where the vendor and the peer-reviewed EJIS study state about 220. These pandemic figures originate substantially in the vendor's marketing case study and are reported here as vendor claims with the 220-versus-230 source tension left unresolved; the roughly 80 percent containment is a completion or non-escalation rate, not a measure of answer accuracy.

    empirical
    • Vendor boost.ai (vendor), How conversational AI is helping Norway's citizens through COVID-19 (NAV case study, 2020) https://boost.ai/case-studies/how-conversational-ai-is-helping-norways-citizens-with-covid/
    • Government evaluation Parmiggiani, Farshchian, Vassilakopoulou, Pappas, Grisot, Frida@work: forskningsprosjekt om betydningen av tillit i bruken av chatboten Frida i NAV (project report for NAV, NTNU / University of Agder / University of Oslo, 2021) https://www.nav.no/_/attachment/download/a9ba64cd-c8ee-4177-a0e7-e1c5ae2749d1:1563c75472fae37937be5b64d6a796569fe35160/Frida@work_sluttrapport.pdf
    • Academic Vassilakopoulou, Haug, Salvesen, Pappas, Developing human/AI interactions for chat-based customer services: lessons learned from the Norwegian government (European Journal of Information Systems, 2022 online first; print 2023, 32(1)) https://www.tandfonline.com/doi/full/10.1080/0960085X.2022.2096490
    • Government NAV, Contact us (nav.no chat section, 2026) https://www.nav.no/kontaktoss/en
  • The best-documented property of NAV's Frida chatbot is its chatbot-to-human handover boundary, which the evidence suggests behaves as a governance-controlled dial: NAV's funded three-university Frida@work project reports that about one in five conversations transferred to a live human advisor under free channel choice, and only about 30 percent of dialogues transferred when NAV removed the explicit choice between the chatbot and human chat, a regime-specific figure that must be read against the interface in force. Independent chat-log studies document irrelevant answers, omitted information, and three classes of domain-knowledge failure, with the most critical failures occurring when a misunderstanding goes undetected inside a conversation the chatbot completed; no per-answer accuracy or error rate has been published, and the Frida@work project found context survives the handover imperfectly, with citizens often unsure whether they are talking to a person or a machine. NAV's own 2025 channel-use analysis found that chatbot visibility appears not to change contact-center inquiry volumes and attributes the steady post-2019 decline to a bundle of causes (self-service improvements, new application systems, SMS notifications, and changed contact-center practices), so the chatbot is not shown to reduce human workload outside the crisis peak.

    empirical
    • Government evaluation Parmiggiani, Farshchian, Vassilakopoulou, Pappas, Grisot, Frida@work: forskningsprosjekt om betydningen av tillit i bruken av chatboten Frida i NAV (project report for NAV, NTNU / University of Agder / University of Oslo, 2021) https://www.nav.no/_/attachment/download/a9ba64cd-c8ee-4177-a0e7-e1c5ae2749d1:1563c75472fae37937be5b64d6a796569fe35160/Frida@work_sluttrapport.pdf
    • Academic Verne, Steinsto, Simonsen, Bratteteig, How Can I Help You? A chatbot's answers to citizens' information needs (Scandinavian Journal of Information Systems, 2022, 34(2)) https://aisel.aisnet.org/sjis/vol34/iss2/7/
    • Academic Simonsen, Steinsto, Verne, Bratteteig, I'm Disabled and Married to a Foreign Single Mother: Public Service Chatbot's Advice on Citizens' Complex Lives (Electronic Participation, ePart 2020, Springer LNCS) https://link.springer.com/chapter/10.1007/978-3-030-58141-1_11
    • Government evaluation McVey, Chatboten Frida og utvikling i kanalbruk hos Nav (Arbeid og velferd nr. 2-2025, NAV analysis journal) https://www.nav.no/no/nav-og-samfunn/kunnskap/analyser-fra-nav/arbeid-og-velferd/arbeid-og-velferd/arbeid-og-velferd-nr.2-2025/chatboten-frida-og-utvikling-i-kanalbruk-hos-nav

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Benefits navigation & public-facing chat domain page.

Levers available here and the patterns behind them

Documented case histories