PAN Lab example
Frida (NAV Norway)
Ask for a human: the handover boundary as a governed surface
This chatbot is the anonymous front door to a national welfare agency: it answers pensions, child support, unemployment and sick-leave questions 24 hours a day, sees no personal information, and decides no case. Modeled on Frida, NAV Norway's citizen chatbot. The thing worth studying is not the answer but the moment a citizen can ask for a human. About four in five conversations end without one, and that number is usually reported as a success - but it is a completion rate, not an accuracy rate, and the log studies found the most dangerous conversations are the ones the chatbot finished while quietly getting something wrong, where the citizen never noticed enough to escalate. The escalation itself is a dial the agency turns, not the citizen: under free choice about one in five conversations reached a person, and when NAV removed the explicit human-chat option the figure moved to about 30 percent. The ones that do cross the boundary arrive degraded - the advisor inherits a transcript that only half survives, and the citizen is often unsure whether they are now talking to a person at all. Nothing here is decided; the guidance reaches the public directly, most of them trusting a fluent reply too much to ask for more. The question is not whether the answer is right - no one has published a per-answer accuracy rate - but who sets how easy it is to reach a human, and whether anyone reads the conversations that ended without one.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Frida-class chatbot handover boundary network: 6 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the chatbot-to-human handover-boundary pattern documented in the Frida (NAV Norway) case file - not a reconstruction of the actual chatbot. The atlas-relevant object is the boundary, not the model behind it: in the documented period Frida is an intent-classification scripted-intent agent on a commercial platform, not a generative system, and it gives anonymous general guidance and makes no benefit decision, so nothing here should be read alongside the eligibility or fraud-scoring deployments. It is deliberately distinct from the verify-before-use benefits copilots (Nava-class, Benefit-Navigator-class) and Caddy, which place a human between the model and the affected person; here no one sits in that seat by design, and the human is reached only by escalation.
- baseline
The handover link is the payload and is drawn as a governance-controlled dial. The escalation fraction is one of the few directly measured, directly governable quantities in the atlas: about one in five conversations transferred to a human under free channel choice, and about 30 percent when NAV removed the explicit choice between the chatbot and human chat. Those figures are regime-specific and must be read against the interface in force, so the model-to-advisors edge runs at baseline to encode the roughly one-in-five free-choice regime while the case file records that NAV, not the citizen, sets how visible the human option is. The peak-load figures (a volume equal to the capacity of roughly 220 to 230 human advisors at week 13 of 2020) carry a documented source tension across the vendor case study, the EJIS article and the NAV-funded report, and are left unresolved.
- baseline
Containment is not resolution, and this shape's honest defining absence is drawn as a latent check: the roughly 80 percent of conversations that end without a human are a completion or non-escalation rate, not an accuracy rate, and no per-answer error rate has been published. Independent chat-log studies document irrelevant answers, omitted information, and three classes of domain-knowledge failure, with the most critical failures occurring when a misunderstanding goes undetected inside a conversation the chatbot completed. The independent model check therefore starts closed - the per-answer accuracy audit the record does not have, even under an unusually research-heavy oversight that studies rather than gates - and a cross-model or oversight-cadence review is what opens it.
- baseline
What crosses the handover is degraded, drawn as the weak store-to-advisors read: the chatbot transcript accompanies the escalated conversation to the advisor, but the NAV-funded three-university project found context survives imperfectly, advisors developed their own workarounds at the boundary, and citizens transferred to a human are often unsure whether they are chatting with a person or a machine. NAV in effect adopted that project's recommendation of an internal chatbot to assist chat employees at handover with an advisor-facing assistant piloted in autumn 2024 (a separate system, named without any model identifier). The store coupling is otherwise protective: the chatbot writes no generated content and no case record, and the trainers' content loop runs the benign direction, updating the scripted intents from mined failures rather than contaminating them.
- assumed
The chatbot's raw error rate is a modeling choice, not a measured per-interaction rate. No per-answer accuracy rate has been published; the reported roughly 80 percent containment is a completion rate, and the vendor figures (270,000-plus inquiries, 80 percent containment, a 220-full-time-employee equivalence, a 250 percent surge) originate in the platform vendor's marketing case study and are labeled as such. The error is held a touch above the verify-before-use copilot family because there is no professional intermediary and the signature failure is the undetected misunderstanding, and below the ungated generative public advisers because in the documented period Frida is a curated scripted-intent library rather than a generative model.
- assumed
Peer pathways are authored on both signs - one public-facing chatbot answers everyone identically, a monoculture in the Lab's qualitative vocabulary, so a wrong answer repeats rather than scatters, and the advisors' own handover workarounds spread across the NKS chat team - and the research-heavy oversight is drawn as a present node that studies rather than gates. These are modeling assumptions, not measurements.
- assumed
The anonymous citizens who use Frida are the operator network in this diagram, as an adoption channel only; no benefit, harm or downstream outcome to any person is computed from anything here. There is no measured demographic disparity in the record and no per-topic (for example benefits-specific) error rate, so equity_observations is deliberately empty and the honesty boundary is carried in prose, never in the dynamics. NAV's own 2025 analysis found chatbot visibility appears not to change contact-center inquiry volumes and attributes the steady post-2019 decline to a bundle of self-service improvements, new application systems, SMS notifications and changed NKS practices, so this Lab does not present Frida as demonstrably reducing human workload outside the crisis peak. Whatever an answer means for the person who acts on it is documented in the case file and measured outside any diagram like this one.
What this example does not show
- Frida is an anonymous, navigation-tier chatbot that gives general guidance and makes no benefit decision: it does not adjudicate eligibility, entitlement or a sanction, so appeal and override constructs from the decision-system cases do not apply, and it should not be read alongside the eligibility or fraud-scoring deployments. The anonymous citizens who use it are modeled here only as an adoption channel; no benefit, harm or downstream outcome to any person is computed from anything here, and whatever an answer means for the person who acts on it is documented in the case file and measured outside any diagram like this one.
- The headline pandemic figures (more than 270,000 inquiries, 80 percent of conversations resolved without a human, a peak equal to roughly 220 full-time employees, a 250 percent surge) originate in the platform vendor's marketing case study and are labeled as vendor claims; the NAV-funded research report independently corroborates the surge and the roughly one-in-five transfer rate but states a capacity of 230 advisors for week 13 of 2020 where the vendor and the EJIS article say 220 - a documented source tension carried unresolved, not a single verified number. The reported roughly 80 percent containment is a completion or non-escalation rate, not an accuracy rate: no per-answer error rate has been published, and the escalation fractions are regime-specific (about one in five under free channel choice, about 30 percent when the human-chat option was hidden), so any modeled handover figure states which interface regime it represents.
- There is no independent audit of the chatbot's answer accuracy in the record and no measured demographic disparity, so equity effects are not modeled; the case's evidence is unusually research-heavy but that oversight studies the tool rather than gating it. NAV's own 2025 channel-use analysis found that chatbot visibility appears not to change contact-center inquiry volumes and attributes the steady post-2019 decline to a bundle of self-service improvements, new application systems, SMS notifications and changed NKS practices, so this scenario does not present Frida as demonstrably reducing human workload outside the crisis peak. Two of the underlying theses were verified through the Norwegian national research archive rather than a full-text fetch (the institutional repositories migrated platforms), and one of the peer-reviewed chat-log studies anonymizes the chatbot under a pseudonym rather than the name Frida; the system identification is sound but the naming caveat is noted.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Frida is the chatbot at the front line of the Norwegian Labour and Welfare Administration's (NAV) anonymous contact-center chat channel; NAV states it launched in summer 2018 and, as of 2026, that citizens first meet Frida (open 24 hours a day) and can ask it for a human advisor on weekdays between 9:00 and 15:00, with the channel anonymous and no personal information visible to NAV. During the COVID-19 lockdown NAV reported a roughly 250 percent surge in inquiries; the platform vendor's case study reports the chatbot answered more than 270,000 coronavirus-related inquiries and that about 80 percent of enquiries were resolved without escalating to a human, and NAV's own funded research report records nearly 11,000 inquiries in Frida on some days between March and May 2020 with a week-13-2020 peak equal to the capacity of about 230 human advisors, where the vendor and the peer-reviewed EJIS study state about 220. These pandemic figures originate substantially in the vendor's marketing case study and are reported here as vendor claims with the 220-versus-230 source tension left unresolved; the roughly 80 percent containment is a completion or non-escalation rate, not a measure of answer accuracy.
empirical- Vendor boost.ai (vendor), How conversational AI is helping Norway's citizens through COVID-19 (NAV case study, 2020) https://boost.ai/case-studies/how-conversational-ai-is-helping-norways-citizens-with-covid/
- Government evaluation Parmiggiani, Farshchian, Vassilakopoulou, Pappas, Grisot, Frida@work: forskningsprosjekt om betydningen av tillit i bruken av chatboten Frida i NAV (project report for NAV, NTNU / University of Agder / University of Oslo, 2021) https://www.nav.no/_/attachment/download/a9ba64cd-c8ee-4177-a0e7-e1c5ae2749d1:1563c75472fae37937be5b64d6a796569fe35160/Frida@work_sluttrapport.pdf
- Academic Vassilakopoulou, Haug, Salvesen, Pappas, Developing human/AI interactions for chat-based customer services: lessons learned from the Norwegian government (European Journal of Information Systems, 2022 online first; print 2023, 32(1)) https://www.tandfonline.com/doi/full/10.1080/0960085X.2022.2096490
- Government NAV, Contact us (nav.no chat section, 2026) https://www.nav.no/kontaktoss/en
The best-documented property of NAV's Frida chatbot is its chatbot-to-human handover boundary, which the evidence suggests behaves as a governance-controlled dial: NAV's funded three-university Frida@work project reports that about one in five conversations transferred to a live human advisor under free channel choice, and only about 30 percent of dialogues transferred when NAV removed the explicit choice between the chatbot and human chat, a regime-specific figure that must be read against the interface in force. Independent chat-log studies document irrelevant answers, omitted information, and three classes of domain-knowledge failure, with the most critical failures occurring when a misunderstanding goes undetected inside a conversation the chatbot completed; no per-answer accuracy or error rate has been published, and the Frida@work project found context survives the handover imperfectly, with citizens often unsure whether they are talking to a person or a machine. NAV's own 2025 channel-use analysis found that chatbot visibility appears not to change contact-center inquiry volumes and attributes the steady post-2019 decline to a bundle of causes (self-service improvements, new application systems, SMS notifications, and changed contact-center practices), so the chatbot is not shown to reduce human workload outside the crisis peak.
empirical- Government evaluation Parmiggiani, Farshchian, Vassilakopoulou, Pappas, Grisot, Frida@work: forskningsprosjekt om betydningen av tillit i bruken av chatboten Frida i NAV (project report for NAV, NTNU / University of Agder / University of Oslo, 2021) https://www.nav.no/_/attachment/download/a9ba64cd-c8ee-4177-a0e7-e1c5ae2749d1:1563c75472fae37937be5b64d6a796569fe35160/Frida@work_sluttrapport.pdf
- Academic Verne, Steinsto, Simonsen, Bratteteig, How Can I Help You? A chatbot's answers to citizens' information needs (Scandinavian Journal of Information Systems, 2022, 34(2)) https://aisel.aisnet.org/sjis/vol34/iss2/7/
- Academic Simonsen, Steinsto, Verne, Bratteteig, I'm Disabled and Married to a Foreign Single Mother: Public Service Chatbot's Advice on Citizens' Complex Lives (Electronic Participation, ePart 2020, Springer LNCS) https://link.springer.com/chapter/10.1007/978-3-030-58141-1_11
- Government evaluation McVey, Chatboten Frida og utvikling i kanalbruk hos Nav (Arbeid og velferd nr. 2-2025, NAV analysis journal) https://www.nav.no/no/nav-og-samfunn/kunnskap/analyser-fra-nav/arbeid-og-velferd/arbeid-og-velferd/arbeid-og-velferd-nr.2-2025/chatboten-frida-og-utvikling-i-kanalbruk-hos-nav
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Benefits navigation & public-facing chat domain page.
Levers available here and the patterns behind them
- Escalate checks — State-feedback vigilance
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Peer sharing rules — Peer-edge governance
- Keep skills sharp — Deskilling-arrest mandate
- Keep prompts neutral — Framing and mirroring reduction
- Gate record entries — Human-in-the-loop write gating
- Vet connections — Connection authorization
- Store less data — Data minimization
- Mark AI-written records — Provenance labeling
- Upgrade model — Improve the model
Documented case histories
- Frida (NAV Norway)
- Nava assistive benefits chatbot
- Caddy adviser copilot at Citizens Advice
- GOV.UK Chat
- Mass.gov Virtual Assistant
- SSA 800-Number Conversational AI Assistant
- EDD Virtual Assistant
- Burokratt
- Singapore's chatbot fleet refresh: eighty scripted engines slated for retirement onto a shared LLM platform
- IRS collection chatbots: expanded and made permanent with no performance measures
- Albert France Services
- Propel in-app SNAP benefits assistant
- GetCalFresh: the nonprofit front door that carried most of California's online SNAP intake
- MyFriendBen benefits screener
- Benefits Data Trust wind-down