PAN Lab example
Burokratt
The network of networks: a federated public-service chatbot
One chat window, many bots behind it. Each Estonian public institution runs its own assistant with its own knowledge base, a central classifier routes a citizen's question to the right one and oversees the handover, and every assistant also reads one shared knowledge module built from the state information portal. Modeled on Estonia's Burokratt national assistant network. Nothing here decides anything: the bots answer questions and hand a stuck chat to a human representative. The trap is the shared link. A single stale or wrong entry in the module every node reads, or a classifier that routes to the wrong bot, is not one institution's error - it is the whole network's answer, and the record shows the interoperability was built faster than anything to evaluate it.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Burokratt-Estonia-class federated public-service chatbot network network: 7 components and 14 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 7 assumed. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the federated chatbot-network pattern documented in the Burokratt case file - many institutional assistants sharing one central routing classifier and one shared knowledge module - not a reconstruction of the actual platform. One representative institution is drawn for the federation. The published PAN model draws two institutions, a higher-uptake and a lower-uptake one, with their own staff and stores; that pair differs in how much the channel is used, not in how it is wired, so the board carries the difference as a documented range rather than as a second copy of the same sub-network.
- assumed
No published session volumes, escalation rates, or answer-accuracy figures for Burokratt were located, so every baseline here is an estimated shape drawn from the architecture and governance record, not a calibration to operational data; the absence of a documented harm or error incident reflects the absence of published evaluation, not evidence of absence.
- assumed
The shared eesti.ee knowledge module grounds every assistant, so staleness or error there reaches the whole network while per-institution stores localise the rest; the central platform team's ingestion is the shared module's inflow and the single most consequential write on the diagram. The record documents shared content being retrieved into answers beside an institution's own content, not copied into institution stores, so the coupling between the two stores runs through the assistant and no direct store-to-store link is drawn; the published PAN model carries that coupling as an explicit estimate.
- assumed
Cross-institution practice diffusion is drawn once, from RIA's central onboarding community to institution staff, the conduit the record names; a return flow from institutions into that community is documented as promoted reuse and is narrated on the same link rather than drawn. The one documented absence - no dedicated algorithmic-oversight body, published evaluation framework, or national-audit report located - is drawn once, as the cross-institution validation of the classifier and shared module that starts closed. A reconciliation of the shared module against institution content and a federation-wide peer cross-check of answers were likewise not located; they are readings of that same absence and are not drawn as further closed checks.
- assumed
The governance node carries the oversight the record documents as real - RIA central platform governance from demo through production and the EU Recovery and Resilience Facility milestone monitoring - as light pathways: the platform's delivery state surfacing to governance, and a stage review over the central team's platform work. This is delivery oversight, drawn deliberately apart from the algorithmic-evaluation absence, which stays a latent check; both intensities are estimated shapes, not calibrations.
- assumed
Served citizens who query the assistants are not in the dynamics. The system makes no eligibility or benefit determination; what propagates here is the quality of an institution's own answers, not any citizen outcome, which is documented in the case file and measured outside any diagram like this one. Adoption is uneven across institutions - about one-third of daily requests in one, minimal in another - and the representative institution is drawn at the exercised end of that range: the routing, handover, trainer-authoring and store-upkeep links sit one step higher than they would read for a minimally used institution, where the escalation and retraining habit gets little practice.
- assumed
The central platform team sees every routed answer pass through the classifier it operates and reads the shared module it maintains. Neither is drawn as its own pathway: no source documents the team adopting or altering an answer in passage, and its influence on the network runs through the ingestion write into the shared module, which is drawn. The published PAN model carries the hub's exposure to routed answers as a modeling device that gives the shared module a live inflow; here the inflow is the ingestion write itself.
What this example does not show
- Served citizens who query the assistants are not modeled here; the Lab models institutional propagation only. The system makes no eligibility or benefit determination, and no citizen outcome is computed from anything in this diagram.
- No published session volumes, escalation-to-human rates, or answer-accuracy figures for Burokratt were located in the public record, and no harm, error-incident, or audit-office report was found; every baseline here is an estimated shape drawn from the architecture and governance record, and the absence of documented incidents reflects the absence of published evaluation, not evidence of absence.
- Adoption is highly variable across institutions and low in some: the independent ethnography documents use differing considerably by institution (minimal in one, about one-third of daily requests in another). A concrete phone-versus-email volume comparison sometimes read into this case belongs to a separate Swedish municipal chatbot studied in the same paper, not to Burokratt, and is not used here.
- The 2026 per-institution AI-agent cooperative network and the Estonian-adapted large language model are a stated plan, not a deployed capability, and are treated as prospective throughout.
- The 53 million euro EU figure is a Recovery and Resilience Facility contribution described as being for two projects altogether under bundled digital-infrastructure and cloud-transition components; it is not a tool-only spend. The nearer tool-scale figures are about 1.5 million euros spent through early 2022 with roughly 13 million euros budgeted over the following four years.
- Two registers coexist in the record and are kept distinct here: a promotional one (a Siri of public services, a UNESCO top-100 listing) and an independent academic one (an FAQ-like function and uneven, in places low, usage); nothing promotional is stated as established performance.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Burokratt is Estonia's national network of public-sector chatbots operated by the Information System Authority: each participating institution runs its own assistant, a central classifier routes a citizen's query between them and oversees the handover, and from 2025 a shared knowledge module built from the eesti.ee state portal feeds cross-domain answers. RIA's page lists 20 participating organisations and trade press reports 18 integrated; an independent 2025 ethnography drawing on twelve insider interviews (conducted in late 2023, when the system spanned ten institutions) found it marketed as advanced AI while functioning much like an FAQ list, with use differing considerably by institution and low in some. No published session volumes, escalation-to-human rates, or answer-accuracy figures, and no dedicated algorithmic-oversight body, published evaluation framework, or national-audit report on the network, were located in the public record.
empirical- Government Information System Authority (RIA), Republic of Estonia, Burokratt (2025) https://www.ria.ee/en/state-information-system/personal-services/burokratt
- Trade press GovInsider, Estonia eyes cross-border interoperability for Burokratt, its Siri of public services (2025) https://govinsider.asia/intl-en/article/estonia-eyes-cross-border-interoperability-for-burokratt-its-siri-of-public-services
- Academic Kaun, Manniste, Public sector chatbots: AI frictions and data infrastructures at the interface of the digital welfare state (New Media and Society, 2025) https://journals.sagepub.com/doi/10.1177/14614448251314394
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Benefits navigation & public-facing chat domain page.
Levers available here and the patterns behind them
- Upgrade model — Improve the model
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Review on schedule — Oversight cadence & retrospectives
- Escalate checks — State-feedback vigilance
- Keep prompts neutral — Framing and mirroring reduction
- Peer sharing rules — Peer-edge governance
- Gate record entries — Human-in-the-loop write gating
- Vet connections — Connection authorization
- Store less data — Data minimization
- Mark AI-written records — Provenance labeling
Documented case histories
- Burokratt
- Nava assistive benefits chatbot
- Caddy adviser copilot at Citizens Advice
- GOV.UK Chat
- Mass.gov Virtual Assistant
- Frida (NAV Norway)
- SSA 800-Number Conversational AI Assistant
- EDD Virtual Assistant
- Singapore's chatbot fleet refresh: eighty scripted engines slated for retirement onto a shared LLM platform
- IRS collection chatbots: expanded and made permanent with no performance measures
- Albert France Services
- Propel in-app SNAP benefits assistant
- GetCalFresh: the nonprofit front door that carried most of California's online SNAP intake
- MyFriendBen benefits screener
- Benefits Data Trust wind-down