Skip to content

PAN Lab example

Albert France Services

Killed without a number: a sovereign adviser assistant that no metric ever measured

This assistant answered a France Services adviser's benefits question with a sourced draft answer from a curated set of official documents, for the adviser to verify, modify and validate before relaying to the citizen. Modeled on Albert France Services. It launched as a sovereign French AI, presented by the Prime Minister, and then gave a wrong answer on identity-card cost in a demonstration before him. But here is the thing worth studying: no error rate, no override count, no usage figure was ever published. Nothing measured whether it was any good. What finally surfaced its failures was not a metric but the workforce — advisers who found it worse than an ordinary search, unions who wrote the malfunctions down, one televised wrong answer. It was quietly stopped, then formally denied generalization, and the only number that ever governed its successor is a cost. The question is not which control failed. It is what you build when nothing was ever watching.

Stylized model of a documented deploymentBenefits navigation & public-facing chat

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Albert-class sovereign adviser assistant (deployment-and-withdrawal) network: 7 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 5 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the sovereign adviser-assistant deployment-and-withdrawal pattern documented in the Albert France Services case file — not a reconstruction of the actual tool or any of its three versions. It is a deliberate sibling of the library's verify-before-use benefits copilots (Nava-class, Benefit-Navigator-class), which put a caseworker between the model and the affected person, and distinct from them: those model US, nonprofit pilots whose risk is a caseworker over-trusting the tool, while this models a sovereign, state-built pilot that was formally denied generalization, whose advisers under-trusted and abandoned it. It is also the near-inverse of the GOV.UK-Chat-class staged-gate assistant: that gate turned on a published (if self-graded) accuracy number, while this deployment published no error number at all.

  • baseline

    The case's defining absence is drawn as a latent check: no instrumented error measurement or independent accuracy audit ever gated a version. No error rate, override count or usage figure was ever published; the deployment carried no metric-mediated detection channel, so its failures surfaced through the operator side — union-documented malfunctions, adviser disaffection, and one televised wrong answer on identity-card cost given in a demonstration before the Prime Minister — rather than through any measurement. The independent model check therefore starts closed, and a cross-model check opens the independent read the deployment never had.

  • baseline

    The payload is that the error-detection channel of last resort was the workforce. The DINUM/ANCT experiment team is drawn as an internal correction channel that had access to the tool but ran no error-rate measurement, iterating three versions qualitatively; the Union and press error dossier is drawn as the operator-side detection that actually fired, aggregating adviser disaffection and a televised error into the reputational signal that drove the withdrawal. The contrast between the body positioned to measure error and not doing so, and the body not positioned to measure it doing the measuring, is the case's lesson and is why the operator-to-detection pathway runs at its normal level while the instrumented independent model check sits empty.

  • assumed

    The store coupling is protective, like the copilots and unlike a contaminating writer: the assistant does not write generated answers into the curated official base (the national operators maintain it), so the contamination pathway is drawn closed at baseline; retrieval grounds each answer in curated official content; and every answer routes the adviser back to the cited official source to check before relaying. These are the model-to-store write held latent, the store-to-model retrieval at baseline, and the store-to-operator route-back at baseline. The near-zero override cost cuts both ways: it bounds contamination when advisers ignore a wrong answer, and it is exactly what let adoption silently collapse.

  • assumed

    The model's raw error rate is a modeling choice, not a measured per-interaction rate; no error rate was ever published for this deployment. It is held a touch above the well-curated verify-before-use copilots because the qualitative record is of a tool advisers found underperformed an ordinary search, with recurring malfunctions and one emblematic demo error, and held near default rather than high because answers were grounded in a curated official base. Whether any given version improved is documented in the case file (three iterated versions), not computed here.

  • assumed

    Peer pathways are authored on both signs. Reinforcing but corrosive: the verdict that Albert answered worse than an ordinary search, and the habit of ignoring it, spread adviser to adviser, so the peer spread here erodes adoption rather than amplifying error — the distinctive twist of an advisory tool with near-zero override cost. Inhibiting: the verify-modify-validate expectation is a shared check. And one sovereign state model answering advisers identically across counters is a monoculture in the Lab's qualitative vocabulary: a wrong answer repeats rather than scatters. These are modeling assumptions, not measurements.

  • assumed

    Served citizens are not in the dynamics; the France Services advisers are the operator network in this diagram, as an adoption channel only. No benefit, harm or downstream outcome to any citizen is computed from anything here, and there is no per-topic (for example benefits-specific) error rate, override count or complaint data in the record — the case's whole point is that no such measurement was ever published. The harm mode on this shape is a wrong sourced-looking answer relayed to a citizen, or an abandoned tool, documented in the case file and measured outside any diagram like this one. This Lab models institutional propagation only, and the cost figures in the record (a union-claimed roughly 1.3 million euro project cost, an AFP-reported roughly 1.2 million euro annual DINUM AI budget with Albert a minimal share) are exogenous, carried in prose, never in the dynamics.

What this example does not show

  • Served citizens, and the benefits or procedures they did or did not ultimately complete, are not modeled here; the Lab models institutional propagation only, and those outcomes are documented in the case file and measured outside any diagram like this one. There is no per-topic (for example benefits-specific) error rate, override count or complaint data — the case's whole point is that no such measurement was ever published.
  • No quantitative performance data exists for this deployment: no error rate, usage volume, adoption frequency or override count was ever published, so nothing here implies a measured error rate existed. The contested figures in the record are carried as claims, not facts: the roughly 1.3 million euro project cost and the September 2025 quiet-stop date are the union Solidaires Finances Publiques' claims (the union inferred the stop from the project no longer appearing in a ministerial working group), while AFP reports DINUM's annual AI budget at about 1.2 million euros since 2024 with Albert France services a minimal share; DINUM itself disputes the failure framing, stating the majority of Albert-brand projects are sustained and fully operational. A full cost-accounted evaluation of the successor was still pending as of mid-2026.
  • The emblematic wrong answer on identity-card cost was given in a demonstration before the Prime Minister; the sources do not hard-date it to the 23 April 2024 televised launch, and French identity-card fee rules vary by case (renewal of an expired card is free, while a 25 euro fee applies to replacement after loss or theft), so it is characterized here as a wrong answer on ID-card cost in the demonstrated case, not as an adjudication of the fee schedule. Mistral AI is named only as the vendor whose models the successor integrates; no AI model identifiers appear here.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • Albert France Services was a sovereign, in-house generative AI assistant built by DINUM with ANCT to help France Services counter advisers answer citizens' benefits and procedure questions from a curated base of official documents, presented by the Prime Minister as a sovereign French AI in April 2024 and, in a demonstration before him, giving a wrong answer on identity-card cost. Piloted from an initial panel of about sixty volunteer advisers to roughly eighty advisers across more than forty counters (forty-eight at final count per AFP) in six departments over three iterated versions, it was, per a January 12, 2026 AFP dispatch, formally not going to be generalized 'in its current form,' a decision DINUM announced on January 9, 2026 while stating that the majority of Albert-brand projects are sustained and fully operational. No error rate, usage volume, or override count for the tool was ever published; AFP reports DINUM's annual AI budget at about 1.2 million euros since 2024 with Albert France Services a minimal share, a figure distinct from and not directly comparable to the union Solidaires Finances Publiques' separate claim of a roughly 1.3 million euro project cost.

    empirical
    • Trade press Weka.fr (AFP dispatch), Albert, l'outil d'IA generative, experimente a France Services ne sera pas generalise (2026) https://www.weka.fr/actualite/administration/article/albert-l-outil-d-ia-generative-experimente-a-france-services-ne-sera-pas-generalise-209194/
    • Trade press Acteurs Publics, Derriere l'echec mediatique d'Albert, un projet d'IA plus global qui s'ancre dans l'Etat (2026) https://acteurspublics.fr/articles/de-chatbot-experimental-a-socle-interministeriel-pour-lia-de-letat-le-parcours-dalbert-ia/
    • Advocacy Solidaires Finances Publiques, Entre ici Albert, au pantheon des IA souveraines (2026) https://solidairesfinancespubliques.org/le-syndicat/dossiers/ia-a-la-dgfip/7192-albert-france-service.html
    • Government France services / ANCT, Experimentation d'un modele d'assistance aux conseillers France services base sur l'intelligence artificielle (2024) https://www.france-services.gouv.fr/actualites/experimentation-dun-modele-dassistance-france-services-IA
  • For Albert France Services no instrumented error-detection channel existed: no error rate, override count, or usage figure was published during the pilot, and the failures that framed the tool surfaced through the operator side, with several unions documenting recurring malfunctions and wrong answers, an investigative-television broadcast in April 2025 (per Solidaires Finances Publiques) featuring unenthusiastic agent testimony, and advisers reporting answers worse than an ordinary search. According to Solidaires Finances Publiques the project had in fact stopped by September 2025 with no announcement, inferred from Albert no longer appearing among projects presented in a ministerial working group (a union claim). Alongside the January 2026 non-generalization decision DINUM migrated the Albert API model aliases off the 'albert-' branding and removed the web-search functionality, retiring legacy aliases by February 15, 2026, while a successor adviser tool that integrates models from the vendor Mistral AI was in test with about 10,000 public agents through June 2026, gated by a summer-2026 evaluation that must notably establish the cost of a generalization.

    empirical
    • Advocacy Solidaires Finances Publiques, Entre ici Albert, au pantheon des IA souveraines (2026) https://solidairesfinancespubliques.org/le-syndicat/dossiers/ia-a-la-dgfip/7192-albert-france-service.html
    • Trade press Next (next.ink), Albert: l'IA souveraine de la Dinum ne sera pas generalisee dans sa forme actuelle (2026) https://next.ink/brief-article/albert-lia-souveraine-de-la-dinum-ne-sera-pas-generalisee-dans-sa-forme-actuelle/
    • Trade press Weka.fr (AFP dispatch), Albert, l'outil d'IA generative, experimente a France Services ne sera pas generalise (2026) https://www.weka.fr/actualite/administration/article/albert-l-outil-d-ia-generative-experimente-a-france-services-ne-sera-pas-generalise-209194/
    • Trade press Acteurs Publics, Derriere l'echec mediatique d'Albert, un projet d'IA plus global qui s'ancre dans l'Etat (2026) https://acteurspublics.fr/articles/de-chatbot-experimental-a-socle-interministeriel-pour-lia-de-letat-le-parcours-dalbert-ia/

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Benefits navigation & public-facing chat domain page.

Levers available here and the patterns behind them

Documented case histories