Skip to content

PAN Lab example

BAMF dialect recognition

One clue among many or the thing that decides

Dialect-recognition AI estimates an asylum applicant's origin from a speech sample, as one clue in a credibility assessment. Modeled on a deployment whose own caseworkers call it a 'rough compass' - imprecise by nature (~20% error for one language, some varieties near-impossible to separate). Used honestly as one input it is defensible; the documented risk is that an imprecise output carries more authority than its accuracy supports, against the applicant's own account. And the applicant, who knows their own origin, often cannot see or contest it. So watch whether reliability bounds the weight, and whether the correction loop reaches the person who holds the truth.

Stylized model of a documented deploymentImmigration & asylum AI

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Origin-signal-class where reliability must bound authority network: 6 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    The speech sample is drawn as its own inbound source at full strength, because the shape of this case is that the applicant supplies the material and never sees what is made of it. The severed contest loop only bites once the source is on the board: the person is simultaneously the origin of the evidence and the party with no access to it. And the agency runs a second system on the same question - device analysis, documented as producing something usable in about a third of cases and actually contradicting a stated identity in about two per cent. Drawing both is what makes the real risk legible: two rough signals pointing the same way read as corroboration, when they may only be two tools inheriting the same assumption about what an origin sounds or looks like. A heavy workload against limited capacity - national in scale, and capacity held above the floor because the fieldwork records caseworkers applying genuine interpretive judgment and calling the tool a rough compass in their own words.

  • baseline

    This models the credibility-signal pattern documented in the case file - not a reconstruction of the actual system. A federal asylum agency uses dialect-recognition AI to estimate an applicant's origin as one input into a credibility assessment. The tool is imprecise - government-reported recognition around 80% for one language (~20% error), some varieties near-impossible to separate - and the agency's own caseworkers describe it as a 'rough compass,' too imprecise to resolve hard cases. Reliability figures and the caseworkers' framing are recorded external findings (investigative reporting, peer-reviewed fieldwork), entered as such.

  • assumed

    The honest 'one clue' framing is credited - the tool does not automate the decision, and the caseworkers who use it know its limits. The documented risk (drawn on the model's self-loop) is that an imprecise output strengthens the state's epistemic authority in the credibility confrontation more than its accuracy warrants: 'the software indicates your speech is not consistent with your claimed origin' is hard for an applicant to rebut and can carry more weight in the room and the record than a 20% error rate supports. Reliability must bound authority - drawn as the empty model-side check.

  • baseline

    The severed correction loop is the empty oversight check: the applicant, who knows their own origin and has the most at stake, is the best error-corrector any such system could have, and asylum procedures routinely sever that loop by not disclosing the AI's role or output in a contestable form. An error a rough-compass reading makes - a childhood across a border, an education in a second dialect - may never surface because the loop that would surface it is closed on the applicant's side. The stakes make both surfaces non-negotiable: a 20% error rate is one thing in a recommendation and another when it helps decide whether a person is returned to a country they fled.

  • assumed

    No asylum outcome and no applicant's credibility is modeled here. This Lab reads institutional propagation only, and the applicant is boundary-only. The tool's reliability figures, the linguists' judgments, the caseworkers' rough-compass framing, and the disclosure gap live in the case file, and are never computed from anything in this diagram; nothing here adjudicates any individual claim.

What this example does not show

  • No asylum outcome and no applicant's credibility is modeled. The Lab reads institutional propagation only; the applicant is boundary-only, and the tool's reliability figures, the linguists' judgments, the caseworkers' rough-compass framing, and the disclosure gap live in the case file, never computed on this diagram — nothing here adjudicates any individual claim.
  • The reliability figures and the rough-compass framing are recorded external findings (investigative reporting, peer-reviewed fieldwork) entered as such; the diagram draws the reliability-bounds-authority control and the applicant's ability to contest as two latent checks, not a computed harm.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A federal asylum agency uses dialect-recognition AI to estimate an applicant's country or region of origin from a short speech sample, as one input into the credibility assessment of their claimed origin. The tool's reliability is limited: government-reported recognition is around 80 percent for Arabic — roughly a 20 percent error rate — and computational linguists judge separating some closely related language varieties close to hopeless. The agency's own caseworkers describe the tool as only a rough compass, too imprecise to resolve the hard cases, and its outputs as clues rather than determinations. Used honestly as one clue among several it is defensible; the documented risk is that an imprecise output acquires more authority than its accuracy supports, in a determination where the state's tool is set against the applicant's own account of who they are.

    empirical
    • Investigative Lulamae, J. (2022, September 5). The BAMF's controversial dialect recognition software: new languages and an EU pilot project. AlgorithmWatch; with Beck, J. (2026), Verfassungsblog legal analysis (https://doi.org/10.59704/b22636dc94f60b29). https://verfassungsblog.de/dialect-recognition-software-dias-law/
    • Peer-reviewed Scheel, S. (2024). Epistemic domination by data extraction: questioning the use of biometrics and mobile phone data analysis in asylum procedures. Journal of Ethnic and Migration Studies, 50(9), 2289-2308. https://doi.org/10.1080/1369183X.2024.2307782 https://pmc.ncbi.nlm.nih.gov/articles/PMC11034547/
  • Two governable surfaces follow from putting a low-reliability signal into a high-stakes credibility determination. First, whether the tool's documented imprecision actually bounds the weight it carries: a rough compass treated as one is honest, but the same output can harden into a credibility finding it cannot support once a phrase like the software indicates a particular origin enters the record and confronts the applicant. Second, whether the applicant can see and contest the signal: in asylum determinations the person with the most at stake and the most knowledge of their own origin is often unable to see or challenge the AI's estimate, so the correction that would catch an error is severed on exactly the side that holds the truth. The governable reading is that reliability must bound authority, and the affected person must be able to contest a signal used against them.

    empirical
    • Peer-reviewed Scheel, S. (2024). Epistemic domination by data extraction: questioning the use of biometrics and mobile phone data analysis in asylum procedures. Journal of Ethnic and Migration Studies, 50(9), 2289-2308. https://doi.org/10.1080/1369183X.2024.2307782 https://pmc.ncbi.nlm.nih.gov/articles/PMC11034547/
    • Investigative Lulamae, J. (2022, September 5). The BAMF's controversial dialect recognition software: new languages and an EU pilot project. AlgorithmWatch; with Beck, J. (2026), Verfassungsblog legal analysis (https://doi.org/10.59704/b22636dc94f60b29). https://verfassungsblog.de/dialect-recognition-software-dias-law/

Where this connects

Institutional pressures in this domain

  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).

All of them in context on the Immigration & asylum AI domain page.

Levers available here and the patterns behind them

Documented case histories