PAN Lab example
BAMF dialect recognition
One clue among many or the thing that decides
Germany's asylum agency uses DIAS, dialect-recognition software, to estimate applicants' origin from their speech. Does its estimate count for more than its accuracy supports?
See more
DIAS is dialect-recognition software run by Germany's federal asylum agency, the BAMF. It analyzes a short recording of an asylum applicant speaking and estimates the country or region their speech is most consistent with. Since 2017 the BAMF has used it to check stated origin when the applicant has no identity documents.
How accurate it is
The government's answers in the Bundestag, Germany's parliament, put recognition for Arabic at about 80 percent in 2017. So roughly one estimate in five was wrong. The figure was 87 percent in 2023, and about 75 percent for other languages in 2022.
Computational linguists judge some closely related varieties close to impossible to separate. Mark Liberman called separating vernacular Persian, Dari, and Pashto "probably pretty much hopeless." The varieties overlap, and a short sample cannot resolve them. The case file concludes that this is not a tool that is usually right with occasional errors. It is often uncertain by the nature of the task.
How it is used
The estimate is one input into a hard and important step: judging whether an applicant's stated origin is credible. That judgment can bear on whether they qualify for protection. Origin is often contested and hard to verify, so the case file calls the intent defensible.
In peer-reviewed fieldwork, BAMF decision-makers called the software "only a rough compass" and "too imprecise to solve the problematic cases." They treat its results as clues, not answers. The case file credits this framing as the responsible one. Used honestly as one clue among several, the tool is defensible.
What can go wrong
The documented risk is that an imprecise result gains more authority than its accuracy supports. Here the state's tool is set against the applicant's own account of who they are. A statement that the software finds their speech inconsistent with their stated origin is hard to rebut. It can carry more weight in the room and in the record than an error rate of about one in five, the government's 2017 figure for Arabic, warrants.
The BAMF also analyzes data from applicants' mobile phones. Fieldwork found this produced usable results in about a third of cases, and refuted a stated identity in about 2 percent.
Who can correct it
The applicant knows their own origin and has the most at stake. Yet applicants are often not shown the software's role or its estimate in a form they can contest. So an error they could explain at once may never surface. It could be a childhood across a border, an education in a second dialect, or a family that spoke differently from the region.
The stakes are what make this matter. As the case file puts it, a 20 percent error rate is one thing in a recommendation. It is another when it helps decide whether a person is returned to a country they fled.
How its use has grown
In July 2022 the BAMF extended the software to Farsi, Dari, and Pashto, and to more varieties of Arabic. It became the basis of a pilot by the EU's asylum agency, the EUAA, in seven countries. They are Austria, Finland, Norway, Sweden, Greece, Switzerland, and Lithuania.
This case starts with one pressure switched on, named Autonomy expands. Here it stands for that growth. On this network it adds to how much the software's estimates are written into the case record. It also adds to how far caseworkers rely on the software, and mistakes build on one another more.
Where the facts come from
The sources for this case are AlgorithmWatch's reporting from 2022 and peer-reviewed fieldwork by Scheel published in 2024. A third is a 2026 legal analysis by Beck in Verfassungsblog, a German law blog. The accuracy figures are the government's own, given in answers in the Bundestag.
What this network is drawn from
This network follows the pattern the case file documents. It does not reconstruct the actual system. It shows the dialect software, the phone data analysis, the applicant's speech sample, the caseworkers, the case record, and the reliability and challenge controls. The applicant is outside the network, and no asylum outcome is computed.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A tool closes a pathway when mistakes stop passing along it. The work along it may go on.
This case has a budget of 11 units. Understand the system, a Lab option that lowers some tools' prices, is not on offer here. Every price below is the full price. Explore (No Targets) sets no targets. The other three levels all ask for the network's mistakes to be contained, meaning corrected rather than building on each other.
Lingering effects are effects that stay after their cause is gone. The Lab's Dynamics menu starts at Both, which includes them, under Explore (No Targets) and Service Targets Only, where you can change it. Service and Safety Targets fixes it at Side-effects, which leaves them out. All Governance Targets fixes it at Both. With lingering effects on, and Understand the system not on offer, three tools work at reduced strength: Store less data, Review the riskiest first, and Check copied records.
Under Service Targets Only, the targets can be met, but no single tool meets them. The cheapest combination costs 4 units: Mark AI-written records with Escalate checks. Three more cost 5 units. They are Gate record entries with Escalate checks, Gate record entries with Mark AI-written records, and Mark AI-written records at its stronger setting with Escalate checks. With lingering effects off, Store less data with Escalate checks or with Mark AI-written records also meets them, for 5 units.
Under Service and Safety Targets, the case is not fully addressable with the available tools. This level also asks you to close every failure pathway. Eight are open before any tool is used. The first four are Origin estimate given to the caseworker, Caseworker weighs the clue, Assessment written to the case record, and Case record material analyzed. The other four are Caseworker reads the case record, Origin estimate entered in the case record, Speech sample analyzed for origin, and Phone findings given to the caseworker. No tool on offer acts on Caseworker weighs the clue, so that pathway stays open whatever you use.
Under All Governance Targets, the case is not fully addressable with the available tools either, for the same reason.
The sources describe no one running two checks on this network: Weight tied to measured accuracy, and Applicant challenge to the estimate. Check with a second model adds the first, and Assign a challenger adds the second. The challenger is a colleague the office names to question the big calls, not the applicant. No tool on offer lets the applicant challenge the estimate. Neither closes a failure pathway, and neither meets the targets at any level on its own.
More is not better here. Using every tool at its strongest setting costs 44 units, four times the budget. It contains the mistakes and closes every failure pathway but Caseworker weighs the clue. It still meets the targets at no level, because the software then contributes too little to the caseworkers' work.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Origin-signal-class where reliability must bound authority network: 6 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 5 assumptions
- assumed
The network shows the applicant's recorded speech as its own source, because the applicant supplies the material the dialect software analyzes. The case file says applicants are often not shown what is made of it in a form they can contest. The network also shows how Germany's federal asylum agency, the BAMF, analyzes data from applicants' phones. Fieldwork found it produced usable results in about a third of cases and refuted a stated identity in about 2 percent. The network assumes two rough results pointing the same way can be read as confirming each other. They may share the same assumption about what an origin sounds or looks like. The workload is assumed heavy against limited caseworker capacity, since the BAMF handles asylum claims nationwide. Caseworkers are assumed to have real room for judgment, because the fieldwork records them calling the software a rough compass.
- baseline
This network follows the pattern the case file documents. It does not reconstruct the actual system. Germany's federal asylum agency uses dialect-recognition software to estimate an applicant's origin, as one input into judging whether their stated origin is credible. Government-reported recognition for Arabic was about 80 percent in 2017, so roughly one estimate in five was wrong. Linguists judge some varieties close to impossible to separate. The BAMF's own caseworkers call the software a rough compass, too imprecise to solve the hard cases. These figures and words are findings from investigative reporting and peer-reviewed fieldwork, entered as recorded.
- assumed
The network credits the caseworkers' practice of treating the estimate as one clue. The software does not make the decision, and the caseworkers who use it know its limits. The documented risk is that an imprecise result gives the caseworker more authority than its accuracy warrants. A statement that the software finds the applicant's speech inconsistent with their stated origin is hard to rebut. It can carry more weight in the room and in the record than an error rate of about one in five, the government's 2017 figure for Arabic, supports. The network places this risk on the link named One tool estimates origin case after case. It adds a check, Weight tied to measured accuracy, for a control that would make the estimate's accuracy limit its weight.
- baseline
The network shows the applicant's chance to challenge the estimate as a check, Applicant challenge to the estimate. The applicant knows their own origin and has the most at stake, so they are the best-placed person to correct an error. The case file says asylum procedures routinely withhold the software's role or estimate in a form the applicant can contest. So an error the applicant could explain may never surface. It could be a childhood across a border or an education in a second dialect. As the case file puts it, a 20 percent error rate is one thing in a recommendation. It is another when it helps decide whether a person is returned to a country they fled.
- assumed
No asylum outcome and no applicant's credibility is computed here. The network shows how mistakes pass between the agency's tools, caseworkers, and records, and the applicant is outside it. The software's reliability figures, the linguists' judgments, the caseworkers' rough-compass words, and the gap in disclosure are recorded in the case file. None is computed from anything in this network, and nothing here judges any individual claim.
What this example does not show
Show all 2 limitations
- This example does not show any asylum outcome or any applicant's credibility. It shows how mistakes pass between the agency's tools, caseworkers, and records, and the applicant stays outside it. The software's reliability figures, the linguists' judgments, the caseworkers' rough-compass words, and the gap in disclosure are recorded in the case file. None is computed in this network, and nothing here judges any individual claim.
- This example does not compute any harm. The reliability figures and the caseworkers' rough-compass words are findings from investigative reporting and peer-reviewed fieldwork, entered as recorded. The network shows two checks the sources do not describe in use: tying the estimate's weight to its accuracy, and the applicant's chance to challenge it. They stand for controls, not for a measured harm.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A federal asylum agency uses dialect-recognition AI to estimate an applicant's country or region of origin from a short speech sample, as one input into the credibility assessment of their claimed origin. The tool's reliability is limited: government-reported recognition was around 80 percent for Arabic in 2017 — roughly a 20 percent error rate — and computational linguists judge separating some closely related language varieties close to hopeless. The agency's own caseworkers describe the tool as only a rough compass, too imprecise to resolve the hard cases, and its outputs as clues rather than determinations. Used honestly as one clue among several it is defensible; the documented risk is that an imprecise output acquires more authority than its accuracy supports, in a determination where the state's tool is set against the applicant's own account of who they are.
empirical- Investigative Lulamae, J. (2022, September 5). The BAMF's controversial dialect recognition software: new languages and an EU pilot project. AlgorithmWatch; with Beck, J. (2026), Verfassungsblog legal analysis (https://doi.org/10.59704/b22636dc94f60b29). https://verfassungsblog.de/dialect-recognition-software-dias-law/
- Peer-reviewed Scheel, S. (2024). Epistemic domination by data extraction: questioning the use of biometrics and mobile phone data analysis in asylum procedures. Journal of Ethnic and Migration Studies, 50(9), 2289-2308. https://doi.org/10.1080/1369183X.2024.2307782 https://pmc.ncbi.nlm.nih.gov/articles/PMC11034547/
Two governable surfaces follow from putting a low-reliability signal into a high-stakes credibility determination. First, whether the tool's documented imprecision actually bounds the weight it carries: a rough compass treated as one is honest, but the same output can harden into a credibility finding it cannot support once a phrase like the software indicates a particular origin enters the record and confronts the applicant. Second, whether the applicant can see and contest the signal: in asylum determinations the person with the most at stake and the most knowledge of their own origin is often unable to see or challenge the AI's estimate, so the correction that would catch an error is severed on exactly the side that holds the truth. The governable reading is that reliability must bound authority, and the affected person must be able to contest a signal used against them.
empirical- Peer-reviewed Scheel, S. (2024). Epistemic domination by data extraction: questioning the use of biometrics and mobile phone data analysis in asylum procedures. Journal of Ethnic and Migration Studies, 50(9), 2289-2308. https://doi.org/10.1080/1369183X.2024.2307782 https://pmc.ncbi.nlm.nih.gov/articles/PMC11034547/
- Investigative Lulamae, J. (2022, September 5). The BAMF's controversial dialect recognition software: new languages and an EU pilot project. AlgorithmWatch; with Beck, J. (2026), Verfassungsblog legal analysis (https://doi.org/10.59704/b22636dc94f60b29). https://verfassungsblog.de/dialect-recognition-software-dias-law/
Where this connects
Institutional pressures in this domain
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
All of them in context on the Immigration & asylum AI domain page.
Levers available here and the patterns behind them
- Store less data — Data minimization
- Assign a challenger — Structured dissent
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Review on schedule — Oversight cadence & retrospectives
- Review the riskiest first — Risk-tiered oversight
- Gate record entries — Human-in-the-loop write gating
- Train the staff — AI literacy & boundary rules
- Mark AI-written records — Provenance labeling
- Escalate checks — State-feedback vigilance
- Upgrade model — Improve the model