PAN Lab example
X Multilingual Hate-Speech Enforcement
The error table the law forced into twenty languages
Automated classifiers enforce a hateful-conduct policy across the 24 official languages of the European Union, on a platform the law obliges to publish its accuracy indicators broken down by language — the only language-indexed error signal in this sector's public record. Modeled on X's enforcement as disclosed in its own mandated DSA filings. The table those filings produced shows appeal rates of 0.0 to 14.0 percent and overturn rates of 0.0 to 72.2 percent across twenty languages; three languages recorded zero appeals and so no error measurement at all; four official languages have no column. Behind the table: 1,352 named reviewers covering 8 of 24 languages, 1,197 of them in English, with the remaining thirteen languages reviewed through translation tools. In the very next period the language table became a country table. So watch the sensor, not just the enforcement: the correction channel is the only error measurement there is, its reach is indexed by the served person's language, and the measurement itself is what decayed. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. With every tool the Lab currently offers, no affordable combination brings this system inside the win condition at those settings. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Enforcement-class with its error signal indexed by language network: 10 components and 25 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 2 assumed · 11 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the language-indexed enforcement pattern documented in the dsa-multilingual-hate-speech case file — not a reconstruction of the actual systems. It is the one deployment in this sector whose error signal is language-indexed in the public record, and it is so only because Regulation (EU) 2022/2065 Articles 15(1)(e) and 42(2) compel that measurement's publication.
- baseline
Demand reads 3 from documented load: automated own-initiative hateful-conduct restricted-reach labels run to tens of thousands per Member State per period, across 24 official languages, on a designated very large online platform. Capacity reads 1 from documented staffing: 1,352 people with primary-language proficiency cover 8 of the 24 languages, 1,197 of them in English; a secondary list of 185, stated to be not distinct from the first, adds three more; thirteen official languages appear in neither list.
- baseline
The two review channels are drawn from the documented review mechanism, never from the overturn table. The operator states that where it needs additional language support it uses translation services and/or machine translation tools, so one desk reads the content and the other reads a rendering — that is the whole modelled difference, and the translation step is drawn as its own model node because it is a documented automated component standing inside the review pathway. The published overturn rates were not used to set any rung: they do not sort by staffing (Swedish, with no named reviewer, at 72.2 percent; Hungarian, also with no named reviewer, at 0.0; English, with 1,197 reviewers, at 40.4), and this network asserts no relation between staffing and error anywhere.
- baseline
The appeal pathway is drawn at the low rung from the deployment's own table: per-language appeal rates against automated hateful-conduct visibility filtering ran 0.0 to 14.0 percent in the period 1 October 2024 to 31 March 2025, with 0.0 percent in Bulgarian, Greek and Lithuanian — languages whose overturn cells are blank, meaning no error measurement exists for them at all. The specialists' per-item correction is drawn substantive (second and third opinions, language-expert specialists, in-house counsel are all documented); what the record shows failing is reach.
- baseline
The overturn rate is treated as a directional, sourced, language-indexed signal and never as a probability or population error rate: it is conditional on appeal, its denominator (the appeal rate) is itself language-dependent, and the same metric family in the same report takes values above 100 percent for other policies, which marks it as a period ratio. The platform's own table footnote on blank cells is ambiguous, and the internally consistent reading (blank overturn cells align with the three 0.0 percent appeal-rate columns) is the one drawn here and quoted in the case file.
- baseline
There is deliberately no pathway from the translation-mediated flow to the primary-language desks. Drawing one would assert that a fluent reader is available for those thirteen languages, which is exactly what the operator's own staffing tables state is not the case; the documented recourse is escalation to specialists with language expertise, drawn as the wider of the two escalation edges.
- baseline
The two latent checks are the absences the record documents. A sampled accuracy audit of enforcement output — an error measurement independent of appeal, which is what Article 15(1)(e)'s 'indicators of the accuracy and the possible rate of error' would require in substance — appears in no reporting period; what is published is conditional on appeal, so zero-appeal cohorts return nothing. And no per-language audit of the moderation label pool's composition or quality is published by this operator or by any platform in the independently audited set.
- baseline
The regulator-side check is drawn present at the low rung, and the rung is derived from the measured behaviour of the disclosure loop rather than from any judgement of the actor: the Commission demonstrably compels (the table exists), demonstrably fines (December 2025, 120 million euros, on transparency obligations — the blue-checkmark design, the advertising repository and researcher data access, not the accuracy tables), and holds open proceedings whose grounds include content moderation (December 2023, extended January 2026). Meanwhile the language-indexed table was replaced by a country table in the very next period and the first harmonised filing covered 7 of 24 languages, with no visible enforcement response to that narrowing in the record read here. No regulator has ruled on this deployment's per-language enforcement accuracy, and nothing here implies one has.
- baseline
The label-pool loop is the structural asymmetry claim this network carries, and it is the operator's own account: models are trained on labels generated by trained human content moderators and on cases those moderators reviewed, and the named language capacity behind that pool is 88.5 percent English — so the thirteen translation-mediated languages run on a pool their own desks contribute least to. The training-pool edges are drawn from those statements; no part of this loop is derived from the overturn table.
- baseline
The egress to the EU statements-of-reasons database is drawn present rather than latent because it is a mandated, continuous, documented crossing: every moderation decision files a statement of reasons under Articles 17 and 24(5), and the database's own landing figure records roughly 42 percent of submitted decisions across all platforms as fully automated. This is the rare boundary crossing that exists because the law requires it, and closing it is not within any actor's authority on this board — which is why no connection-closing lever is offered in the paired scenario.
- assumed
Served people are not in the dynamics. The people whose posts are moderated, the languages they write in, the reach of their posts and the outcomes of their appeals are boundary quantities recorded in the case file — including the published per-language rates, which are recorded there as external observations with their denominators. No protected characteristic is represented anywhere, no content is described, and no served-person outcome is computed from anything drawn here.
- baseline
Every per-language magnitude in this network's anchors is the platform's own legally mandated self-report, filed under Commission scrutiny and exposed to fining power, and is entered as vendor-tier evidence. The independent sources in this record — the harmonised-report audit (an unrefereed preprint, tiered accordingly) and civil-society monitoring — verify the reporting ecosystem, not the enforcement system's ground truth. No independent measurement of this deployment's per-language enforcement accuracy exists in the record read here.
- baseline
Two things this network once drew as pathways of their own are carried in the description of a neighbouring element, and neither fact has left this page. The first is the playbooks: reviewers work from playbooks of colloquial terms and phrases updated to reflect EU languages and trends, so the label and phrase material the primary-language desks write into the training pool also reads back into their review. The record documents the playbooks, not that read-back as a route by which enforcement error travels, so it rides on the reviewer-labels pathway rather than as a separate return arrow. The second is the single model family: the same models enforce the same policy in all 24 official languages at once, which is what the one classifier element stands for, and the pool-to-classifier pathway carries the consequence that a gap in the pool repeats across a whole language cohort. The source model this network is derived from draws the first as a narrow read-back at an estimated width and has no counterpart to the second.
What this example does not show
- Every per-language magnitude here is the platform's own claim in a legally mandated self-report, filed under Commission scrutiny — vendor-tier evidence entered as such. The independent sources verify the reporting ecosystem (what platforms disclosed, and how the disclosures compare), not the enforcement system's ground truth. No independent measurement of this deployment's per-language enforcement accuracy exists in the record read here.
- The overturn rate is conditional on appeal, its denominator (the appeal rate) is itself language-dependent, and the same metric family in the same report exceeds 100 percent for other policies — it is a period ratio, never a probability or population error rate. The published overturn rates also do not sort by reviewer staffing (Swedish, with no named reviewer, at 72.2 percent; Hungarian, also with no named reviewer, at 0.0; English, with 1,197 reviewers, at 40.4), and neither this scenario nor its diagram asserts any relation between staffing and error.
- The December 2025 fine of 120 million euros (EUR) against this operator concerns transparency obligations — deceptive verification design, the advertising repository, and researcher data access — and is not a ruling on moderation accuracy, hate-speech enforcement, or the language tables. The Commission proceedings opened in December 2023, whose grounds include content moderation, remain open with no published findings, and are framed here as exactly that.
- Jurisdiction is EU-anchored, not US: this measurement exists because of an EU legal obligation (DSA Articles 15(1)(e) and 42(2)) and has no US analogue in this sector's record.
- No served-person outcome is modeled. The people whose posts are moderated, the languages they write in, and the outcomes of their appeals are boundary quantities recorded in the case file — including the published per-language rates, recorded there with their denominators as external observations. No protected characteristic appears anywhere in this record, and none is inferred.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Under Digital Services Act Articles 15(1)(e) and 42(2) — which require indicators of the accuracy and possible rate of error of automated content moderation, broken down by each official Member State language — X published, for the period 1 October 2024 to 31 March 2025, appeal and overturn rates for automated hateful-conduct visibility filtering across twenty named EU languages: appeal rates from 0.0 to 14.0 percent and overturn rates from 0.0 percent (Hungarian) to 72.2 percent (Swedish), with English at 40.4 percent. The independent audit of the sector's harmonised filings — a different set of reports from the April 2025 table above — found that of the eight largest EU platforms, four reported language-wise classification metrics (Facebook and Instagram across all 24 official languages, LinkedIn 22, X 7), three substituted country, and one reported none. Those are self-reported classification metrics; this atlas, not the audit, reads the April 2025 table as the only language-indexed error signal of its kind, because it indexes an outcome of the correction channel — appeals and overturns — rather than a platform's own accuracy figure. The figures are the platform's own legally mandated self-report, filed under Commission scrutiny, and no independent measurement of the deployment's per-language enforcement accuracy exists in the record.
empirical- Government Regulation (EU) 2022/2065 of the European Parliament and of the Council of 19 October 2022 on a Single Market For Digital Services (Digital Services Act), OJ L 277/1, Article 42(2) and Article 15(1)(e) https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32022R2065
- Vendor X Internet Unlimited Company (2025, April). DSA Transparency Report covering 1 October 2024 to 31 March 2025 (Indicators of Accuracy for Content Moderation; Linguistic Expertise of our Content Moderation Team) https://transparency.x.com/dsa-transparency-report-2025-april.html (archived 6 June 2026: https://web.archive.org/web/20260606101141/https://transparency.x.com/dsa-transparency-report-2025-april.html ; earlier capture 16 July 2025: https://web.archive.org/web/20250716055102/https://transparency.x.com/dsa-transparency-report-2025-april.html)
- Reference Trujillo, A., Tessa, B., & Cresci, S. (2026). Disarranged Harmonization of Transparency Reporting by Social Media Platforms Under the Digital Services Act. arXiv:2605.17655v1 [cs.CY] https://arxiv.org/abs/2605.17655
As of March 2025, X named 1,352 content moderators with primary-language professional proficiency covering 8 of the 24 official EU languages — 1,197 of them (88.5 percent) in English — plus a secondary list of 185 people, stated to be not distinct from the first, adding three more languages; thirteen official languages appear in neither list. For those, the company states: 'In situations where we need additional language support, we use translation services and/or machine translation tools, to investigate and address challenges in additional languages.' Flagged content is either human-reviewed before action or auto-actioned on the model's historical accuracy, and the models are trained on labels generated by the company's own trained moderators. Civil-society review found the same reviewer-language asymmetry sector-wide (for example Meta at 1 named Maltese reviewer against 3,110 Spanish) and concluded there is 'no clarity on moderation outcomes per country or per language.'
empirical- Vendor X Internet Unlimited Company (2025, October). DSA Transparency Report covering 1 April to 30 June 2025 (Indicators of Accuracy; Linguistic Expertise of our Content Moderation Team) https://transparency.x.com/dsa-transparency-report-2025-october.html (archived 21 July 2026: https://web.archive.org/web/20260721152250/https://transparency.x.com/dsa-transparency-report-2025-october.html ; earlier capture 30 October 2025: https://web.archive.org/web/20251030011908/https://transparency.x.com/dsa-transparency-report-2025-october.html)
- Advocacy International Network Against Cyber Hate (2025). Overview of the latest transparency reports under the DSA (executive summary brief) https://www.inach.net/wp-content/uploads/Transparency-Reports-2024_2025-DSA-2-1.pdf
The published per-language error indicator is an overturn rate conditional on appeal: its denominator, the appeal rate, varies by language from 0.0 to 14.0 percent, and the same metric family exceeds 100 percent for other policies in the same report, marking it as a period ratio rather than a probability or population error rate. The three languages with 0.0 percent appeal rates (Bulgarian, Greek, Lithuanian) carry no overturn rate at all — no error measurement exists for them — and four official languages (Croatian, Irish, Maltese, Slovak) are absent from the table entirely. The platform's own footnote on blank cells ('Cells that are blank mean that there was no enforcement. For cells containing 0.0% value, there were no cases of successful appeals or overturns') is ambiguous, and the reading used here — blank overturn cells aligning with the three zero-appeal columns — is the internally consistent one. The published overturn rates do not sort by reviewer staffing (Swedish, with no named reviewer, at 72.2 percent; Hungarian, also with no named reviewer, at 0.0; English, with 1,197 reviewers, at 40.4), and no monotone relation between staffing and measured error is supported by the record.
empirical- Vendor X Internet Unlimited Company (2025, April). DSA Transparency Report covering 1 October 2024 to 31 March 2025 (Indicators of Accuracy for Content Moderation; Linguistic Expertise of our Content Moderation Team) https://transparency.x.com/dsa-transparency-report-2025-april.html (archived 6 June 2026: https://web.archive.org/web/20260606101141/https://transparency.x.com/dsa-transparency-report-2025-april.html ; earlier capture 16 July 2025: https://web.archive.org/web/20250716055102/https://transparency.x.com/dsa-transparency-report-2025-april.html)
- Vendor X Internet Unlimited Company (2025, October). DSA Transparency Report covering 1 April to 30 June 2025 (Indicators of Accuracy; Linguistic Expertise of our Content Moderation Team) https://transparency.x.com/dsa-transparency-report-2025-october.html (archived 21 July 2026: https://web.archive.org/web/20260721152250/https://transparency.x.com/dsa-transparency-report-2025-october.html ; earlier capture 30 October 2025: https://web.archive.org/web/20251030011908/https://transparency.x.com/dsa-transparency-report-2025-october.html)
In the reporting period immediately following the twenty-language table (1 April to 30 June 2025), X published the same visibility-filtering accuracy indicators broken down by country rather than by language, and its first harmonised-template filing covered language-wise metrics for 7 of the 24 official languages per the independent audit; independent analysis of the sector's reports separately found their figures disconnected and their category vocabularies inconsistent across platforms. Every moderation decision files a statement of reasons to the Commission's public DSA Transparency Database, whose landing figure records roughly 42 percent of submitted decisions across all platforms as fully automated. Enforcement against the operator is active but distinct from these tables: the Commission's formal proceedings opened 18 December 2023 — content moderation and dissemination of illegal content among the grounds — remain open with no published findings and were extended on 26 January 2026 (recommender systems, plus a new investigation into the platform's integrated generative AI assistant), and the EUR 120 million fine of 5 December 2025, the first DSA non-compliance decision, concerns transparency obligations (deceptive verified-checkmark design under Article 25(1), the advertising repository under Article 39, researcher data access under Article 40(12)) — not moderation accuracy, hate-speech enforcement, or the language tables.
empirical- Vendor X Internet Unlimited Company (2025, October). DSA Transparency Report covering 1 April to 30 June 2025 (Indicators of Accuracy; Linguistic Expertise of our Content Moderation Team) https://transparency.x.com/dsa-transparency-report-2025-october.html (archived 21 July 2026: https://web.archive.org/web/20260721152250/https://transparency.x.com/dsa-transparency-report-2025-october.html ; earlier capture 30 October 2025: https://web.archive.org/web/20251030011908/https://transparency.x.com/dsa-transparency-report-2025-october.html)
- Reference Trujillo, A., Tessa, B., & Cresci, S. (2026). Disarranged Harmonization of Transparency Reporting by Social Media Platforms Under the Digital Services Act. arXiv:2605.17655v1 [cs.CY] https://arxiv.org/abs/2605.17655
- Government European Commission (2023, December 18). Commission opens formal proceedings against X under the Digital Services Act (press release) https://digital-strategy.ec.europa.eu/en/news/commission-opens-formal-proceedings-against-x-under-digital-services-act
- Government European Commission (2025, December 5). Commission fines X EUR 120 million under the Digital Services Act (press release) https://digital-strategy.ec.europa.eu/en/news/commission-fines-x-eu120-million-under-digital-services-act
- Reference eucrim, Max Planck Institute for the Study of Crime, Security and Law (2026). Overview of the Latest Developments Under the Digital Services Act: November 2025 - February 2026 https://eucrim.eu/news/overview-of-the-latest-developments-under-the-digital-services-act-november-2025-february-2026/
- Reference Ohnesorge, J. (2025, September 25). The DSA's transparency reports. Alexander von Humboldt Institute for Internet and Society (HIIG), Digital Society Blog, DOI 10.5281/zenodo.17201618 https://www.hiig.de/en/analysis-of-the-dsas-transparency-reports/
- Government European Commission. DSA Transparency Database (statements of reasons for content moderation decisions) https://transparency.dsa.ec.europa.eu/
Where this connects
Institutional pressures in this domain
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
All of them in context on the Content moderation & editorial AI domain page.
Levers available here and the patterns behind them
- Upgrade model — Improve the model
- Review the riskiest first — Risk-tiered oversight
- Review on schedule — Oversight cadence & retrospectives
- Understand the system — Understand the system
- Escalate checks — State-feedback vigilance
- Check copied records — Reconcile copied records
- Gate record entries — Human-in-the-loop write gating
- Mark AI-written records — Provenance labeling
- Assign a challenger — Structured dissent
- Gate vendor updates — Vendor quality gate
- Check with a second model — Cross-model verification
Documented case histories
- X Multilingual Hate-Speech Enforcement
- The errors that became visible when the reviewers went home
- The most built-out correction structure and the reach it doesn't have
- The byline nobody was behind
- A staff byline the AI wrote and the review it implied
- StopNCII & Take It Down
- X Community Notes (crowd annotation)
- GIFCT hash-sharing database
- Google CSAM detection and total account closure
- Meta cross-check: the enforcement-exemption tier
- The CyberTipline: triage under a rule against looking
- Sama Nairobi: the review workforce as the governed subsystem
- TikTok EU and UK trust-and-safety staffing substitution
- The score is published and the service cannot act on it
- YouTube Content ID