Skip to content

Domain Atlas

Content moderation & editorial AI

This domain covers two AI deployments that both decide what the public sees: automated content moderation on platforms, and AI-drafted editorial content in newsrooms. In moderation, machine classifiers remove content proactively — often before any user has seen it — at a scale no human review could match, and the governing fact is that the human review and appeals path is the error-correction loop, not an optional add-on. A platform's own natural experiment made this concrete: when human reviewers were sent home during the pandemic and the platform deliberately chose over-enforcement, automated removals more than doubled, appeals roughly doubled, and the reinstatement rate on appeal jumped from about a quarter to about half — direct evidence that the automation was making roughly twice the rate of catchable errors, visible only because the appeals queue caught them. Two things follow. Proactive removal acts before anyone sees the content, so an over-broad takedown is invisible unless an appeals path surfaces it — and in documented cases automated removal destroyed evidence of war crimes with archival access declined, an irreversible action with no correction loop at all. And over-enforcement versus under-enforcement is a chosen trade-off: when you cannot review everything, you are choosing which error to make, and that choice is a governance decision, not a technical default. In the newsroom, AI-drafted articles published under a human byline without disclosure are an accountability failure of a different shape — one outlet's audit found it had to correct a large share of its AI-written articles, so the review that a byline implies was not actually performed. The Lab networks model only the deploying organization — its classifiers or drafting tools, its reviewers, editors, and appeals functions, and its enforcement or publication records; the people whose content is moderated or who read the articles sit outside the dynamics, and no user or reader outcome is computed on any diagram.

Use cases

What AI is doing here

Automated content enforcement

Predictive

Machine classifiers and hash-matching that remove content proactively — often before any user has seen it — at a scale no human review could match, where the removal happens before there is any signal it was wrong, so an over-broad takedown is invisible unless an appeals path surfaces it.

Appeals & the error-correction loop

Predictive

The human review, appeals queue, and external-oversight functions that catch the errors automated enforcement makes at scale — the error-correction loop a natural experiment showed is load-bearing, since removing human review roughly doubled the rate of removals later reinstated on appeal.

Editorial AI drafting

Generative

AI that drafts published editorial content, often under a human byline — where the byline implies a review the reader trusts, and an outlet's own audit finding it had to correct a large share of its AI-written articles shows that review was not actually performed and the AI use was not disclosed.

Case files

What has gone wrong and right

Documented deployments, presented as model organizations calibrated to the evidence, with full citations.

X Community Notes (crowd annotation)

United States (X Corp., Bastrop, Texas), operating globally: available in more than 60 countries with notes detected in 103 languages. The scoring code and its documentation are published by the operator at github.com/twitter/communitynotes under Apache-2.0, and five data files are released daily. The only binding regulatory scrutiny of the mechanism is European: the European Commission opened formal Digital Services Act proceedings on 18 December 2023 whose grounds expressly include the effectiveness of this system in the EU, under Articles 34(1), 34(2) and 35(1), against a platform designated a Very Large Online Platform on 25 April 2023 with 112 million monthly active users in the EU. That limb remains open with no published finding. The Commission's first non-compliance decision against the platform, a 120 million euro fine of 5 December 2025, rests on the blue checkmark design, the advertisement repository and researcher data access, and concerns this mechanism nowhere.

The one large-scale content-moderation system in this atlas whose enforcement action takes nothing down. Volunteer contributors write a note on a post; other contributors rate it; an open-source bridging algorithm publishes the note only when raters whose past ratings point in opposite directions both call it helpful, and the whole note-and-rating corpus is released daily so anyone can grade the result. What appears is usually right — 97.5 percent of a physician-and-virologist-reviewed sample of displayed notes was entirely accurate, and an attached note cut reposts by 46 to 61 percent in two independent causal designs. What appears is also late and rare: 87.7 percent of notes never clear the consensus gate, a purposive sample of 2024 US election posts found 209 of 283 with accurate notes that were never shown, and the median note lands after most of the resharing is done — which is why the same studies put the system-wide effect at roughly 12 to 15 percent. No litigation anywhere concerns it, and the only regulator that has ever named it opened a proceeding in December 2023 and has published no finding.

Explore this deployment in the PAN Lab →

GIFCT hash-sharing database

United States. The Global Internet Forum to Counter Terrorism is a US-incorporated independent nonprofit, founded in 2017 by Facebook, Microsoft, Twitter and YouTube and constituted as a standalone organisation in 2020, with predominantly US-headquartered members. Its effect is global, because matching runs on member platforms worldwide. No regulator supervises the consortium as such. The instruments that do bind fall on the MEMBERS individually: Regulation (EU) 2021/784 on addressing the dissemination of terrorist content online (adopted 29 April 2021, applicable from 7 June 2022, one-hour removal orders on hosting providers), the EU Digital Services Act, and the UK Online Safety Act 2023. The Christchurch Call, launched 15 May 2019 and supported by 55 governments plus the European Commission and 19 online service providers as of August 2026, names 'the expansion and use of shared databases of hashes and URLs' among its industry commitments and is politically binding on nobody.

Thirty-nine platforms share one index of terrorist and violent extremist content: approximately 2.4 million hashes covering approximately 408,000 unique and distinct items at the end of 2025, on the consortium's own count, with three quarters of the behaviourally labelled entries in the taxonomy's least determinate category, 'glorification of terrorist acts'. One company's judgement about one item on its own service becomes an input to every other member's upload comparison. Only the company that contributed a hash may take it out; the Global Internet Forum to Counter Terrorism, which owns the taxonomy, is recorded by its own commissioned reviewers as lacking the ability to modify even a label. And because the database stores hashes rather than content, there is nothing for an outside auditor to inspect — a point the peer-reviewed literature and the consortium's own reports make in almost the same words. The one erroneous entry ever described in public was two hashes of a music video that was not violent, graphic or explicit, caught because a second company happened to look.

Explore this deployment in the PAN Lab →

Google CSAM detection and total account closure

United States. The reporting duty is federal, under 18 U.S.C. section 2258A, which requires a provider to report an apparent violation as soon as reasonably possible after obtaining actual knowledge and which simultaneously disclaims at subsection (f)(3) any duty to monitor any user or to affirmatively search, screen or scan; failure-to-report penalties run from $600,000 to $1,000,000 depending on the offence and the provider's user base. 47 U.S.C. section 230(c)(2)(A) immunises voluntary good-faith restriction of objectionable material. The two documented cases sit in San Francisco, California and Houston, Texas, and both were investigated and closed by local police. The quantitative appeal telemetry in this file comes from a DIFFERENT jurisdiction and a different service scope: Google Ireland Limited's annual reports under Regulation (EU) 2021/1232 and, from the 2025 reporting period, Regulation (EU) 2024/2916, which cover Google's messaging and mail services in the European Union and NOT Google Photos or YouTube. Two United States decisions on this deployment shape both went the operator's way on different questions: a pleading-stage dismissal of a different, self-represented plaintiff's challenge to a closure (D.D.C., 26 July 2024) and a state supreme court holding that Google scanned as a private actor rather than as a government agent (Wisconsin, 24 February 2026).

A machine-learning classifier built to surface never-before-seen child sexual abuse material flagged two fathers' photographs of their toddlers' genital infections, taken because a nurse and a pediatrician asked for them. A trained specialist reviewer confirmed each flag against the federal definition, Google filed a CyberTipline report, and Google closed each man's entire account: a decade of Gmail, Drive documents, the whole photographic record of a child's early years, contacts, and in one case phone service from Google Fi. Police investigated both men and cleared both; one report concluded that 'the incident did not meet the elements of a crime and that no crime occurred'. Both appealed with the clearance in hand and neither was reinstated. Google's own regulated filings explain why the appeal could not absorb it: reinstatements come 'not due to an error in detection or a content-level false positive' but from context about intent, and its reports to the clearinghouse are, in the company's words, 'one-way reporting'. Nothing in the pipeline malfunctioned. The variable that would have changed the answer lived outside every input the system had.

Explore this deployment in the PAN Lab →

Meta cross-check: the enforcement-exemption tier

United States (Meta Platforms, Inc., Menlo Park, California), operating a global programme across Facebook and Instagram. There is no litigation and no regulatory adjudication of this programme anywhere. The reviewer is the Oversight Board, a quasi-judicial body Meta established and funds through an irrevocable trust: its decisions on individual cases bind, its policy recommendations do not, and it holds no audit, subpoena or compulsion power. It published a policy advisory opinion on the programme on 6 December 2022 at Meta's own request, and Meta published a 60-day response on 6 March 2023. European Union exposure is Digital Services Act-shaped and adjacent rather than about this programme: on 24 October 2025 the European Commission issued preliminary findings, conducted with Coimisiun na Mean, on notice-and-action mechanisms, dark patterns, appeal mechanisms and researcher data access, with exposure up to 6 per cent of worldwide annual turnover if confirmed. Preliminary findings do not prejudge the outcome.

The atlas's inverse moderation case. Every other deployment in this domain fails by acting on content it should have left alone; this one fails by routing some accounts' identified content out of enforcement into a queue that was known to be under-resourced, so the protection and the error land on different populations by construction. Listed accounts' flagged posts stayed fully up through as many as five layers of human reconsideration — a mean of more than five days, about twelve for United States content, about seventeen for Afghanistan and Syria, and 222 days at the extreme — while everyone else's likely-mistaken removals went through unreviewed when the same overstretched reviewers ran out of hours. Meta's external review body found the programme had never checked whether the exception path was actually more accurate than the ordinary one, cleared the backlog by asking, and was then refused on exactly the five recommendations that would have let anyone outside see who was protected or count what the delay cost.

Explore this deployment in the PAN Lab →

The CyberTipline: triage under a rule against looking

United States: a federal statutory clearinghouse under 18 U.S.C. § 2258A, operated by a congressionally authorised private non-profit and funded principally through the Department of Justice's Office of Juvenile Justice and Delinquency Prevention. The flow is global: in calendar year 2025, 77.1 per cent of reports resolved to a location outside the United States, and reports were made available to law enforcement in 170 countries and territories as well as to 61 United States Internet Crimes Against Children task forces and to federal agencies. The Fourth Amendment doctrine that decides what the clearinghouse's own analysts may examine is circuit-dependent and unsettled: the Second, Fourth and Ninth Circuits, and the Tenth in the alternative holding that opened the question, on one side; the Fifth and Sixth Circuits and the Supreme Court of Wisconsin on the other. A certiorari petition was filed on 14 April 2026 and distributed on 17 June 2026 for the conference of 28 September 2026, and as of 28 August 2026 the Supreme Court has neither granted nor denied it. The REPORT Act (Pub. L. 118-59, 7 May 2024) added minor sex trafficking and enticement to the reporting duty, extended preservation from 90 days to one year, raised failure-to-report penalties to between $600,000 and $1,000,000 scaled by offence and user base, and extended the clearinghouse's limited liability to contracted vendors. No court or regulator supervises the clearinghouse's performance; what the courts supervise is its epistemic reach.

The atlas's statutory-funnel case, and the only deployment in it whose central constraint is constitutional rather than technical. Federal law makes reporting mandatory, detection voluntary, and every field that would make a report usable optional at the sender's sole discretion; it then requires the clearinghouse to make every report it receives available to law enforcement, with no discretion to filter. So a private non-profit triages 21 million reports a year on information the senders may withhold — and, because a court held it part of the government for search purposes, its analysts open only the files a platform employee already opened. In one appellate record a platform reviewer opened 31 of 156 flagged files and the clearinghouse employee opened exactly those 31; the report then went to the wrong county and sat for half a year. The channel that would tell anyone which reports were worth the hours exists, is well designed and is voluntary: federal agencies returned 7,085 pieces of feedback on 3.4 million reports.

Explore this deployment in the PAN Lab →

Sama Nairobi: the review workforce as the governed subsystem

Kenya, and this is not a United States deployment. The Employment and Labour Relations Court at Nairobi and the Court of Appeal at Nairobi. The employer of record is Samasource Kenya EPZ Limited, trading as Sama, a Kenyan export processing zone company whose parent is headquartered in San Francisco; Meta Platforms, Inc. and Meta Platforms Ireland Limited are named as respondents and dispute being employers. Two petitions were brought: Petition E071 of 2022, filed on 10 May 2022, and Constitutional Petition E052 of 2023, brought by the dismissed reviewers; they were later consolidated. The litigation is ACTIVE WITH NO MERITS DETERMINATION. On 20 September 2024 the Court of Appeal issued two judgments the same day that went in opposite directions: [2024] KECA 1262 dismissed the client's jurisdiction appeals with costs, so the claims proceed in a Kenyan labour court, and [2024] KECA 1152 allowed the appeals against the interim ruling of 2 June 2023 and set it aside in its entirety together with all consequential orders, including the order requiring proper medical, psychiatric and psychological care. Rulings expected on 12 February 2026 were not delivered and the court adjourned on notice without fixing a date. The deployment itself is concluded: the employer left content moderation in 2023 and the work moved to a successor supplier; in April 2026 the client ended its remaining annotation contract at the same delivery centre.

The domain's inversion, and the reason it earns a file. Everywhere else in content moderation the human reviewer is the remedy — the correction channel that catches what the classifier — the software that sorts items into policy categories — gets wrong. Here the review workforce is where the harm lands. Approximately 200 reviewers covered a sub-continent across roughly eleven African languages, working to a reported fifty-second average handling time and an eighty-four per cent accuracy score enforced against the same person from opposite directions, with the one hour of weekly wellness break approved by the managers who owned the throughput. Throughput was metered continuously; psychological load was measured once, four years later, by a hospital psychiatrist, for a court. The party that set the targets, wrote the policy, supplied the tool and composed the queue is not the party that employed, insured or medically supported the people who met them, and it defends the claims by saying so. Four years after the first filing, a Kenyan appellate court has confirmed the claims may be heard against the foreign client and has set aside every interim protection the workers had won — on the same day — and no merits ruling has been delivered.

Explore this deployment in the PAN Lab →

TikTok EU and UK trust-and-safety staffing substitution

European Union-anchored, with the United Kingdom as a named second jurisdiction. Not a United States deployment. The EU limb runs under Regulation (EU) 2022/2065, the Digital Services Act, enforced for very large online platforms by the European Commission alone, with Coimisiún na Meán as the Digital Services Coordinator of Ireland, TikTok's country of establishment; the designated filer is TikTok Technology Limited, Dublin, designated a VLOP on 25 April 2023 at 135.9 million EU monthly active recipients. Moderation sites in the record sit at Dublin, Amsterdam and Berlin, with German labour law and the Berlin Labour Court governing the Berlin dispute, and Coimisiún na Meán separately operating a binding Online Safety Code for video-sharing platforms whose Part B applied from 21 July 2025. The UK limb runs through TikTok Information Technologies UK Limited at Canary Wharf, London, regulated by Ofcom since 2021 under the Video Sharing Platform regime and now under the Online Safety Act 2023, with UK employment law, ACAS and the House of Commons Science, Innovation and Technology Committee in the record. The parent is ByteDance. The US national-security and divestiture material that surrounds that ownership concerns a different jurisdiction and a different question and is outside this file entirely.

A platform under two open European Commission investigations into whether it manages its systemic risks well enough reduced its EU content-moderation workforce from more than six thousand people at the end of 2023 to 4,596 at the end of June 2025, while its European audience grew about a quarter. It told a House of Commons committee that of approximately 430 London roles at risk, roughly a third were the people who labelled the training data for the automated moderators replacing them — because, it said, progress in the development of those models had significantly decreased the need for that labelling — and another significant proportion were ancillary roles including the teams that train moderators on the Community Guidelines. The Commission had already put on the record that the Digital Services Act 'does not prescribe any specific rules about the resources to be dedicated to content moderation'. Neither proceeding names moderation staffing as a ground. The parties who actually extracted a public account of what was happening were a German union that struck for four days, a London union-recognition ballot the redundancy notices preceded by eight days, and a select committee that asked for the safety risk assessment, was not given one, and published that fact.

Explore this deployment in the PAN Lab →

The score is published and the service cannot act on it

United States for the infrastructure and the non-profit operator — the Wikimedia Foundation, San Francisco, owns and runs the servers — and NOT for the governance. The models serve more than 250 Wikipedia language editions and the better-calibrated multilingual model covers 47 of them; the operator's own reverting agent runs on twelve Wikipedias and not on English; and the crowd audit that best demonstrates the governance mechanism happened on Spanish Wikipedia. Bot approval, bot blocking, the granting and revocation of the rollback right, threshold selection and tool adoption are all exercised by the self-governing volunteer community of each language edition. Litigation posture: None. There is no court, no regulator, no consent order and no statutory transparency mandate anywhere in this record, and that is a substantive structural finding about this governance arrangement rather than an absence of controversy.

A machine-learning service scores essentially every Wikipedia edit for damage as it is saved and publishes the result to anyone who asks, without a credential — and has no way of its own to revert, block, warn or report anything. Acting on the score is somebody else's job, under somebody else's governance: a volunteer-run bot on English Wikipedia holding user groups the community granted and can revoke, an operator-built agent that twelve other Wikipedias have each switched on for themselves, and volunteer patrollers working score-filtered queues at a threshold their wiki chose from a table of published consequences. The measured results cut both ways and both are public. Showing the flag narrowed the gap between how registered and unregistered editors are treated; the classifier behind it flags unregistered editors at a disparate impact ratio of 20.02 against a base-rate ratio of 7.93. There is no regulator, no court, no consent order and no reporting duty anywhere in this record — every instrument here exists because the operator and the communities chose it, and the one that made threshold-setting participatory was removed in an infrastructure migration as an old, unused feature.

Explore this deployment in the PAN Lab →

YouTube Content ID

United States for the operator — YouTube, LLC, a Delaware company with its principal place of business in San Bruno, California, per its own federal complaint — and for the litigation and criminal record cited (N.D. Cal., D. Neb., D. Ariz.). The deployment is global and its published figures are worldwide with no country breakdown, a disclosure gap the contemporaneous academic-policy analysis flags by name. No regulator has ordered a change to this system and no court has found it unlawful anywhere as of 28 August 2026. The one United States case that attacked its access structure produced no finding: class certification was denied on 22 May 2023 and, on 12 June 2023, the day trial was to begin, the parties stipulated to dismissal with prejudice of all claims raised or that could have been raised. The European dimension is real and is context rather than a United States regulatory fact: Article 17 of the Copyright in the Digital Single Market Directive and the Digital Services Act supply reporting and redress obligations in Europe, and on the analysis of the Communia and Kluwer commentators they are part of why this transparency report exists at all. In the United States the report is voluntary.

A fingerprint comparison runs over every video uploaded to YouTube and checks it against reference files supplied by rights-holders admitted through an access gate. On a match, that rights-holder's standing instruction fires automatically — block the video, take its advertising revenue, or track its figures — with no case-by-case human decision on the claiming side. YouTube processed 2,502,941,368 such claims in calendar 2025 on its own count, 99.48 per cent of every copyright action taken on the platform that year, and over 90 per cent of claims take the money rather than the video. About half a per cent of claims are ever disputed, and a dispute is answered by the party that made the claim: it has thirty days, and its silence releases the claim. Every rung the uploader climbs raises the chance that party converts the matter into a legal removal request, which carries a copyright strike, and three strikes in ninety days ends the account. Of 45,724 failed appeals in one half-year, 13,841 produced a removal and 31,883 ended because the uploader cancelled the appeal or deleted the video. The whole funnel is published by the deployer, voluntarily, and it contains no measure of the claims that were wrong and never contested.

Explore this deployment in the PAN Lab →

System map

Who is in the system and what pushes on it

Who is in the system

  • Frontline workers. Caseworkers, screeners, eligibility staff — the operator network whose judgment the system augments or erodes.
  • Supervisors & QA. The institutional correction layer: overrides, second reads, quality review.
  • Agency leadership. Owns procurement, policy, and the authority map; answers for the system publicly.
  • Served people & families. Those the decisions land on. Deliberately outside the PAN dynamics — their outcomes are measured, never simulated.
  • Regulators & oversight bodies. Boards, auditors, data-protection officers, inspectorates — external correction capacity.
  • Advocates & community organizations. Surface harms institutions do not see; historically the earliest accurate signal.

Dominant pressures

  • Reviewer bottleneck. One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Austerity & recovery incentives. Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Vendor opacity. The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift. The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance. Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

Governance

Questions leaders should be asking

  1. 1. Automated enforcement removes content at a scale no human review could match, and often before anyone has seen it — so is the human review and appeals path resourced as the error-correction loop it actually is, or treated as an optional add-on to an automated decision that is really the decision?
  2. 2. When you cannot review everything, over-enforcement and under-enforcement are a chosen trade-off — you are deciding which error to make — so is that choice being made deliberately and owned as a governance decision, or defaulting to whatever the classifier does at the threshold someone set once?
  3. 3. A proactive takedown acts before any user sees the content, so an over-broad removal is invisible unless an appeals path surfaces it — and some removals (evidence of atrocities, for instance) are irreversible with no preservation path; is anyone measuring the errors the automation makes before they are seen, and preserving what cannot be un-removed?
  4. 4. An AI-drafted article published under a human byline implies a review that the byline stands behind — so when a large share of such articles later needs correction, was the review actually performed, and is the use of AI disclosed to the reader who trusts the byline?

For the actions behind these questions, see the Practice Library.

Seeing your organization in this domain? Mapping its actual pathways, pressures, and correction capacity is engagement work.

Work With Paramerge