Skip to content

PAN Lab example

Community Notes on X, formerly Birdwatch on Twitter

The crowd writes the correction; the operator sets the bar

A post goes up. Somebody thinks it is misleading. In every other deployment in this domain the next thing that happens is that a classifier scores the post and something is removed, restricted or downranked. Here nothing is. A volunteer writes a note — a paragraph of context, usually with a cited source URL — and proposes it. Other volunteers rate that note. Every hour a matrix-factorization model is re-trained from scratch over the whole rating matrix, and if the note's fitted intercept reaches 0.40 with a latent factor magnitude under 0.50, the note appears beneath the post. The post itself is untouched. Modeled on the documented record of X's Community Notes, launched as Birdwatch in January 2021. The operator states the boundary in terms: notes rated helpful by enough contributors from different points of view will appear directly on posts, and beyond that, notes do not affect display of posts or enforcement of the platform's rules. It also disclaims the decision. It does not write, rate or moderate notes, and the mechanism does not work by majority rules — a note requires agreement between contributors who have sometimes disagreed in their past ratings. That requirement is arithmetic, not a slogan: intercept terms are regularized five times more strongly than the factor terms, which the operator's own documentation says is what makes a note need raters with diverse factors before it gets a label. Read what that buys and what it costs, because both are measured. It buys accuracy. In a randomly sampled set of notes on popular COVID-19 vaccine posts, evaluated with an infectious-disease physician and a virologist, 97.5 percent were entirely accurate and 0.5 percent inaccurate. It buys a real effect once a note lands: a synthetic-control study of 40,078 posts found 46.1 percent fewer reposts in post-attachment growth, and an independent difference-in-differences study of 237,180 cascades found 61.2 percent less subsequent spread and 94.3 percent higher odds the author deletes the post. What it costs is time and coverage. The same synthetic-control study measured its effect over the whole life of the post and it falls to 11.6 percent; the same difference-in-differences study puts the system-wide reduction in engagement at 14.9 percent, because notes often appear too late to intervene in the early and most viral stage of diffusion. Roughly half of a post's reposts happen in its first five hours and 80 percent within sixteen, and the median half-life of a post's impressions is about 79.5 minutes. Published estimates of how long a note takes run from a 14.3-hour median in one 2021-2023 sample to a 65.7-hour mean over 1.8 million notes — different clocks on different samples, and there is no single number. Notes attached beyond 47 hours are associated with a lifetime repost reduction of 0.1 percent, statistically indistinguishable from nothing, and with view and reply growth going up. And most notes never arrive. Across four years 87.7 percent stayed unrated and 8.3 percent were shown; only 13.55 percent of posts that had a note proposed ever carried one. The gate fails hardest where the stakes are highest, because the claims raters from opposite directions cannot agree on are the contested ones: an advocacy study of a purposive sample of 283 misleading 2024 election posts found 209 with accurate notes that were never shown, and the peer-reviewed effect is significantly weaker for political content and for influential accounts. Everything you have just read was measured by outsiders, from the operator's own daily public release, without permission — and none of it has produced a documented change to the corpus, the scorer or the thresholds. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets with the nine units you are given. On this board more budget is exactly what opens it. Every legal combination inside nine was checked — 8,708 of them under Service and Safety Targets, none winning. Raise the allowance one unit at a time and nothing changes at ten, eleven, twelve or thirteen. At fourteen, thirteen combinations win. The All Governance Targets level turns at sixteen. Here the obstacle is money. The cheapest winning set under Service and Safety Targets is five instruments: marking what a machine wrote, raising the bar when raters disagree about a note's quality, rules for who may write and rate, granting each connection out of the record explicitly, and keeping less of what the record keeps. Nothing on this board is out of reach. What nine buys is what this deployment is documented purchasing, and the distance between nine and fourteen is the finding: the instruments that would close it exist, they belong to the party that runs the scorer, and the record shows them unbought. Explore and Service Targets Only can be won, and cheaply: one instrument, costing two of your nine.

Stylized model of a documented deploymentContent moderation & editorial AI

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Rater-consensus-class contextual annotation network: 13 components and 25 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 1 assumed · 10 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • baseline

    D48-derived new org (Phase 6, content-moderation-editorial). REGISTER FIRST, because it governs every value here: this deployment's enforcement action is an ADDITION. A note appears on a post that stays fully available, and the operator's own documentation states that beyond appearing, notes do not affect display of posts or enforcement of the platform's rules. Nothing on this diagram removes, restricts, downranks or blocks anything, and no node classifies a post. What the machine scores is a NOTE, and what it decides is that note's visibility. The one documented consequence outside the display channel is the platform owner's October 2023 announcement that a noted post becomes ineligible for creator ad-revenue sharing, which sits in tension with the operator documentation and is carried in the scenario as a tension rather than resolved. Claims of algorithmic reach penalties for noted posts appear only on low-tier marketing pages, are supported by no credible source, and appear nowhere in this bundle.

  • baseline

    TOPOLOGY. Thirteen nodes, all documented, none decorative. TWO models because the record documents two machines doing different jobs: the published scorer that decides a note's visibility, and the automated writer agents admitted since July 2025 to propose candidate notes, which the pilot's own evaluation counts separately and which reach the displayed state at twice the human rate. THREE operator classes because the sources document three groups with different authority: the volunteers who write notes, the contributors whose ratings the scorer converts into the decision, and the company that owns every parameter and decides no case. THREE stores because the note-and-rating corpus, the display-status-and-timing record and the published scorer configuration are three different things the record measures separately, and the wiring between them is the deployment. ONE input source because the note-request file is a real, published, separately released demand signal. ONE reviewer because the record documents one external measurement channel, wired in on one read of the daily released record, which is the single act every one of its measurements rests on. ONE guardrail because the automated writer's test-mode admission screen is a real bounded automated screen with published pass marks. ONE worklist because the pool of notes awaiting ratings held 87.7 percent of every note written over four years. ONE external boundary because the daily release is a genuine one-way crossing out of the operator's control.

  • baseline

    ABSENCES ARE DERIVED TOO, and four of them are load-bearing. There is NO enforcement node, because nothing downstream acts on a note: the note is the whole action, and the operator's documentation says so in terms. There is NO appeal or contest channel and none is drawn, because there is none to draw — a post author's only documented channel is to request additional review of a note or report it, the remedy for a wrong or missing note is more ratings rather than adjudication, and post authors are served people who are boundary-only on every Lab diagram. There is NO retriever, because nothing retrieves; the scorer reads the whole matrix. And there is NO pathway from any operator class into either model: the operator's only route to the decision is an edit to a published configuration file that the machine then reads, which is why this network carries no operator-to-model edge at all. That absence is what makes the operator's power structural rather than per-case.

  • baseline

    THE REGULATOR IS DOCUMENTED AND IS DELIBERATELY NOT DRAWN, and this is the sharpest judgement in the file. The European Commission opened formal proceedings on 18 December 2023 whose grounds expressly include the effectiveness of this mechanism in the EU under Articles 34 and 35, and that limb remains open with no published finding; the Commission's first non-compliance decision against the platform, in December 2025, rests on entirely separate obligations. A reviewer node needs a documented inbound read, and the record documents none: no measurement of this deployment by the authority has been published, and the proceeding's own grounds cite the mechanism rather than any reading of its data. Drawing the authority as a reviewer would put an oversight channel on the diagram that has produced nothing, and drawing a filing pathway into it would cite the platform's OTHER Commission limb, which belongs to a different deployment and a different case file. The proceeding is carried in the scenario limitations and on the one check edge out of the measurement channel instead.

  • baseline

    WHERE THE LAB SHAPE DIVERGES FROM THE PAN SHAPE, and nothing is asserted here that the PAN file does not already record. Five divergences. First, the PAN entry carries the external measurement channel and the regulator as ONE user class; the Lab draws the measurement channel alone, for the reason above. Second, the PAN entry folds the note-request signal into the corpus-to-writer read; the Lab draws it as its own input source with two pathways, so the signal and the fact that nothing routes on it are two distinguishable statements rather than one clause. Third, PAN has no edge kind for a check and no external-boundary kind: the outside channel's reach into the parameter owner is a PAN peer edge redrawn here as a check, because a channel that improves the deployment is inhibiting in the Lab's vocabulary and reinforcing in PAN's, and its width still comes from PAN on the stated mapping; the reconciliation, the admission screen, the read of the rule and the daily release have no PAN counterpart and each says so on its own line. Fourth, PAN carries the admission screen and the rating backlog as attributes rather than components; the Lab draws them as the mediator nodes they are, which changes no flow because mediators carry none. Fifth, the Lab draws nine of PAN's twenty-seven edges inside a surviving pathway rather than as pathways of their own: the two store-to-store aggregations, three of the external reader's four inbound edges, the operator's and the writers' read-backs, the configuration read by the automated writers, and the raters' return to the writers. Each of those is named, with where its fact now lives, in the re-derivation entry below.

  • baseline

    BASELINES, and exactly how far the PAN org carries them. The PAN entry for this deployment holds twenty-seven edges. Eighteen of this network's twenty-five pathways have a one-to-one counterpart among them, and every one of those mirrors that edge's width on a single fixed four-rung mapping, stated in the derivation comment beside this network and applied with no exceptions. Nine further PAN edges are folded into a surviving pathway, which keeps its own counterpart's width; the folded edge's width is narrated in the survivor's copy and never added to it. The remaining seven are derived from the cited record directly. Four contrasts are load-bearing. The corpus reaches the scorer at the top rung while the scorer's output reaches the raters one rung lower, by published design, because status changes are delayed so a rater's judgement forms without knowing where a note stands. The measurement channel reads the released record at the middle rung, the same record the operator reads at the low rung, and writes back on one pathway that runs at zero, which is the file's central claim stated as a picture. The rule's self-loop runs at the top rung against an independent read of it at the low rung, because a rule everyone can read has still been evaluated only in simulation. And the request signal reaches writers at the low rung while its routing into the machinery runs at zero, which is the coverage failure drawn as two acts rather than narrated as a complaint.

  • baseline

    DEMAND 3 / CAPACITY 1. Demand 3 comes from the published corpus rather than any operator claim: 1,614,743 notes on 1,016,673 distinct posts by 227,702 contributors over four years, in 103 languages across more than 60 countries, on a platform the Commission's own designation puts at 112 million monthly active users in the EU alone — and nothing selects which posts the system looks at. Capacity 1 rests on three measured things and deliberately not on a fourth. It rests on throughput: roughly 908 ratings are consumed per displayed human-written note and a note needs at least five ratings even to leave the unrated pool. It rests on the operator's own published list of challenges, which names rater burden from high volumes of low-quality notes. And it rests on labour supply: publication measurably raises a first-time contributor's retention, so a publication rate reported as low and declining erodes the population that would raise it. It does NOT rest on the 87.7 percent of notes that never leave the unrated pool, because that share is what the consensus gate produces rather than what the crowd could not get to, and it is drawn on the pathways instead.

  • baseline

    EVIDENCE STATUS, labelled where it is used, because this record mixes tiers and because its three causal estimates measure three different things. PEER-REVIEWED CAUSAL: a synthetic-control study of 40,078 posts reports 46.1 percent fewer reposts in post-attachment growth and 11.6 percent over the post's whole lifespan, and a difference-in-differences study of 237,180 cascades reports 61.2 percent less subsequent spread against a 14.9 percent system-wide reduction in total engagement. PEER-REVIEWED NULL: an earlier program-level evaluation of the roll-outs found no aggregate reduction and attributed it to display latency exceeding the diffusion half-life. These are different estimands and their headline numbers are never pooled anywhere in this bundle; the conditional and the unconditional figure always travel together. PEER-REVIEWED ACCURACY: 97.5 percent of a randomly sampled set of displayed notes on one topic were entirely accurate — one topic, one period, and it says nothing about notes that were never displayed. WORKING PAPER: the 15.5-hour mean and 14.3-hour median latency pairing, tiered below the published items. PREPRINT: the whole-corpus statistics, the participation structure and the sustainability finding. SIMULATION: the manipulation result, which is an agent-based study of the published rule and not evidence that manipulation occurred. ADVOCACY, PURPOSIVE SAMPLE: the 74 percent election-coverage figure, attributed on every use. OPERATOR-REPORTED: the pilot-era figures that people were 25 to 34 percent less likely to like or repost, superseded for causal purposes and used only as the operator's own claim.

  • baseline

    THERE IS NO SINGLE LATENCY FIGURE, and any copy that states one is wrong. Published central estimates of the delay run from a 14.3-hour median in one 2021-2023 sample, through roughly 23 hours implied by a 2023 synthetic-control sample's speed quartiles, 24.29 and 26 hours in two later reports, and 2.23 to 2.85 days by rollout period, to a 65.7-hour mean over 1.8 million notes. They differ by clock (post to note creation, post to display, note creation to first status), by sample (all notes, displayed notes, notes near the threshold) and by period. Every use in this bundle carries the figure with its clock and its sample, or states the range. The one quantity that is NOT a latency is the automated-writer pilot's 6.0-hour median time to verdict: its clock starts at note creation, and the source says so, so it is never compared with the post-to-display figures.

  • assumed

    SERVED PEOPLE ARE NOT IN THE DYNAMICS, and here that boundary has an unusual consequence worth stating rather than hiding. The entire measured effect of this deployment runs through readers deciding not to repost — and readers, annotated post authors, and the people exposed during the hours before a note appears are served people who appear nowhere on this diagram. So the pathway that carries what this system actually does to the world is the one pathway the Lab may not draw. Nothing here computes a reader's decision, a post's reach, an author's deletion, or a harm to anyone. Every engagement figure in this bundle is a recorded external observation from a published study, and the two coverage figures — 13.55 percent of noted posts ever carrying a note, and 74 percent of accurate notes on a purposive sample of election claims never shown — are recorded external observations of coverage, never of error and never model-derived.

  • baseline

    RE-DERIVED AT THE COARSEST FAITHFUL GRANULARITY (2026-09-22). This board was first drawn with thirty-four pathways and is now drawn with twenty-five. No node was merged or removed, no surviving pathway changed width, and no documented fact left the page: nine pathways that drew an act already drawn elsewhere were folded into the pathway that draws it. The hourly conversion of ratings into status is drawn once, as the rating matrix into the scorer and the scorer's write of the status, and the return of the status history into the next hour's training rides the same read because the status history is one of the corpus's own files. The outside channel's separate reads of the display decision, the corpus and the status history are drawn as its one read of the daily release, because the record documents one download and not three arrangements. Its read of the published rule is drawn on the independent read of the rule, because both drew the same agent-based study. The operator's read-back of its own constants rides its edit of them. The writers' read of the notes already on a post rides their write of notes. The published rules that bound the automated writers ride the quota pathway. And the raters' verdict returning to writers rides the display outcome reaching writers, because what reaches a writer is whether the note was shown. Each folded pathway's width is narrated in the survivor's copy rather than added to it.

What this example does not show

  • LITIGATION AND REGULATORY POSTURE, carried BYTE-EXACT from the evidence dossier's own status field, which is where this deployment's posture lives because the dossier records no separate litigationPosture: "Active and expanding. The mechanism is operating, its code and data remain public, it has been copied by at least one other very large platform, and it is under an open EU regulatory proceeding that names it. No litigation anywhere concerns it." The dossier's own advisory correction extends that with two facts and no more: the open European Commission proceeding of 18 December 2023 names the system by name and remains without any published finding, and the Commission's December 2025 fine against the platform concerns entirely different obligations.
  • The European Commission's formal Digital Services Act proceedings of 18 December 2023 expressly name this mechanism among the grounds, assessing the effectiveness of measures taken to combat information manipulation on the platform under Articles 34(1), 34(2) and 35(1). That limb has produced no preliminary finding and no decision. The Commission's July 2024 preliminary findings and its 120 million euro non-compliance decision of 5 December 2025 rest entirely on the blue checkmark design, the advertisement repository and researcher data access, and touch this mechanism nowhere. The fine must never be attached to this system. Nothing here may be read as an adjudicated finding about the mechanism, because there is none.
  • The three causal studies do not measure the same thing and their headline numbers are never pooled. The 46.1 percent figure is a reduction in post-attachment repost GROWTH under synthetic control over 40,078 posts in a 2023 window; the 61.2 percent figure is a reduction in subsequent spread under difference-in-differences over 237,180 cascades; the 49.1 percent figure is a difference-in-differences estimate in a working paper. Their whole-lifespan and system-wide counterparts are 11.6 percent, 14.9 percent and 16.34 percent. Both numbers travel together everywhere in this bundle, and each is named with its estimand. An earlier peer-reviewed program-level evaluation of the roll-outs found NO aggregate reduction at all and attributed the null to display latency exceeding the diffusion half-life; that is not a contradiction of the later work, it is a different estimand, and the discrepancy is itself the finding.
  • There is no single latency figure for this deployment and any copy that states one is wrong. Published central estimates run from a 14.3-hour median in one 2021-2023 sample, through roughly 23 hours implied by a 2023 sample's speed quartiles, 24.29 and 26 hours in two later reports, and means of 2.23 to 2.85 days by rollout period, to a 65.7-hour mean over 1.8 million notes. They differ by clock, by sample and by period, and each use here carries its clock and its sample or states the range. The automated-writer pilot's 6.0-hour median time to verdict is NOT one of these: its clock starts at note creation, its own source says so, and it is never compared with a post-to-display figure.
  • Effect estimates are period-bound and the system has changed underneath them. The synthetic-control data ends in June 2023, and its own authors note subsequent changes including faster scoring and automatic propagation of notes to matching images and links, and say those changes are likely to lead to significant additional reductions. Nothing in this bundle presents a 2023-window estimate as the current performance of the system.
  • The publication-rate figures have different denominators and are not interchangeable: 8.3 percent of NOTES reached the displayed state and 13.55 percent of noted POSTS ever carried one over the four-year corpus; other samples put the note figure at 11.3 and 11.5 percent, one report at about 10 percent and declining, and a newspaper count at under 9 percent of notes written in 2024. Every use states whether the denominator is notes or posts.
  • The 74 percent election-coverage figure is advocacy-organisation research on a purposive, non-random sample of misleading election posts, not a population estimate, and it is attributed on every use. Its publisher's site serves a captcha interstitial to automated requests, so the figure is carried through trade coverage that quotes it verbatim. The accompanying total-views figure is reported inconsistently across summaries and is not carried anywhere in this bundle.
  • The manipulation result is a SIMULATION of the published scorer on synthetic data and not evidence that manipulation occurred. No source in this record documents a successful real-world coordinated suppression campaign. The operator itself names coordinated manipulation as a crucial risk for open rating systems; the platform's owner has publicly claimed the system was gamed, which is an unverified assertion by an interested party and is not repeated here.
  • There is a live tension in the record about whether a note carries a sanction, and this bundle flags it rather than resolving it. The operator's documentation states that notes rated helpful by enough contributors from different points of view will appear directly on posts and that beyond that, notes do not affect display of posts or enforcement of the platform's rules; one study's identification strategy relies explicitly on the platform not using notes to reduce visibility. Against that, the platform's owner announced on 29 to 30 October 2023 that posts carrying a note become ineligible for creator ad-revenue sharing. Both facts are carried. Claims of large algorithmic reach penalties for noted posts appear only on low-tier marketing pages, are supported by no credible source, and appear nowhere here.
  • The accuracy measurement is a research letter on a randomly sampled set of notes about ONE topic in one period, evaluated with an infectious-disease physician and a virologist. It is strong evidence that displayed notes are usually accurate. It is not a general accuracy rate for the program, and it says nothing at all about the notes that were never displayed — which are 87.7 percent of them.
  • Another company's adoption of this model, and that company's Oversight Board advisory opinion of 26 March 2026, are CONTEXT ONLY. The Board is that company's body, its opinion is non-binding, it addresses that company's expansion plans rather than this deployment, and nothing here may imply that any oversight body has ruled on this system. The adoption itself is described as begun and under review rather than complete or global, because that is what the record says.
  • Operator-reported effect figures — that people who saw bridging-selected annotations were 25 to 34 percent less likely to like or repost, and 20 to 40 percent less likely to agree with the substance of a misleading post — are vendor-tier, come from platform-run experiments during the pilot period, and are superseded for causal purposes by the independent studies. They appear only as the operator's own claim, labelled.
  • Served people are not modeled. No reader's decision, no post's reach, no author's deletion and no exposure during the latency window is computed from anything on this diagram, and the boundary has an unusual bite here: the entire measured effect of this deployment runs through readers changing their own behaviour, and readers appear nowhere. Every engagement figure in this bundle is a recorded external observation from a published study, and the two coverage figures are recorded external observations of coverage, never of error.
  • This board is one of two on the same platform in this catalogue and asserts nothing about the other. The sibling deployment is in-house classifiers performing hateful-conduct detection with visibility filtering and removal, a paid primary-language reviewer workforce and an appeals ladder, and it exists as a case because of statutory language-indexed disclosure duties. This one adds context to a post that stays fully available, has no appeal body at all, and is evidenced by the operator's own open code and daily corpus. The December 2023 Commission proceeding lists both limbs; each case file cites only its own, and neither may claim a finding.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • Community Notes, launched as Birdwatch in January 2021 and deployed worldwide from 11 December 2022, is a crowd-annotation system on X in which volunteer contributors write contextual notes on specific posts and other contributors rate those notes, with an open-source matrix-factorization model deciding which notes appear. An independent parse of the operator's own daily public corpus for 23 January 2021 to 23 January 2025 counts 227,702 unique contributors writing 1,614,743 notes on 1,016,673 distinct posts, detected in 103 languages, with the program available in more than 60 countries. Participation is highly unequal and largely monolingual: the top 10 percent of contributors wrote 58 percent of all notes, a Gini coefficient of 0.68; one apparently automated account wrote 33,186 notes; and only about 16 percent of note authors ever wrote in more than one language. Contributor admission is by three published criteria — an account at least six months old, a verified phone number from a trusted carrier not associated with another Community Notes account, and no recent notice of violations of the platform's rules — with random selection from country-specific waitlists where applicants exceed available slots. All contributions are pseudonymous and publicly visible. Around 22 to 26 May 2025 notes stopped appearing in users' feeds for several days following a data-centre outage, with the operator's engineering account acknowledging continuing issues and the Community Notes account stating on 26 May that it was working to get notes appearing normally.

    empirical
    • Academic Mohammadi, S., et al. (2026). From Birdwatch to Community Notes, from Twitter to X: four years of community-based content moderation (arXiv:2510.09585v4) https://arxiv.org/abs/2510.09585
    • Academic Razuvayevskaya, O., et al. (2025). Timeliness, consensus, and composition of the crowd: Community Notes on X (arXiv:2510.12559) https://arxiv.org/abs/2510.12559
    • Trade press TheJournal.ie (2025, May). Community Notes vanishes from X feeds, raising serious questions amid ongoing EU probe https://www.thejournal.ie/x-community-notes-disappeared-from-platform-factchecking-misinformation-elon-musk-6717349-May2025/
  • The enforcement action here is an addition, not a subtraction, and the operator's own documentation states it. Its published FAQ says notes rated helpful by enough contributors from different points of view will appear directly on posts, and that beyond that, notes do not affect display of posts or enforcement of X's Rules. Its introduction page states that X does not write, rate, or moderate notes, except where a note itself violates the platform's rules, and that the mechanism does not work by majority rules: a note requires agreement between contributors who have sometimes disagreed in their past ratings. Nothing is removed, downranked, or restricted by the note itself, and the entire measured effect runs through readers changing their own behaviour. There is no appeal body, no human reviewer of last resort, and no escalation queue; a post author's only documented channel is to request additional review of a note or report it, and the remedy for a wrong or missing note is more ratings rather than adjudication. One documented consequence sits outside the display channel and in tension with the operator documentation: on 29 to 30 October 2023 the platform's owner announced that posts carrying a Community Note become ineligible for creator ad-revenue sharing, framing it as maximizing the incentive for accuracy over sensationalism and asserting that attempts to weaponize notes to demonetize people would be immediately obvious because the code and data are open. Both statements are carried here and the tension between them is flagged rather than resolved. Claims of large algorithmic reach penalties for noted posts appear only on low-tier marketing pages, are supported by no credible source, and are not asserted anywhere in this atlas.

    empirical
    • Vendor X Corp. Community Notes documentation, About: FAQ (repository documentation source) https://raw.githubusercontent.com/twitter/communitynotes/main/documentation/about/faq.md
    • Trade press Bell, K. (2023, October 30). X won't pay creators for tweets that get fact checked with community notes. Engadget https://www.engadget.com/x-wont-pay-creators-for-tweets-that-get-fact-checked-with-community-notes-174206477.html
    • Academic Renault, T., Restrepo Amariles, D., & Troussel, A. (2024). Collaboratively adding context to social media posts reduces the sharing of false news (arXiv:2404.02803; HEC Paris Research Paper LAW-2024-1519). Working paper, no journal publication located https://arxiv.org/abs/2404.02803
  • A note is displayed only when a published consensus gate is cleared, and most notes never clear it. The operator's ranking documentation gives the rule: a matrix factorization fits a global intercept, a per-rater intercept and factor, and a per-note intercept and factor; intercept terms are regularized at 0.15 against 0.03 for the factor terms, five times more strongly, which the documentation says is what requires that notes are rated by raters with diverse factors before a note gets a label. A note is Currently Rated Helpful at intercept 0.40 or above with latent factor magnitude under 0.50, raised to 0.50 by a tag-outlier filter; at least five ratings are needed to leave the Needs More Ratings state; the model is re-trained from scratch every hour, with status changes deliberately delayed so as not to influence independent raters. Measured against the whole public corpus for 2021 to 2025, 87.7 percent of notes remained in Needs More Ratings and 8.3 percent reached Helpful, and only 13.55 percent of posts with at least one proposed note ever received a Helpful note. Independent samples put the note-level rate at 11.3 percent, at 11.5 percent, and at about 10 percent and declining, and a newspaper analysis counted roughly 79,000 of more than 900,000 notes written in 2024 shown publicly, under 9 percent; these denominators are notes rather than posts and are not interchangeable. The gate is weakest where the stakes are highest. The Center for Countering Digital Hate reported on 30 October 2024 that 209 of 283 sampled misleading US-election posts, 74 percent, had accurate notes that were never shown to all users; that is advocacy-organisation research on a purposive, non-random sample rather than a population estimate, its publisher's site is captcha-blocked to automated requests, and the figure is carried through trade coverage quoting it verbatim. An archival analysis of more than 1.8 million notes finds 69 percent of noted posts receiving conflicting classifications from contributors, and about 68 percent annotated as not needing a note at all. A regression-discontinuity analysis finds that having a note published raises the retention of first-time contributors, so a low and declining publication rate erodes the contributor population that would raise it.

    empirical
    • Vendor X Corp. Community Notes documentation, Under the Hood: Ranking Notes (repository documentation source, Apache-2.0) https://github.com/twitter/communitynotes/blob/main/documentation/under-the-hood/ranking-notes.md
    • Academic Mohammadi, S., et al. (2026). From Birdwatch to Community Notes, from Twitter to X: four years of community-based content moderation (arXiv:2510.09585v4) https://arxiv.org/abs/2510.09585
    • Academic Razuvayevskaya, O., et al. (2025). Timeliness, consensus, and composition of the crowd: Community Notes on X (arXiv:2510.12559) https://arxiv.org/abs/2510.12559
    • Academic Arjmandi-Lari, M., Mantzarlis, A., & Stafford, T. (2025). Threats to the sustainability of Community Notes on X (arXiv:2510.00650, submitted to ICWSM) https://arxiv.org/abs/2510.00650
    • Advocacy Center for Countering Digital Hate (2024, October 30). Rated Not Helpful: How X's Community Notes system falls short on misleading election claims (publisher site captcha-blocked; figures carried through Social Media Today's 30 October 2024 write-up at https://www.socialmediatoday.com/news/reports-find-community-notes-failing-address-misinformation-x-formally-twitter/731558/ ) https://counterhate.com/research/rated-not-helpful-x-community-notes/
  • The deployment's documented failure mode is timing, and there is no single latency figure for it. Published central estimates of the delay between a post and a note appearing beneath it are: a mean of 15.5 hours and a median of 14.3 hours in a 2021-2023 difference-in-differences sample; quartile boundaries of 12, 23, and 47 hours from post creation to note attachment in a March-June 2023 synthetic-control sample, implying a median near 23 hours; means of 2.85 days after the US rollout and 2.23 days after the worldwide rollout, with the shortest display delay observed anywhere in that dataset being 80.2 minutes; an average of 26 hours in a whole-corpus parse; and a mean of 65.7 hours across 1.8 million notes. These differ by clock — post to note creation, post to display, or note creation to first status — by sample, and by period, and each use must carry its clock and its sample or state the range. Against that clock, roughly 50 percent of a post's reposts occur in its first 5 hours and 80 percent within 16, and the median half-life of a post's impressions is about 79.5 minutes. The dose-response is measured and monotone: notes attached in the 1-12, 12-23, 23-47, and 47-plus hour quartiles are associated with lifetime repost reductions of 24.9, 12.3, 4.3 and 0.1 percent, the last statistically indistinguishable from nothing, and in that slowest quartile view growth rose 13.6 percent and reply growth 27.0 percent, consistent with a late note drawing attention back to a stale post. The earliest rigorous evaluation, a difference-in-differences and regression-discontinuity analysis of the US and worldwide roll-outs, found no evidence that introducing Community Notes reduced aggregate engagement with misleading posts and attributed the null to display latency exceeding the diffusion half-life.

    empirical
    • Academic Renault, T., Restrepo Amariles, D., & Troussel, A. (2024). Collaboratively adding context to social media posts reduces the sharing of false news (arXiv:2404.02803; HEC Paris Research Paper LAW-2024-1519). Working paper, no journal publication located https://arxiv.org/abs/2404.02803
    • Academic Slaughter, I., Peytavin, A., Ugander, J., & Saveski, M. (2025). Community notes reduce engagement with and diffusion of false information online. Proceedings of the National Academy of Sciences, 122(38), e2503413122 https://arxiv.org/abs/2502.13322
    • Academic Chuai, Y., Tian, H., Proferes, N., Zhang, K., & Lenzini, G. (2024). Did the roll-out of Community Notes reduce engagement with misinformation on X/Twitter? Proceedings of the ACM on Human-Computer Interaction, 8, CSCW2, Article 428 https://doi.org/10.1145/3686967
    • Academic Mohammadi, S., et al. (2026). From Birdwatch to Community Notes, from Twitter to X: four years of community-based content moderation (arXiv:2510.09585v4) https://arxiv.org/abs/2510.09585
    • Academic Razuvayevskaya, O., et al. (2025). Timeliness, consensus, and composition of the crowd: Community Notes on X (arXiv:2510.12559) https://arxiv.org/abs/2510.12559
  • Once a note is attached the mechanism works, and every effect figure must carry both its conditional and its unconditional counterpart because the three causal studies measure different estimands and their headline numbers must never be pooled. A synthetic-control study of 40,078 posts for which notes were proposed between 16 March and 23 June 2023, of which 6,757 (16.9 percent) received a Helpful note, estimated post-attachment GROWTH reductions of 46.1 percent in reposts, 44.1 percent in likes, 21.9 percent in replies and 13.5 percent in views, and WHOLE-LIFESPAN reductions of 11.6, 13.3, 6.9 and 5.5 percent respectively; it also found noted content's repost cascades becoming less deep and less structurally viral than matched counterfactuals, and it measured no deletion outcome at all. An independent difference-in-differences study of 237,180 fact-checked cascades reposted more than 431 million times estimated a 61.2 percent reduction in subsequent spread and a 94.3 percent increase in the ODDS that the author deletes the post, against a SYSTEM-WIDE reduction of 14.9 percent in total engagement with misleading posts, stating that notes often appear too late to intervene in the early and most viral stage of diffusion, and reporting the effect as significantly weaker for posts from influential accounts and for political content. A working paper on about 285,000 notes reports 49.1 percent fewer retweets by difference-in-differences and 52.4 percent by pre-treatment outcome matching, with overall reductions of 16.34 percent in retweets, 11.75 percent in replies and 16.87 percent in quotes once publication delay is accounted for, and a deletion rate of 15.8 percent just above the 0.4 helpfulness threshold against 8.6 percent just below it — a relative gap across a threshold, not a causal probability increase. The notes themselves are usually right: in a randomly sampled set of notes on popular COVID-19 vaccine posts evaluated with an infectious-disease physician and a virologist, 97.5 percent were entirely accurate, 2 percent partially accurate, and 0.5 percent inaccurate, with 49 percent citing highly credible sources and 44 percent moderately credible ones. That accuracy measurement covers one topic in one period and says nothing about the notes that were never displayed. Operator-reported pilot figures — that people who saw bridging-selected annotations were 25 to 34 percent less likely to like or repost, and 20 to 40 percent less likely to agree with the substance of a potentially misleading post — are vendor-tier claims from platform-run experiments and are superseded for causal purposes by the independent studies.

    empirical
    • Academic Slaughter, I., Peytavin, A., Ugander, J., & Saveski, M. (2025). Community notes reduce engagement with and diffusion of false information online. Proceedings of the National Academy of Sciences, 122(38), e2503413122 https://arxiv.org/abs/2502.13322
    • Academic Chuai, Y., et al. (2026). Community-based fact-checking reduces the spread of misleading posts on X (formerly Twitter). Nature Communications, 17, 4070 https://doi.org/10.1038/s41467-026-72597-0
    • Academic Renault, T., Restrepo Amariles, D., & Troussel, A. (2024). Collaboratively adding context to social media posts reduces the sharing of false news (arXiv:2404.02803; HEC Paris Research Paper LAW-2024-1519). Working paper, no journal publication located https://arxiv.org/abs/2404.02803
    • Academic Allen, J., Desai, N., Namazi, A., Leas, E., Dredze, M., Smith, D. M., & Ayers, J. W. (2024). Characteristics of X (formerly Twitter) Community Notes addressing COVID-19 vaccine misinformation. JAMA, 331(19), 1670 https://doi.org/10.1001/jama.2024.4800
  • The decision rule and the whole operating record are public, and that is why an independent evidence base for this deployment exists. The scoring code, its documentation, and its note-writer interface template sit in a public repository under the Apache-2.0 licence, and five data files — notes, ratings, note status history, contributor enrolment, and note requests — are released daily on a best-effort basis, cumulative, containing only items created up to 48 hours before release, with each participant carrying a Community-Notes-specific pseudonymous identifier that is stable across handle changes; deleted content disappears from the downloads while the note status history retains participant identifiers and status timelines. Three independent research teams produced causal estimates of the system's effects from exactly that data without the operator's permission, a civil-society organisation measured coverage failure on a purposive sample of election claims, a university team parsed four years of the corpus, and an agent-based study probed the manipulation surface of the deployed rule itself rather than a guess at it, reporting that under polarization and in-group rating preference the published scorer suppresses a substantial fraction of genuinely helpful notes and that a coordinated minority of 5 to 20 percent of raters could strategically suppress targeted helpful notes — a simulation on synthetic data, not a measurement that manipulation occurred. The operator's own published list of challenges names four risks it watches: coordinated manipulation as a crucial risk for open rating systems, outcomes dominated by a simple majority or biased by the distribution of contributors, harassment of contributors, and rater burden from high volumes of low-quality notes. Speed and coverage — the two properties every independent measurement identifies as the deployment's failure mode — are not among them. The openness is also what made the mechanism copyable: on 7 January 2025 Meta announced it would end third-party fact-checking in the United States in favour of a community-notes-style model, and on 18 March 2025 it began testing on Facebook, Instagram, and Threads with about 200,000 signed-up contributors, stating it would use X's open source algorithm as the basis of its rating system and that, unlike the fact-check labels it replaced, notes would provide extra context but would not impact who can see the content or how widely it can be shared. That adoption is context about another company and is described as begun and under review rather than complete or global.

    empirical
    • Vendor X Corp. Community Notes documentation, Under the Hood: Download Data (repository documentation source) https://github.com/twitter/communitynotes/blob/main/documentation/under-the-hood/download-data.md
    • Vendor X Corp. Community Notes documentation, About: Challenges (repository documentation source) https://raw.githubusercontent.com/twitter/communitynotes/main/documentation/about/challenges.md
    • Academic Truong, B. H., Wu, X., Flammini, A., Menczer, F., & Stewart, A. J. (2025). Community Notes are vulnerable to rater bias and manipulation (arXiv:2511.02615). Agent-based simulation of the published scorer https://arxiv.org/abs/2511.02615
    • Vendor Kaplan, J. (2025, January 7). More Speech and Fewer Mistakes. Meta Newsroom. https://about.fb.com/news/2025/01/meta-more-speech-fewer-mistakes/
  • Since 1 July 2025 the writing side of the mechanism has admitted automated agents while the deciding side has stayed human. The operator's published interface documentation states that automated writers propose notes while humans still decide what is helpful enough to show, and that ratings come from regular contributors, that is humans, whose input ultimately determines which notes show. Admission is earned in test mode against an automated evaluator that screens URL validity and whether a note addresses a claim rather than an opinion: at least 95 percent of the candidate's most recent 50 test notes must score high on URL validity and at least 30 percent high on the claim-versus-opinion measure. Daily writing limits start at 10 and scale between 2 and 500 with measured helpfulness. A think-tank evaluation computed from the operator's public downloads for September 2025 to March 2026 counts 27 automated writer accounts enrolled and 24 active, submitting 31,464 notes: 7.4 percent of note volume but 13.9 percent of displayed notes, reaching Currently-Rated-Helpful at 18.0 percent against 8.9 percent for human writers, rated Not Helpful at 2.3 percent against 4.1 percent, and consuming roughly 304 ratings per displayed note against 908 for human-written ones, with monthly output growing from 93 notes in September to 8,109 in February. Its median time-to-verdict of 6.0 hours against 6.3 for humans is measured from note creation to first non-Needs-More-Ratings status and is therefore not comparable to the post-to-display latencies measured elsewhere. The evaluation explicitly measured no engagement outcome, so nothing follows from it about whether automated notes reduced spread, and it records that the enrolment criteria for automated accounts are undisclosed. The operator's interface documentation does not state that automated notes are visibly labelled to readers, although press coverage of the July 2025 launch reported that they would be; that labelling is therefore press-attributed rather than operator-confirmed.

    empirical
    • Vendor X Corp. Community Notes documentation, API overview (automated note-writer pilot) https://raw.githubusercontent.com/twitter/communitynotes/main/documentation/api/overview.md
    • Trade press Purnell, S. (2026, June 16). AI writers on Community Notes: an evaluation of seven months of data. R Street Institute https://www.rstreet.org/research/ai-writers-on-community-notes-an-evaluation-of-seven-months-of-data/
  • One regulator has named this mechanism and none has published a finding on it. The European Commission's first formal proceedings under the Digital Services Act, opened 18 December 2023 against a platform designated a Very Large Online Platform on 25 April 2023 with 112 million monthly active users in the EU, list among their grounds the effectiveness of measures taken to combat information manipulation on the platform, notably the effectiveness of X's so-called Community Notes system in the EU and the effectiveness of related policies mitigating risks to civic discourse and electoral processes, against Articles 34(1), 34(2) and 35(1). That limb remains open. The Commission's preliminary findings of July 2024 were limited to the blue-checkmark design, the advertisement repository, and researcher data access, so no preliminary finding has ever issued on the Community Notes limb; and the Commission's first non-compliance decision against the platform, a 120 million euro fine of 5 December 2025, rests on those same three grounds and concerns this mechanism nowhere. That fine must never be attached to Community Notes. A Commission spokesperson declined to comment on the mechanism in May 2025 because of the ongoing proceedings while confirming the investigation into its effectiveness. No litigation of any kind concerns this deployment. On 26 March 2026 the Oversight Board issued a policy advisory opinion on Meta's plans to expand community notes beyond the United States, concluding that delays in note publication, the limited number of published notes, and the model's dependence on the broader information environment's reliability raise serious doubts about the extent to which community notes can meaningfully address misinformation linked to harm, and recommending that countries with repressive human-rights records, imminent major elections, histories of coordinated disinformation networks, active crises or conflicts, unsupported language complexity, or persistent internet-access obstacles be omitted or delayed. That opinion is non-binding, it is that company's own body, it addresses that company's expansion plans rather than X's system, and its evidence is X's published performance figures. It is context here and is not a ruling about this deployment.

    empirical
    • Government European Commission (2023, December 18). Commission opens formal proceedings against X under the Digital Services Act (press release) https://digital-strategy.ec.europa.eu/en/news/commission-opens-formal-proceedings-against-x-under-digital-services-act
    • Government European Commission (2025, December 5). Commission fines X EUR 120 million under the Digital Services Act (press release) https://digital-strategy.ec.europa.eu/en/news/commission-fines-x-eu120-million-under-digital-services-act
    • Trade press TheJournal.ie (2025, May). Community Notes vanishes from X feeds, raising serious questions amid ongoing EU probe https://www.thejournal.ie/x-community-notes-disappeared-from-platform-factchecking-misinformation-elon-musk-6717349-May2025/

Where this connects

Institutional pressures in this domain

  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).

All of them in context on the Content moderation & editorial AI domain page.

Levers available here and the patterns behind them

Documented case histories