PAN Lab example
Wikipedia's edit-scoring service (ORES, now Lift Wing)
The score is published, the threshold belongs to the governed
A scoring service computes a damage probability for essentially every edit as it is saved, publishes it to anyone who asks without a credential, and has no way at all to act on it. Modeled on Wikipedia's edit-quality scoring, its successor revert-risk models, and the separately governed agents that consume them. Everything a platform pipeline holds in one place is pulled apart here: the classifier is the operator's, the acting agent is a volunteer's or a community's, the threshold belongs to the wiki that will live with it, and the correction channel is a public wiki page. The numbers are published before the choice is made — the strictest quality filter on English Wikipedia is right more than nine times in ten and finds about a tenth of the problem edits; the broadest catches about 82 percent and is right about 15 percent of the time. So does the cost of the trade: one agent's maintainer moved his false-positive budget from 0.25 percent to 0.1 and gave up roughly fifteen points of catch rate to do it. Both fairness results are public too, and they disagree. Showing the flag narrowed the gap between how registered and unregistered editors are treated; the classifier behind it flags unregistered editors at more than twice their already higher base rate. Watch what happens to the measurement, not just to the enforcement: no regulator, court or reporting duty exists anywhere in this record, so every one of these instruments survives only as long as somebody keeps choosing it. Before you pick a target level: Explore and Service Targets Only can be won, and cheaply: one instrument, costing two of your ten. Service and Safety Targets can be won, but only by spending all ten of your units, and only one combination of instruments does it: four of them, one at full strength. All Governance Targets cannot be won on this budget, and the obstacle is money. Every gate but the pathway gate can be cleared under All Governance Targets for two of your ten. Clearing that gate too takes thirteen units, and at thirteen only two combinations win, each of five instruments with one at full strength. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Scoring-service class whose operating point is set by the governed network: 12 components and 28 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 2 assumed · 8 published baseline · 3 measured. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the open-infrastructure scoring pattern documented in the wikipedia-ores case file, not a reconstruction of the actual systems. It is the one deployment in this sector where the score is published for any item to anyone without a credential, the operating point is chosen by the community that will live with it, the models and their training data are open with their biases stated on the operator's own cards, and the service that computes the score has no way to act on it. Those four properties are the reason it is on the board.
- baseline
Demand reads 3 from a published volume: English Wikipedia alone took between 2.59 and 2.94 million non-bot edits to content pages per month over the twelve months to July 2026, on the order of 85,000 to 95,000 human content edits a day on one of more than 250 editions the models serve, with a scoring job starting for nearly every edit saved. The operator's own arithmetic on the all-editions figure is that reviewing roughly 290,000 edits a day at ten a minute would be about 483 volunteer labour hours daily.
- measured
Capacity reads 3 because the counterfactual was measured rather than reasoned, and the measurement has two limits that are stated rather than hidden. Across four outages of the fastest automated tier in the first half of 2011, median time to undo an edit rose from 744 to 1,286 seconds and the geometric mean from 941 to 1,674, while the proportion of revisions eventually undone was statistically indistinguishable at chi-square 0.64, p equals 0.43: the human tiers absorbed the work more slowly. This quality-control system also predates the scoring service by a decade and the service was added to it rather than put in its place. The limits: the experiment removed ONE automated tier on ONE wiki at a scale and tooling mix that no longer obtain, and its authors declined to price what the workaround cost the editors who performed it. No measurement exists for removing the score itself.
- baseline
Every rung on this board follows one stated mapping from the PAN org's own per-edge rates, so the widths are arithmetic rather than taste: a rate of 0.55 or above is drawn 3, 0.30 to 0.54 is drawn 2, 0.15 to 0.29 is drawn 1, and a channel the record documents as absent is drawn 0. Six pathways carry no PAN rate because PAN has no vocabulary for them — the two input hops and the three inhibiting kinds — and each of those names its own frequency evidence in the derivation comment beside it.
- baseline
The score reaches the thing that acts on it through the public record and through nothing else, and that routing is the deployment's signature rather than a modelling convenience. The operator built the service to decouple curating training data, building models, auditing predictions and building the interfaces or bots that act on them, and it declined to gate access on its own stated reasoning that an open interface offers no barrier at which it could selectively decide who may request a score. The agent that reverts reads the same record an outside researcher reads.
- baseline
No external-boundary node is drawn, and the omission was tested rather than assumed. An egress pathway drains the Privacy gauge by construction, and this deployment's outbound flow carries no identifiable client data: scores are computed on public revisions and are themselves public, and the tools that consume them are built on that public record. Drawing an egress here would assert a data-protection concern the record denies. What the record does document is a measured disparate impact against unregistered editors, which is carried where it belongs — on the input pathway, as the one privacy-flagged edge on this board.
- measured
The two fairness measurements point in opposite directions and this network carries both, because choosing between them would misrepresent the record. A regression discontinuity across twenty-three language editions measured what the published flag does to the HUMAN decision and found it raised revert probability far more for registered editors, from 4.6 to 14.3 percent at the middle cutoff, than for unregistered ones, from 13.5 to 19.2 percent, narrowing a pre-existing gap; it also lowered the rate at which reverts of unregistered editors were themselves contested. A separate peer-reviewed evaluation measured the CLASSIFIER's own disparate impact at 20.02 against a base-rate ratio of 7.93 in the same data. A biased classifier whose published score reduced a larger human bias is the honest reading, and neither half is drawn without the other.
- measured
One pathway is drawn at 0 and it is the transparency mechanism the peer-reviewed record identifies as the core of participatory machine learning here. A machine-readable query let a wiki's tool developer ask the service for the best filter rate available at a stated recall, for their own wiki and their own model version; the published worked example on English Wikipedia was a threshold of 0.32 giving a filter rate of 0.89, a false-positive rate of 0.087, precision of 0.23 and recall of 0.75. Verified live on 2026-08-28, that query answers that the parameter is not supported by this endpoint any more. Nobody decided to end participatory threshold-setting: the query that implemented it was removed during an infrastructure migration, under the operator's own description of dropping some very old and not used features. Community threshold-setting survives through the other routes drawn here.
- baseline
The correction channel is drawn narrow, and the reason is what it is rather than how often it is used. A reverted editor is told a machine did it, in a message that concedes the model sometimes undoes good edits, and is pointed at a public page where the full list of reported errors also sits; creating that page is a required step of deploying the operator's agent. But it is a wiki page, not a case with a disposition, so this deployment publishes no reversal rate on contested reverts and cannot supply the overturn statistic a mandated transparency report supplies. Its correction channel is more open and less measured than a regulated platform's.
- baseline
Each documented mechanism is drawn once, and several reads are carried on the pathway that already holds them rather than drawn a second time. The outside comparison of the deployed model with a deliberately unfair alternative was computed from the same published record every consumer reads, and tool builders build against that record without asking the model's operator. The measured bias finding reached the agents' maintainers as a recommendation on the model card, which nothing obliges them to follow. The agents' maintainers tune against the reported errors on the public page and the catch rate they measure on their own reviewed slice, and both halves of that trade are written on the page. The disagreement sample the audit tooling draws is the reconciliation between the score record and the label pool, and the training provenance printed on each model card is part of that record. The volunteer bot's dataset is labelled by volunteers answering its maintainers' open call. A patroller's personal minimum score reaches that patroller's own queue and nobody else's. A change to the operating point arrives in an administrator's watchlist like any other edit. Retraining on reported errors is a stated intention with no shipped mechanism in the record.
- baseline
Two nodes carry no pathway of their own, as the schema's mediator rule provides. The queue is the score-filtered patrolling backlog, a throughput backlog with a published volume behind it, and it is tethered to the pathway from the scoring service to the patrollers. The exclusion filters are the acting agent's published carve-out list, applied to its own output before it acts, and the inhibiting pathway they stand for is drawn on the agent itself.
- assumed
Served people are not in the dynamics. The editors whose contributions are scored, undone, warned or thanked are boundary quantities recorded in the case file, including the measured newcomer costs of this quality-control system in the years before the scoring service existed. Those newcomer figures describe the harm this deployment was built in response to and are never attributed to it. No protected characteristic is represented anywhere on this board; the one demographic-adjacent property drawn is whether an editor is registered, which is an input feature of the deployed model and the axis on which its disparate impact was independently measured.
- baseline
The evidence tiers behind this board are mixed and are labelled where they matter. The acting agent's accuracy statistics are its volunteer maintainers' own published figures computed on their own held-out, human-reviewed slice: methodologically described, but self-reported rather than independently audited. The caution-level precision table is the builders' pre-deployment testing on 22 completed review spreadsheets covering over 600 edits across six projects, not production performance. The peer-reviewed base leans on a small overlapping author group with a long employment relationship to the operator; the evaluation that measures the deployed model unfavourably and the audit probe are the clearest checks outside that lineage. There is no regulator, court, consent order or statutory reporting duty anywhere in this record, which means every accountability artefact here exists because the operator and the communities chose it.
What this example does not show
- The Lab models institutional propagation through the operator network. It does not model editors, articles, readers, or what any individual edit contained, and it computes no outcome for anyone whose contribution was scored, undone or thanked.
- The service this is derived from no longer exists as infrastructure. The original servers were retired at the end of their planned lifespan in early 2024; what runs today is a successor serving platform with a compatibility endpoint in front of it, and the older edit-quality models are deprecated in favour of a revert-risk family. The board is a shape, not a live wiring diagram.
- The two fairness findings measure different layers and neither refutes the other. One is a causal estimate of what showing the flag does to a human moderator's decision; the other is a measurement of the classifier's own error distribution. Presenting either alone would misstate the record, and the Lab draws both.
- The newcomer figures in the case file describe this quality-control system in the years 2006 to 2011, before the scoring service existed. They are the harm the service was built in response to and are never a measured effect of it.
- The acting agent's accuracy statistics are its volunteer maintainers' own, computed on their own held-out human-reviewed slice, and the caution-level precision table is the builders' pre-deployment testing on 22 review spreadsheets covering over 600 edits across six projects. Both are methodologically described and neither is independently audited.
- Precision and recall here are per-wiki and per-threshold and travel badly. The figures quoted are English Wikipedia's; on Polish Wikipedia the comparable filter captures 91 percent of problem edits against 34 percent for the English one, which is why that community does not carry the broad filter at all.
- There is no litigation, regulator, court, consent order or statutory reporting duty anywhere in this record. That absence is a finding about this governance arrangement rather than an absence of harm, and it means every accountability instrument drawn here could be withdrawn the same way it was granted.
- The evidence base leans on a small overlapping group of authors, one of whom was employed by the operator for most of the period covered. The independent evaluation that measures the deployed model unfavourably and the audit probe are the clearest checks outside that lineage, and the Lab leans on them where it can.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
Wikipedia's edit-scoring service publishes a damage or revert-risk probability for essentially every edit as it is saved and has no path of its own by which it can act on one. Aaron Halfaker and R. Stuart Geiger, writing the system up for CSCW in 2020, describe it as built to decouple four activities normally performed by the same engineers: curating training data, building models, auditing predictions, and building the interfaces or bots that act on predictions. The operator did not gate access either, on its own stated reasoning that 'Given the open API, there is no barrier where we can selectively decide who can request a score from a classifier.' Acting on the score is done by separately governed agents: a volunteer-run bot on English Wikipedia, an operator-built agent local administrators switch on, tool-assisted patrollers working score-ranked queues, and ordinary editors with score-driven filters enabled in their own preferences. At the time of that paper the estate ran roughly 110 classifiers across 44 languages in four families; the ORES infrastructure has since been retired and what runs today is Lift Wing, with ores-legacy.wikimedia.org as a compatibility endpoint in front of it, serving a revert-risk family that replaced the older edit-quality models. Verified live on 28 August 2026: the legacy host answers, self-identifies as the 'ORES legacy service', and still returns scores, and the Lift Wing endpoint returns a revert-risk score for an arbitrary revision without any credential.
empirical- Academic Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189
- Vendor Albon, C., Wikimedia Foundation (2023, August 3). ORES To Lift Wing Migration (wikitech-l announcement), with the Wikitech ORES and Lift Wing pages and a live probe of both endpoints on 2026-08-28 showing the threshold-statistics query removed https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/EK65B7QCQHEG37C2ERPIUSP64OX3ZEUJ/
The scores, the models, the training-data provenance, and the deliberation about all three are public here, and that openness is what makes every outside measurement in this case possible. Scores are queryable for any revision by anyone without a credential, and were from 2015; the models and their code ship under the Apache 2.0 licence and the training pipelines were built as reproducible Makefiles so a third party can rebuild an equivalent model. The operator publishes model cards as a standing institutional practice, written to the Mitchell et al. framework, on a public index of proposed, production, and deprecated models; each card is asked to state why the model was made, its proper and improper uses, and its evaluation scores especially for marginalised groups, and each card page carries a talk page, so the documentation is also the venue where it is argued about. Every operating point's precision and recall is published in patroller-facing help text before anyone chooses among them. Against that openness sit two documented limits: historical predictions were retained only until the end of 2019, and because no operator adjudicates anything, a false-positive report is a public wiki page rather than a case with a disposition, so no reversal rate on contested reverts is published the way an overturn rate appears in a mandated transparency report.
empirical- Academic Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189
- Vendor Wikimedia Foundation. Machine learning models: Production language-agnostic revert risk and Production multilingual revert risk (Meta-Wiki model cards stating the deployed model's own bias against new users, temporary accounts and IP edits) https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk
- Vendor Wikimedia Foundation. Help:New filters for edit review, Quality and Intent Filters (MediaWiki.org; per-threshold precision and recall stated to patrollers, and the per-wiki divergence) https://www.mediawiki.org/wiki/Help:New_filters_for_edit_review/Quality_and_Intent_Filters
- Academic Levonian, Z., Hagen, L., Li, L., Lilleboe, J., Wastvedt, S., Halfaker, A., & Terveen, L. (2024). ORES-Inspect: A technology probe for machine learning audits on enwiki. Wiki Workshop 2024 (arXiv:2406.08453) https://arxiv.org/abs/2406.08453
Every operational lever in this deployment is held by the governed community rather than by the model's operator, and the record shows each of them being exercised. A bot may not edit at all before English Wikipedia's Bot Approvals Group has approved it and it has run a trial, and any administrator may block one that malfunctions. The rollback right that the fastest patrolling tool requires is granted and revoked by administrators, and that tool's own documentation states it 'is not intended for new Wikipedia users' and that misuse 'may result in revocation of rollback permissions or being blocked from editing'; the cross-wiki patrolling tool gates its global queue on global rollback, steward or global sysop rights, or at least 1,000 global edits and no active block, and its documentation records no scoring integration. Automoderator, the operator's own reverting agent, 'will not begin running until a local administrator turns it on', its caution level is written to MediaWiki:AutoModeratorConfig.json through Special:CommunityConfiguration — an ordinary watchlistable wiki page — and creating a false-positive reporting page is a required step of deployment. It is live on twelve Wikipedias (Indonesian, Turkish, Ukrainian, Vietnamese, Afrikaans, Bengali, Azerbaijani, Chinese, Spanish, Italian, Dutch, and Albanian, first Turkish in June 2024) and not on English, whose project page records the team's own position that 'If English Wikipedia editors don't want to use Automoderator, that's fine!' When a volunteer wired a score directly to automatic reversion on Spanish Wikipedia, the community stopped it without the operator: PatruBOT was crowd-audited on ordinary wiki pages, judged to be erring too often, and had its account blocked by an administrator — an episode the ORES team recorded as 'entirely a community governed activity that required no intervention of our team or the Wikimedia Foundation staff', with a successor, SeroBOT, later resuming at a higher confidence threshold.
empirical- Academic Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189
- Vendor English Wikipedia. Wikipedia:Bot policy and Wikipedia:Huggle; and Wikimedia Meta-Wiki, SWViewer (the community-granted and community-revocable permission to act quickly on a flagged edit) https://en.wikipedia.org/wiki/Wikipedia:Bot_policy
- Vendor Wikimedia Foundation. Moderator Tools/Automoderator and Moderator Tools/Automoderator/Testing (MediaWiki.org; the community-switch, the watchlistable configuration, the required false-positive page, and the caution-level table from the builders' own pre-deployment testing on 22 spreadsheets over about 600 edits) https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator
The acting agent's control variable is a human-chosen error budget rather than a score, and both halves of the trade it buys are published. ClueBot NG's documentation states that 'The threshold is not randomly chosen by a human, but is instead calculated to match a given false positive rate... A human selects a false positive rate, which is the percentage of constructive edits incorrectly classified as vandalism.' At the current 0.1 percent setting the bot catches approximately 40 percent of vandalism; at the previous 0.25 percent setting it caught approximately 55 percent, so roughly fifteen points of catch rate were given up to halve the wrongful-reversion rate. Chosen instead for total accuracy, it classifies over 90 percent of edits correctly. Post-processing filters on self-reverts, edit count, and warning share cut actual false positives below the stated rate before any revert happens. These figures are the volunteer maintainers' own, computed on a held-out, human-reviewed slice of their own dataset: methodologically described and not independently audited. The agent's scale and standing are separately verifiable: registered 20 October 2010, 6,682,890 edits as of 28 August 2026, holding the bot, reviewer, rollbacker, and autoconfirmed user groups, all community-granted, and stoppable by any administrator editing a run page to 'False'. The operator's own agent applies its own published carve-outs, never reverting administrators, global sysops, stewards or bots, self-reverts, reverts of its own actions, or new page creations.
empirical- Vendor English Wikipedia. User:ClueBot NG/Documentation (volunteer maintainers' own published statistics on their own held-out data; the false-positive-budget mechanism, the 0.1 and 0.25 per cent settings, and the dataset-bias concession) https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation
- Vendor Wikimedia Foundation. Moderator Tools/Automoderator and Moderator Tools/Automoderator/Testing (MediaWiki.org; the community-switch, the watchlistable configuration, the required false-positive page, and the caution-level table from the builders' own pre-deployment testing on 22 spreadsheets over about 600 edits) https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator
The deployed classifier's own bias against anonymous editors was measured independently and is large. Mykola Trokhymovych, Muniza Aslam, Ai-Jou Chou, Ricardo Baeza-Yates, and Diego Saez-Trumper reported at KDD in 2023 a Disparate Impact Ratio of 20.02 for the deployed ORES model against a base-rate ratio of 7.93 in the same data, meaning it flagged anonymous editors well out of proportion even to their genuinely higher revert rate; the successor multilingual model reaches 9.54 with the same user features and 1.98 to 3.08 without them. The deployed model's area under the curve was 0.84 with precision at recall 0.75 of 0.22 on an unbalanced holdout, against 0.75 and 0.07 for a deliberately unfair baseline that simply reverts every anonymous edit. Model accuracy correlates negatively with a language edition's share of anonymous editors in both systems. The operator states the same limitation in its own documentation: the model card for the language-agnostic revert-risk model — the model its own reverting agent uses — says the model 'may exhibit bias against edits from new users, temporary accounts, or IP edits' and recommends the multilingual model instead for anonymous edits in the languages that model covers, while forbidding use of its predictions as ground truth for training other models and forbidding scoring a page's first revision.
empirical- Academic Trokhymovych, M., Aslam, M., Chou, A.-J., Baeza-Yates, R., & Saez-Trumper, D. (2023). Fair multilingual vandalism detection system for Wikipedia. KDD '23 (arXiv:2306.01650) https://arxiv.org/abs/2306.01650
- Vendor Wikimedia Foundation. Machine learning models: Production language-agnostic revert risk and Production multilingual revert risk (Meta-Wiki model cards stating the deployed model's own bias against new users, temporary accounts and IP edits) https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk
Publishing the flag demonstrably changed the human decision, and the measured direction was toward more equal treatment — a separate finding from the classifier's own bias, at a different layer, and the two must be carried together. Nathan TeBlunthuis, Benjamin Mako Hill, and Aaron Halfaker ran a regression discontinuity across 23 Wikipedia language editions from January 2019 to March 2020, using the RCFilters threshold cutoffs as the discontinuity and revisions within 0.03 of a cutoff. At the 'maybe damaging' cutoff, being flagged raised revert probability from 13.5 to 19.2 percent for unregistered editors and from 4.6 to 14.3 percent for registered ones; at 'likely damaging', from 33.5 to 50.2 percent and from 15.5 to 44.5 percent respectively. Because flagging moved the under-scrutinised group more than the over-scrutinised one, aggregated across thresholds it INCREASED demographic parity between registered and unregistered editors. It also lowered the odds that a revert was itself controversial for unregistered editors, from 3.08 to 2.81 percent at 'likely damaging' and 3.33 to 2.92 percent at 'very likely damaging', which the authors read as evidence that flagging lowered the decision system's false-positive rate. The same authors state plainly in the same paper that 'ORES encodes biases against unregistered editors and - to a lesser extent - against editors without user pages'. The honest reading of the two findings together is a biased classifier whose published score reduced a larger human bias, and neither half stands alone.
empirical- Academic TeBlunthuis, N., Hill, B. M., & Halfaker, A. (2021). Effects of Algorithmic Flagging on Fairness: Quasi-experimental Evidence from Wikipedia. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1, Article 56 (arXiv:2006.03121) https://arxiv.org/abs/2006.03121
- Academic Trokhymovych, M., Aslam, M., Chou, A.-J., Baeza-Yates, R., & Saez-Trumper, D. (2023). Fair multilingual vandalism detection system for Wikipedia. KDD '23 (arXiv:2306.01650) https://arxiv.org/abs/2306.01650
The transparency mechanism the peer-reviewed record identifies as the core of participatory machine learning here was removed in an infrastructure migration as an unused feature, and nobody decided against it. ORES exposed threshold optimisations in a machine-readable format so a wiki's tool developer could ask for the maximum filter rate at a stated recall for their own wiki and their own model version; the published worked example on English Wikipedia was a threshold of 0.32 giving a filter rate of 0.89, a false-positive rate of 0.087, precision of 0.23 and recall of 0.75. Probed live on 28 August 2026, that query returns: 'model_info query parameter is not supported by this endpoint anymore.' Chris Albon announced on the wikitech-l list on 3 August 2023 that the ORES API endpoint would move onto Lift Wing by 30 September 2023, that the Foundation wanted zero traffic on the old endpoint by January 2024, that 'The servers that run ORES are at the end of their planned lifespan and so to save cost we are going to shut them down in early 2024', and that 'The ores-legacy endpoint is not a 100% replacement for ores, we removed some very old and not used features.' The migration remains an open programme: a Phabricator task opened on 5 March 2026 records the deprecation guidance as scattered across three wiki pages and 'difficult to find and to maintain', and the MediaWiki modernization page carries its own warning that it 'contains outdated information that may not accurately reflect the current state of Wikimedia ML systems'. The published-score channel survived the migration; the published-fitness-statistics channel did not survive on this endpoint.
empirical- Vendor Albon, C., Wikimedia Foundation (2023, August 3). ORES To Lift Wing Migration (wikitech-l announcement), with the Wikitech ORES and Lift Wing pages and a live probe of both endpoints on 2026-08-28 showing the threshold-statistics query removed https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/EK65B7QCQHEG37C2ERPIUSP64OX3ZEUJ/
- Academic Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189
The correction channel for a wrongly reverted editor is an ordinary public wiki page rather than a ticket nobody outside can see, and its openness is exactly what makes it unmeasured. The acting agent's documentation tells anyone reverted in error to redo the edit, remove the warning, and report the false positive, and points at a page that also holds the full public list of reported false positives; that page was reachable when checked on 28 August 2026. The operator's own agent makes creating such a page a required step of deployment and links it from the talk-page message, from the page history, and from the user's contributions beside the ordinary Undo and Thank actions, in a default message that reads: 'Because the model I use is not perfect, it sometimes reverts good edits. If you believe the change you made was constructive, please report it here.' The agent's team states an intention to investigate retraining on reported false positives. What does not exist is a disposition: because no operator adjudicates anything, a false-positive report is a wiki page rather than a case, so this deployment publishes no reversal rate on contested reverts and cannot supply the overturn statistic that a mandated platform transparency report supplies. Its correction channel is more open and less measured than a regulated one.
empirical- Vendor English Wikipedia. User:ClueBot NG/Documentation (volunteer maintainers' own published statistics on their own held-out data; the false-positive-budget mechanism, the 0.1 and 0.25 per cent settings, and the dataset-bias concession) https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation
- Vendor English Wikipedia. User:ClueBot NG/FalsePositives and the ClueBot NG dataset review interface on Toolforge (a public list of reported errors and an open interface for labelling training data) https://en.wikipedia.org/wiki/User:ClueBot_NG/FalsePositives
- Vendor Wikimedia Foundation. Moderator Tools/Automoderator and Moderator Tools/Automoderator/Testing (MediaWiki.org; the community-switch, the watchlistable configuration, the required false-positive page, and the caution-level table from the builders' own pre-deployment testing on 22 spreadsheets over about 600 edits) https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator
The scale this deployment exists to address is published, and so is the labour arithmetic behind it. Queried from the Wikimedia Analytics REST API on 28 August 2026, English Wikipedia alone took between 2.59 and 2.94 million non-bot edits to content pages per month across the twelve months to July 2026 — 2,839,865 in July 2026 — on the order of 85,000 to 95,000 human content edits a day on one of more than 250 language editions the models serve, against an operator-stated base rate of fewer than 5 problem edits in 100. The ORES paper's all-editions figure was about 290,000 edits a day, and its arithmetic is that reviewing that at an aggressive ten revisions a minute is about 483 volunteer labour hours daily, which a model filtering ninety percent of the stream reduces to about 48.3 — turning 240 volunteers at two hours a day into 24, and, for a small wiki, turning the task into one or two part-time volunteers. The service that does the filtering ran at 50 to 125 external requests a minute in steady state with bursts to 400 to 500 a second, precaching requests roughly an order of magnitude higher because a scoring job starts for nearly every edit, an approximately 80 percent cache hit rate, and most predictions computed in about a second. The team behind it 'never had more than 3 paid staff and 3 volunteers at any time, and no more than 2 requests typically in progress simultaneously', against roughly 66,000 monthly active English Wikipedia editors at the time.
empirical- Vendor Wikimedia Foundation. Analytics REST API, monthly user edits to content pages for en.wikipedia.org (2.59 to 2.94 million per month in the year to July 2026) https://wikimedia.org/api/rest_v1/metrics/edits/aggregate/en.wikipedia.org/user/content/monthly/2025010100/2026080100
- Academic Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189
- Vendor Wikimedia Foundation. Help:New filters for edit review, Quality and Intent Filters (MediaWiki.org; per-threshold precision and recall stated to patrollers, and the per-wiki divergence) https://www.mediawiki.org/wiki/Help:New_filters_for_edit_review/Quality_and_Intent_Filters
Four outages of the fastest automated tier in the first half of 2011 give this deployment a measured counterfactual, and what it measures is latency rather than coverage. R. Stuart Geiger and Aaron Halfaker, at WikiSym 2013, analysed ClueBot NG's downtime on 15 to 18 February, 13 to 17 March, 29 March to 7 April, and 15 April to 1 May 2011. Comparing only Wednesdays and Thursdays in order to control for the weekly editing rhythm, median time-to-revert rose from 744 seconds with the bot running to 1,286 seconds with it down, and the geometric mean from 941 to 1,674 seconds. The proportion of edits that were reverts fell significantly during downtime (chi-square 115.9, p<0.001), but the proportion of revisions EVENTUALLY reverted did not differ (chi-square 0.64, p=0.43): the human tiers absorbed the work at a slower rate rather than losing it. The authors were careful not to read this as the bot being dispensable, asking instead what the workaround cost the editors who performed it. The same paper documents the reviewer tiers whose different clocks make that absorption possible: fully automated bots reverting within seconds, tool-assisted humans mostly within a minute, manual browser reverts between a minute and a day, and batch scripts on an idiosyncratic scatter. The measurement is excellent evidence for the shape of the effect and weak evidence for its present magnitude: it concerns one bot on one wiki at a scale and tooling mix that no longer obtain.
empirical- Academic Geiger, R. S., & Halfaker, A. (2013). When the Levee Breaks: Without Bots, What Happens to Wikipedia's Quality Control Processes? WikiSym 2013 https://stuartgeiger.com/wikisym13-cluebot.pdf
Auditing this deployment is a tooled activity for the governed population rather than a privilege of the operator, and its limit is a retention decision. ORES-Inspect, described by Zachary Levonian, Lauren Hagen, Lu Li, Jada Lilleboe, Solvejg Wastvedt, Aaron Halfaker, and Loren Terveen at the 2024 Wiki Workshop, is an open-source Toolforge interface that lets any editor sample two disagreement quadrants — 'Unexpected Reverts', where the model called an edit fine and the community reverted it, and 'Unexpected Consensus', where the model called an edit damaging and the community left it standing — across the 35.6 million non-bot English Wikipedia edits of 2019, with the prediction as it was made at the time, and turn a single noticed misclassification into a quantified false-positive or false-negative rate for a chosen slice such as newcomers, LGBT-history pages, or stubs. The design targets exactly the loop the deployment carries: reverts become training labels for the next model, so an over-flagging threshold can teach itself. The same paper records that historical predictions were retained only until the end of 2019, so an audit of what the model said at the moment of a past edit cannot run past that year even though the live scores have always been public.
empirical- Academic Levonian, Z., Hagen, L., Li, L., Lilleboe, J., Wastvedt, S., Halfaker, A., & Terveen, L. (2024). ORES-Inspect: A technology probe for machine learning audits on enwiki. Wiki Workshop 2024 (arXiv:2406.08453) https://arxiv.org/abs/2406.08453
- Academic Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189
The label definition itself is contestable by the governed population here, and the training data is collected in the open. Tzu-Sheng Kuo, Aaron Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein, and Haiyi Zhu reported at CHI 2024 on Wikibench, a system that puts AI evaluation-data curation through Wikipedia's ordinary talk-page and consensus machinery; the authors report that datasets curated this way 'can effectively capture community consensus, disagreement, and uncertainty' and that participants used it to refine label definitions, set data inclusion criteria, and author data statements. The acting agent's own training labels come from a public Toolforge review interface whose maintainers state their aim plainly — 'We need volunteers to help review edits and classify them as either vandalism or constructive. We hope to eventually completely replace our current dataset with a random sampling of edits, reviewed and classified by volunteers' — with each edit in the trial slice behind their published statistics reviewed by at least two humans, and with the same documentation conceding that 'Our current dataset has some degree of bias, as well as some inaccuracies.' The older edit-quality models were trained on per-wiki volunteer labelling campaigns; the successor language-agnostic model was trained on published MediaWiki History and Wikitext History tables for January 2022 to January 2023 excluding bot edits on a 70/30 split, and the multilingual model on 8.6 million revisions from January to July 2022, sampled up to 300,000 per language, with a 17 percent unregistered-edit rate and an 8 percent revert rate in the training data.
empirical- Academic Kuo, T.-S., Halfaker, A., Cheng, Z., Kim, J., Wu, M.-H., Wu, T., Holstein, K., & Zhu, H. (2024). Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia. CHI 2024 (arXiv:2402.14147) https://arxiv.org/abs/2402.14147
- Vendor English Wikipedia. User:ClueBot NG/Documentation (volunteer maintainers' own published statistics on their own held-out data; the false-positive-budget mechanism, the 0.1 and 0.25 per cent settings, and the dataset-bias concession) https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation
- Vendor Wikimedia Foundation. Machine learning models: Production language-agnostic revert risk and Production multilingual revert risk (Meta-Wiki model cards stating the deployed model's own bias against new users, temporary accounts and IP edits) https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk
- Academic Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189
The costs this deployment was designed against were measured on Wikipedia itself before the scoring service existed, and are never a measured effect of it. Aaron Halfaker, R. Stuart Geiger, Jonathan T. Morgan and John Riedl reported in 2013 that the share of good-faith newcomers whose first-session edit was reverted rose from 6.1 percent in the first half of 2006 to 18.2 percent in the first half of 2007, that two-month survival of those newcomers fell from 25.6 percent to 11.7 percent within a year and did not recover, and that tool-mediated rejection of them rose from about 0 percent in 2006 to about 40 percent in 2010, with both rejection and tool-mediated rejection significant negative predictors of newcomer survival. The harm sat in the interaction as much as in the classification: reciprocation of a reverted newcomer's attempt to open a discussion averaged 7 percent for editors using Huggle, about 30 percent for Rollback, 53 percent for Twinkle, and 56 to 67 percent for manual reverters, and 2,250 discussion attempts from 918 registered editors were addressed to an algorithmic editor that could not reply. The 2020 ORES paper concedes that after this research 'the often-hostile quality control processes that were designed over a decade ago remain largely unchanged'. Against that, the same score is also used prosocially: the 'Very likely good' filter, about 99 percent precise at over 90 percent recall, is documented as a way to find good-faith newcomers to thank, and a Wiki Education tool asks the article-quality model to re-score a student's draft with one more citation, header, or image in order to recommend the most productive next edit.
empirical- Academic Halfaker, A., Geiger, R. S., Morgan, J. T., & Riedl, J. (2013). The Rise and Decline of an Open Collaboration System: How Wikipedia's Reaction to Popularity Is Causing Its Decline. American Behavioral Scientist, 57(5), 664-688 https://stuartgeiger.com/papers/abs-rise-and-decline-wikipedia.pdf
- Academic Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189
- Vendor Wikimedia Foundation. Help:New filters for edit review, Quality and Intent Filters (MediaWiki.org; per-threshold precision and recall stated to patrollers, and the per-wiki divergence) https://www.mediawiki.org/wiki/Help:New_filters_for_edit_review/Quality_and_Intent_Filters
There is no litigation, no regulator, no court, no consent order, and no statutory transparency mandate anywhere in this deployment's record, and that absence is a structural finding rather than an absence of controversy. Litigation posture: None. Every accountability artefact here — the open scoring interface, the published operating points, the model cards stating their own biases, the public false-positive pages, the community audit tooling — exists because the operator and the self-governing volunteer communities chose it, and could be withdrawn the same way; the removal of the machine-readable threshold-statistics query in the 2023 to 2024 infrastructure migration is a small, dated instance of exactly that. Two further honesty notes belong with any description of the strength of this record. First, this is not a solved or harm-free deployment: its own operators concede the hostile quality-control processes documented in 2013 remain largely unchanged, the deployed model's card warns it may be biased against new users, temporary accounts, and unregistered editors, and the acting agent's maintainers concede their dataset carries bias and inaccuracies. Second, the peer-reviewed evidence base leans on a small overlapping author group — Halfaker and Geiger appear on the ORES paper, the outage paper, and the newcomer-decline paper, and Halfaker also co-authors the flagging-fairness paper and the audit probe, having been a Wikimedia Foundation employee for most of the period covered — with the KDD evaluation and the audit probe the clearest checks outside that lineage, and the KDD paper the one that measures the deployed model unfavourably.
empirical- Academic Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189
- Vendor Wikimedia Foundation. Machine learning models: Production language-agnostic revert risk and Production multilingual revert risk (Meta-Wiki model cards stating the deployed model's own bias against new users, temporary accounts and IP edits) https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk
- Vendor English Wikipedia. User:ClueBot NG/Documentation (volunteer maintainers' own published statistics on their own held-out data; the false-positive-budget mechanism, the 0.1 and 0.25 per cent settings, and the dataset-bias concession) https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation
- Academic Trokhymovych, M., Aslam, M., Chou, A.-J., Baeza-Yates, R., & Saez-Trumper, D. (2023). Fair multilingual vandalism detection system for Wikipedia. KDD '23 (arXiv:2306.01650) https://arxiv.org/abs/2306.01650
- Academic Levonian, Z., Hagen, L., Li, L., Lilleboe, J., Wastvedt, S., Halfaker, A., & Terveen, L. (2024). ORES-Inspect: A technology probe for machine learning audits on enwiki. Wiki Workshop 2024 (arXiv:2406.08453) https://arxiv.org/abs/2406.08453
- Vendor Albon, C., Wikimedia Foundation (2023, August 3). ORES To Lift Wing Migration (wikitech-l announcement), with the Wikitech ORES and Lift Wing pages and a live probe of both endpoints on 2026-08-28 showing the threshold-statistics query removed https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/EK65B7QCQHEG37C2ERPIUSP64OX3ZEUJ/
Where this connects
Institutional pressures in this domain
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
All of them in context on the Content moderation & editorial AI domain page.
Levers available here and the patterns behind them
- Review the riskiest first — Risk-tiered oversight
- Gate record entries — Human-in-the-loop write gating
- Pause AI on alarms — Deployment circuit-breaker
- Understand the system — Understand the system
- Check copied records — Reconcile copied records
- Check with a second model — Cross-model verification
- Upgrade model — Improve the model
- Verify output — Put a verifier on the agent
- Escalate checks — State-feedback vigilance
- Review on schedule — Oversight cadence & retrospectives
- Mark AI-written records — Provenance labeling
- Peer sharing rules — Peer-edge governance
Documented case histories
- The score is published and the service cannot act on it
- The errors that became visible when the reviewers went home
- The most built-out correction structure and the reach it doesn't have
- The byline nobody was behind
- A staff byline the AI wrote and the review it implied
- StopNCII & Take It Down
- X Multilingual Hate-Speech Enforcement
- X Community Notes (crowd annotation)
- GIFCT hash-sharing database
- Google CSAM detection and total account closure
- Meta cross-check: the enforcement-exemption tier
- The CyberTipline: triage under a rule against looking
- Sama Nairobi: the review workforce as the governed subsystem
- TikTok EU and UK trust-and-safety staffing substitution
- YouTube Content ID