Skip to content

Domain Atlas / Content moderation & editorial AI

Case fileUnited States for the infrastructure and the non-profit operator — the Wikimedia Foundation, San Francisco, owns and runs the servers — and NOT for the governance. The models serve more than 250 Wikipedia language editions and the better-calibrated multilingual model covers 47 of them; the operator's own reverting agent runs on twelve Wikipedias and not on English; and the crowd audit that best demonstrates the governance mechanism happened on Spanish Wikipedia. Bot approval, bot blocking, the granting and revocation of the rollback right, threshold selection and tool adoption are all exercised by the self-governing volunteer community of each language edition. Litigation posture: None. There is no court, no regulator, no consent order and no statutory transparency mandate anywhere in this record, and that is a substantive structural finding about this governance arrangement rather than an absence of controversy.medium deployment

The score is published and the service cannot act on it

Explore this deployment in the PAN Lab ↗

In the PAN Lab, the readouts of this case's model organization carry a shaded evidence band whose width follows the least-established class among the modeling inputs the readings rest on.

The least-established input behind this case's model organization's readings is an assumption, not a measurement. Evidence base: 2 assumed · 7 published baseline · 3 measured.

Wikipedia's edit-scoring service publishes a damage or revert-risk probability for essentially every edit as it is saved and has no path of its own by which it can act on one. Aaron Halfaker and R. Stuart Geiger, writing the system up for CSCW in 2020, describe it as built to decouple four activities normally performed by the same engineers: curating training data, building models, auditing predictions, and building the interfaces or bots that act on predictions. The operator did not gate access either, on its own stated reasoning that 'Given the open API, there is no barrier where we can selectively decide who can request a score from a classifier.' Acting on the score is done by separately governed agents: a volunteer-run bot on English Wikipedia, an operator-built agent local administrators switch on, tool-assisted patrollers working score-ranked queues, and ordinary editors with score-driven filters enabled in their own preferences. At the time of that paper the estate ran roughly 110 classifiers across 44 languages in four families; the ORES infrastructure has since been retired and what runs today is Lift Wing, with ores-legacy.wikimedia.org as a compatibility endpoint in front of it, serving a revert-risk family that replaced the older edit-quality models. Verified live on 28 August 2026: the legacy host answers, self-identifies as the 'ORES legacy service' and still returns scores, and the Lift Wing endpoint returns a revert-risk score for an arbitrary revision without any credential.[2]

What happened

A scoring service computes a damage or revert-risk probability for essentially every edit to every supported Wikipedia as the edit is saved, caches it, and publishes it over an interface that asks for no credential. Nothing in that service reverts, blocks, warns or reports anything. The design intent was explicit: Aaron Halfaker and R. Stuart Geiger, writing up the system for CSCW in 2020, describe it as built to decouple four activities normally done by the same engineers — curating training data, building models, auditing predictions, and building the interfaces or bots that act on predictions. The operator did not gate access either, and said why: "Given the open API, there is no barrier where we can selectively decide who can request a score from a classifier."

Acting on the score is a separate job under separate governance, and there are four kinds of actor doing it. A fully automated bot, ClueBot NG, has edited English Wikipedia since October 2010 and had made 6,682,890 edits as of 28 August 2026, holding the bot, reviewer, rollbacker and autoconfirmed user groups — all of them community-granted and community-revocable. Automoderator, built by the Wikimedia Foundation's Moderator Tools team, scores every main-namespace edit and reverts above a threshold; it is deployed on twelve Wikipedias, and not on English. Tool-assisted patrollers work score-ranked queues in Huggle and similar tools. And ordinary editors watch recent changes with the score-driven filters switched on in their own preferences.

The threshold is the actual policy in a damage-detection system, and here it is chosen by the community that will live with it. On English Wikipedia the operating points are published to patrollers in ordinary help text, with both of their costs stated: the strictest quality filter is "right more than 90% of the time" but "finds only about 10% of all the problem edits", while the broadest is "right only about 15% of the time" and catches "about 82% of problem edits", against a stated base rate of "fewer than 5 in 100". The same architecture yields different arrangements on different wikis, because the model and the threshold are both local: on Polish Wikipedia the "Likely have problems" filter captures 91 percent of problem edits against 34 percent for the corresponding English one, and Polish Wikipedia therefore "does not need - or have" the broad, noisy filter English Wikipedia relies on.

The volunteer bot's control variable is not a score at all but an error budget. Its documentation states that "The threshold is not randomly chosen by a human, but is instead calculated to match a given false positive rate... A human selects a false positive rate, which is the percentage of constructive edits incorrectly classified as vandalism." At the current setting of 0.1 percent the bot catches approximately 40 percent of vandalism; at the previous 0.25 percent setting it caught approximately 55. Fifteen points of catch rate were spent, deliberately and in public, to halve the wrongful-reversion rate. Both halves of that trade are published, which is rare enough in this domain to be worth saying plainly. These are the maintainers' own figures, computed on their own held-out, human-reviewed slice: methodologically described, and not independently audited.

The Foundation's own reverting agent was built so that it would not be an operator enforcement node. "After deployment, Automoderator will not begin running until a local administrator turns it on." Its configuration lives at a Community Configuration page that writes an ordinary, watchlistable wiki page, so a change to the operating point appears in the same feed as an edit to an article. Creating a false-positive reporting page is a required step of deployment. It never reverts administrators, global sysops, stewards or bots, never reverts self-reverts or reverts of its own actions, and never touches new page creations. And the choice of threshold is put to a community as a table of consequences rather than a slider with no units: scores above 0.99 at about 100 percent precision and roughly 152 English Wikipedia reverts a day, 0.985 at about 95 percent and 350, 0.98 at about 93 percent and 680, 0.975 at about 82 percent and 1,077. The builders withdrew the least cautious option themselves for falling below their own 90 percent accuracy target. That table is their pre-deployment testing, aggregated from 22 completed review spreadsheets covering over 600 edits across six projects, not measured production performance.

When a volunteer did wire a score straight to automatic reversion, the community stopped it and the model's operators were not involved. PatruBOT auto-reverted Spanish Wikipedia edits above a threshold its developer had set too low; Spanish Wikipedians crowd-sourced an audit of its errors on ordinary wiki pages, reached a consensus that it was making too many mistakes, and an administrator blocked the bot's account. The ORES team wrote that this "was entirely a community governed activity that required no intervention of our team or the Wikimedia Foundation staff", and recorded that a successor, SeroBOT, later resumed auto-reverting at a higher confidence threshold.

Two measurements of fairness exist here, they point in opposite directions, and the record only makes sense if both are held at once. Nathan TeBlunthuis, Benjamin Mako Hill and Aaron Halfaker ran a regression discontinuity across 23 language editions between January 2019 and March 2020, using the filter cutoffs as the discontinuity. Being flagged raised revert probability from 13.5 to 19.2 percent for unregistered editors and from 4.6 to 14.3 percent for registered ones at the "maybe damaging" cutoff, and from 33.5 to 50.2 and 15.5 to 44.5 percent respectively at "likely damaging" — so the flag moved the under-scrutinised group more than the over-scrutinised one and narrowed a pre-existing gap. It also lowered the odds that a revert of an unregistered editor was itself contested, from 3.08 to 2.81 percent at one cutoff and 3.33 to 2.92 at another. Those authors state in the same paper that "ORES encodes biases against unregistered editors and - to a lesser extent - against editors without user pages". Mykola Trokhymovych and colleagues, at the KDD conference in 2023, measured that classifier bias directly: a Disparate Impact Ratio of 20.02 for the deployed model against a base-rate ratio of 7.93 in the same data, with the successor multilingual model at 9.54 with the same user features and 1.98 to 3.08 without them. The deployed model's area under the curve was 0.84 with precision at recall 0.75 of 0.22; a deliberately unfair baseline that simply reverts every unregistered edit scores 0.75 and 0.07. A biased classifier whose published score reduced a larger human bias is the honest summary, and it is why where the threshold sits, and who sets it, is the thing that matters here.

The costs this arrangement is designed against were measured on Wikipedia itself before any of it existed. Aaron Halfaker, R. Stuart Geiger, Jonathan T. Morgan and John Riedl reported in 2013 that the share of good-faith newcomers whose first-session edit was reverted rose from 6.1 percent in the first half of 2006 to 18.2 percent a year later, that two-month survival of those newcomers fell from 25.6 to 11.7 percent and did not recover, and that tool-mediated rejection of them rose from roughly zero in 2006 to about 40 percent in 2010. The harm was in the interaction as much as the classification: reverted newcomers who tried to open a discussion were answered 7 percent of the time by editors using Huggle, about 30 percent for Rollback, 53 percent for Twinkle and 56 to 67 percent for manual reverters, and 2,250 discussion attempts were addressed to an algorithmic editor that could not reply. This is the harm the scoring service was built in response to, not a measured effect of it — and the ORES paper concedes that after that research "the often-hostile quality control processes that were designed over a decade ago remain largely unchanged".

Auditing here is a tooled activity for the governed population rather than a privilege of the operator. ORES-Inspect, described by Zachary Levonian and colleagues in 2024, lets any editor sample the two disagreement quadrants — edits the model called fine that the community reverted, and edits it called damaging that the community left standing — across 35.6 million non-bot English Wikipedia edits from 2019, with the prediction as it was made at the time, and turn one noticed error into a quantified rate for a slice of their choosing. Wikibench, a CHI 2024 field study by Tzu-Sheng Kuo and colleagues, put evaluation-data curation through Wikipedia's ordinary talk-page and consensus machinery and reported that participants refined label definitions, set inclusion criteria and authored data statements. The label definition itself is contestable, and the bot's maintainers concede their own dataset "has some degree of bias, as well as some inaccuracies" while running a public interface asking volunteers to help replace it.

The operator publishes model cards as a standing practice, written to the Mitchell et al. framework, on wiki pages that carry talk pages. The card for the language-agnostic revert-risk model — the model Automoderator actually uses — says it "may exhibit bias against edits from new users, temporary accounts, or IP edits" and recommends the multilingual model instead for anonymous edits in the languages that model covers. It also forbids using the model as ground truth for training other models, and forbids scoring a page's first revision. Coverage is the honest limit of the whole transparency story: the better-calibrated model — the one whose scores match the rate at which flagged edits really do turn out to be reverted — covers 47 languages against more than 250 editions, and across both systems model accuracy correlates negatively with a language edition's share of anonymous editors, so the communities that lean most on anonymous contribution get the least accurate scoring.

There is one measured counterfactual for what an automated damage-detection channel actually contributes, and it is worth the space. ClueBot NG went down four times in the first half of 2011 — 15 to 18 February, 13 to 17 March, 29 March to 7 April, and 15 April to 1 May. Comparing only Wednesdays and Thursdays to control for the weekly editing rhythm, Geiger and Halfaker found median time-to-revert rose from 744 seconds to 1,286 and the geometric mean from 941 to 1,674 while the bot was down. Significantly fewer reverts happened in those windows, chi-square 115.9, p<0.001 — but the proportion of revisions eventually reverted was statistically indistinguishable, chi-square 0.64, p=0.43. The human tiers absorbed the work more slowly. The authors declined to read that as the bot being dispensable, asking instead what the workaround cost the editors who performed it.

ORES no longer exists as infrastructure and has not since early 2024. Chris Albon announced on the wikitech-l list on 3 August 2023 that the API endpoint would move onto Lift Wing by 30 September 2023, that the Foundation wanted zero traffic on the old endpoint by January 2024, and that "The servers that run ORES are at the end of their planned lifespan and so to save cost we are going to shut them down in early 2024". What runs today is Lift Wing, a model-serving platform on Kubernetes, with ores-legacy.wikimedia.org as a compatibility shim in front of it. Verified live on 28 August 2026: the host answers, self-identifies as the "ORES legacy service", and still returns scores from the old model names; the Lift Wing endpoint returns revert-risk scores for an arbitrary revision without any credential. The published-score channel survived the migration intact.

One channel did not, and it is the one the peer-reviewed record identifies as the core of participatory machine learning here. ORES exposed threshold optimisations in a machine-readable format, so a wiki's tool developer could ask, in effect, for the maximum filter rate at recall of at least 0.75 for their wiki and their model version; the published worked example on English Wikipedia was a threshold of 0.32 giving a filter rate of 0.89, a false-positive rate of 0.087, precision of 0.23 and recall of 0.75. That query now returns: "model_info query parameter is not supported by this endpoint anymore." The migration announcement had warned that "The ores-legacy endpoint is not a 100% replacement for ores, we removed some very old and not used features." Nobody voted to end participatory threshold-setting. The mechanism that implemented it was removed as an unused feature when the servers reached end of life, and a Phabricator task opened on 5 March 2026 records that the deprecation documentation is still scattered across three pages and "difficult to find and to maintain". The historical audit trail is thinner than the live one for the same kind of reason: ORES-Inspect notes that historical predictions were retained only until the end of 2019, so an audit of what the model said at the moment of a past edit stops there.

The scale problem all of this exists to address has not gone away. English Wikipedia alone took between 2.59 and 2.94 million non-bot edits to content pages per month over the twelve months to July 2026, on the order of 85,000 to 95,000 a day on one language edition. The operator's own arithmetic on the all-editions figure is that reviewing roughly 290,000 edits a day at an aggressive ten a minute would be about 483 volunteer labour hours daily, which a model that filters ninety percent of the stream reduces to about 48. Against that stands a team that, on its own account, "never had more than 3 paid staff and 3 volunteers at any time, and no more than 2 requests typically in progress simultaneously".

Litigation posture: None. There is no court, no regulator, no consent order and no statutory transparency mandate anywhere in this record. That is the structural fact that makes this case worth reading beside every other deployment in the domain — and it cuts both ways, because every accountability artefact here exists because the operator and the communities chose it, and could be withdrawn the same way. The retirement of the threshold-statistics query is a small, dated, documented instance of exactly that.

The sociotechnical reading

Most cases in this atlas sit inside one organisation: a model, the people who act on it, the records it writes, and someone whose job is to check. This one has the same four parts and puts each of them under a different authority, and almost everything interesting follows from where those lines fall.

Start with what the automated element cannot do. The scoring service has no enforcement path. It computes a probability and writes it to a public record, and that is the end of its powers. The thing that acts on the score reads it out of the same public record an outside researcher reads. There is no private wire between the classifier and the actor, which means the interface between "what the machine thinks" and "what happens to an edit" is an artefact anybody can inspect, argue with, and copy. Read against the closed pipelines elsewhere in this domain — where the classifier, the threshold, the acting agent and the appeal channel all sit inside one operator and the public record is whatever that operator chose to publish — this is the same task with the joints exposed.

The second structural fact is that the operating point is chosen by the people who will live with it, three times over: once for the wiki, once in a tool configuration page anybody may edit and watch, and once by each patroller who sets a personal minimum. That is a governance arrangement rather than a feature, and it is what makes the published precision and recall load-bearing rather than decorative. Numbers that describe a choice you cannot make are trivia; numbers that describe a choice you must make are the choice. The bot maintainer's error budget is the sharpest version: pick a maximum false-positive rate and let the threshold be computed from it, so the question a human answers is "how much wrongful reversion is acceptable" rather than "what number feels right".

Third, the checking here runs sideways rather than downward. The people who are governed are the people who hold the levers: a bot may not edit until a local approvals group has cleared it and it has run a trial; any administrator may block one that misbehaves; the permission the fastest tools require is granted and revoked by the same community; the operator's own agent will not start until a local administrator turns it on. Four of the five governance hops in the underlying PAN model are checks between classes rather than influence between them. That is not a milder version of a regulator; it is a different shape entirely, and its failure modes are different too. It has no floor. Nothing compels a volunteer maintainer to act on a published finding, and nothing preserves an instrument that stops being convenient to run.

Fourth, the record loops back on itself in a way this atlas sees everywhere and can rarely watch. Reverts become the labels the next model trains on, so an operating point that surfaces one kind of edit teaches the next model to surface it again. Here the check on that loop exists and is public — sample the two quadrants where the model and the community disagreed, and read what the model said at the time against what people then did — and its limit is a retention decision rather than a design one: predictions were kept only through the end of 2019.

The two fairness measurements are the hardest thing in the file to hold steady, and the Lab network draws both. TeBlunthuis, Hill and Halfaker measured the effect of the FLAG on the human decision and found it raised revert probability far more for registered editors than for unregistered ones, narrowing a pre-existing profiling gap; Trokhymovych and colleagues measured the CLASSIFIER's own disparate impact at 20.02 against a base rate of 7.93. Neither refutes the other. They are measurements of different layers, and the useful lesson is that a biased instrument placed inside a well-arranged decision process can improve the process's fairness while remaining biased — which is an argument for arranging the process, not for excusing the instrument.

What this deployment cannot supply is the statistic every regulated platform in this domain does supply. Because no operator adjudicates anything, a false-positive report is a wiki page rather than a case with a disposition, and there is no published reversal rate on contested reverts. Its correction channel is more open and less measured than a mandated one: any reverted editor is told a machine did it, in a message that concedes the model sometimes undoes good edits, and can add their case to a public list that anybody can read — and nobody publishes what share of that list was upheld.

Finally, the pressure this network carries is not a scandal but an attrition. One infrastructure decision, taken to save cost at the end of a server's planned life, narrowed two things at once: the sensor, because the machine-readable fitness-statistics query that made threshold-setting participatory was dropped as an old, unused feature; and the model estate, because roughly 110 classifiers in four families across 44 languages gave way to a revert-risk family whose language-agnostic member runs anywhere and whose better-calibrated multilingual member covers 47 of more than 250 editions. Nobody decided against participatory machine learning. The thing that implemented it was removed while nobody was looking at that particular line item, and there was no regulator, court or reporting duty anywhere in the arrangement whose job it would have been to notice.

Read this board against the enforcement pipelines beside it and the question it answers is a specific one: which harms in those networks come from the classifier, and which come from the governance wrapped around it. This deployment has a measurably biased classifier and it publishes the measurement. What it does not have is a closed pipeline.

The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.

Grounding sources for this case

The same sources that ground this model organization in the PAN library: evaluations, government documents, investigative reporting, and advocacy documentation, each labeled by tier.

halfaker2020GroundingAcademicSave

Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189

https://arxiv.org/abs/1909.05189

Grounds: model org: wikipedia_ores

Topics: co-design

geiger2013GroundingAcademicSave

Geiger, R. S., & Halfaker, A. (2013). When the Levee Breaks: Without Bots, What Happens to Wikipedia's Quality Control Processes? WikiSym 2013 https://stuartgeiger.com/wikisym13-cluebot.pdf

https://stuartgeiger.com/wikisym13-cluebot.pdf

Grounds: model org: wikipedia_ores

teblunthuis2021GroundingAcademicSave

TeBlunthuis, N., Hill, B. M., & Halfaker, A. (2021). Effects of Algorithmic Flagging on Fairness: Quasi-experimental Evidence from Wikipedia. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1, Article 56 (arXiv:2006.03121) https://arxiv.org/abs/2006.03121

https://arxiv.org/abs/2006.03121

Grounds: model org: wikipedia_ores

Topics: algorithmic-fairness

halfaker2013GroundingAcademicSave

Halfaker, A., Geiger, R. S., Morgan, J. T., & Riedl, J. (2013). The Rise and Decline of an Open Collaboration System: How Wikipedia's Reaction to Popularity Is Causing Its Decline. American Behavioral Scientist, 57(5), 664-688 https://stuartgeiger.com/papers/abs-rise-and-decline-wikipedia.pdf

https://stuartgeiger.com/papers/abs-rise-and-decline-wikipedia.pdf

Grounds: model org: wikipedia_ores

trokhymovych2023GroundingAcademicSave

Trokhymovych, M., Aslam, M., Chou, A.-J., Baeza-Yates, R., & Saez-Trumper, D. (2023). Fair multilingual vandalism detection system for Wikipedia. KDD '23 (arXiv:2306.01650) https://arxiv.org/abs/2306.01650

https://arxiv.org/abs/2306.01650

Grounds: model org: wikipedia_ores

levonian2024GroundingAcademicSave

Levonian, Z., Hagen, L., Li, L., Lilleboe, J., Wastvedt, S., Halfaker, A., & Terveen, L. (2024). ORES-Inspect: A technology probe for machine learning audits on enwiki. Wiki Workshop 2024 (arXiv:2406.08453) https://arxiv.org/abs/2406.08453

https://arxiv.org/abs/2406.08453

Grounds: model org: wikipedia_ores

kuo2024GroundingAcademicSave

Kuo, T.-S., Halfaker, A., Cheng, Z., Kim, J., Wu, M.-H., Wu, T., Holstein, K., & Zhu, H. (2024). Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia. CHI 2024 (arXiv:2402.14147) https://arxiv.org/abs/2402.14147

https://arxiv.org/abs/2402.14147

Grounds: model org: wikipedia_ores

albon2023GroundingVendorSave

Albon, C., Wikimedia Foundation (2023, August 3). ORES To Lift Wing Migration (wikitech-l announcement), with the Wikitech ORES and Lift Wing pages and a live probe of both endpoints on 2026-08-28 showing the threshold-statistics query removed https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/EK65B7QCQHEG37C2ERPIUSP64OX3ZEUJ/

https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/EK65B7QCQHEG37C2ERPIUSP64OX3ZEUJ/

Grounds: model org: wikipedia_ores

englishwikipediaGroundingVendorSave

English Wikipedia. User:ClueBot NG/Documentation (volunteer maintainers' own published statistics on their own held-out data; the false-positive-budget mechanism, the 0.1 and 0.25 per cent settings, and the dataset-bias concession) https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation

https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation

Grounds: model org: wikipedia_ores

englishwikipediaaGroundingVendorSave

English Wikipedia. User:ClueBot NG/FalsePositives and the ClueBot NG dataset review interface on Toolforge (a public list of reported errors and an open interface for labelling training data) https://en.wikipedia.org/wiki/User:ClueBot_NG/FalsePositives

https://en.wikipedia.org/wiki/User:ClueBot_NG/FalsePositives

Grounds: model org: wikipedia_ores

englishwikipediabGroundingVendorSave

English Wikipedia. Wikipedia:Bot policy and Wikipedia:Huggle; and Wikimedia Meta-Wiki, SWViewer (the community-granted and community-revocable permission to act quickly on a flagged edit) https://en.wikipedia.org/wiki/Wikipedia:Bot_policy

https://en.wikipedia.org/wiki/Wikipedia:Bot_policy

Grounds: model org: wikipedia_ores

englishwikipediacGroundingVendorSave

English Wikipedia. Wikipedia:Huggle/Config (the community-editable configuration page wiring the score into the patrolling queue, with the scoring flag and the amplifier weight) https://en.wikipedia.org/wiki/Wikipedia:Huggle/Config

https://en.wikipedia.org/wiki/Wikipedia:Huggle/Config

Grounds: model org: wikipedia_ores

wikimediafoundationaGroundingVendorSave

Wikimedia Foundation. Moderator Tools/Automoderator and Moderator Tools/Automoderator/Testing (MediaWiki.org; the community-switch, the watchlistable configuration, the required false-positive page, and the caution-level table from the builders' own pre-deployment testing on 22 spreadsheets over about 600 edits) https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator

https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator

Grounds: model org: wikipedia_ores

wikimediafoundationbGroundingVendorSave

Wikimedia Foundation. Machine learning models: Production language-agnostic revert risk and Production multilingual revert risk (Meta-Wiki model cards stating the deployed model's own bias against new users, temporary accounts and IP edits) https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

Grounds: model org: wikipedia_ores

Seeing your organization in this case file?

The histories here are documented after the harm. Mapping a live deployment's pathways and pressures, before the incident report, is engagement work: intake, diagnosis, prescription, and monitoring, with every limitation stated.

Sources & Evidence

Claims made on this page and what supports them. The full registry lives in Evidence.

EmpiricalWikipedia's edit-scoring service publishes a damage or revert-risk probability for essentially every edit as i…

Wikipedia's edit-scoring service publishes a damage or revert-risk probability for essentially every edit as it is saved and has no path of its own by which it can act on one. Aaron Halfaker and R. Stuart Geiger, writing the system up for CSCW in 2020, describe it as built to decouple four activities normally performed by the same engineers: curating training data, building models, auditing predictions, and building the interfaces or bots that act on predictions. The operator did not gate access either, on its own stated reasoning that 'Given the open API, there is no barrier where we can selectively decide who can request a score from a classifier.' Acting on the score is done by separately governed agents: a volunteer-run bot on English Wikipedia, an operator-built agent local administrators switch on, tool-assisted patrollers working score-ranked queues, and ordinary editors with score-driven filters enabled in their own preferences. At the time of that paper the estate ran roughly 110 classifiers across 44 languages in four families; the ORES infrastructure has since been retired and what runs today is Lift Wing, with ores-legacy.wikimedia.org as a compatibility endpoint in front of it, serving a revert-risk family that replaced the older edit-quality models. Verified live on 28 August 2026: the legacy host answers, self-identifies as the 'ORES legacy service' and still returns scores, and the Lift Wing endpoint returns a revert-risk score for an arbitrary revision without any credential.

halfaker2020GroundingAcademicSave

Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189

https://arxiv.org/abs/1909.05189

Grounds: model org: wikipedia_ores

Topics: co-design

albon2023GroundingVendorSave

Albon, C., Wikimedia Foundation (2023, August 3). ORES To Lift Wing Migration (wikitech-l announcement), with the Wikitech ORES and Lift Wing pages and a live probe of both endpoints on 2026-08-28 showing the threshold-statistics query removed https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/EK65B7QCQHEG37C2ERPIUSP64OX3ZEUJ/

https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/EK65B7QCQHEG37C2ERPIUSP64OX3ZEUJ/

Grounds: model org: wikipedia_ores

EmpiricalThe scores, the models, the training-data provenance and the deliberation about all three are public here, and…

The scores, the models, the training-data provenance and the deliberation about all three are public here, and that openness is what makes every outside measurement in this case possible. Scores are queryable for any revision by anyone without a credential, and were from 2015; the models and their code ship under the Apache 2.0 licence and the training pipelines were built as reproducible Makefiles so a third party can rebuild an equivalent model. The operator publishes model cards as a standing institutional practice, written to the Mitchell et al. framework, on a public index of proposed, production and deprecated models; each card is asked to state why the model was made, its proper and improper uses, and its evaluation scores especially for marginalised groups, and each card page carries a talk page, so the documentation is also the venue where it is argued about. Every operating point's precision and recall is published in patroller-facing help text before anyone chooses among them. Against that openness sit two documented limits: historical predictions were retained only until the end of 2019, and because no operator adjudicates anything, a false-positive report is a public wiki page rather than a case with a disposition, so no reversal rate on contested reverts is published the way an overturn rate appears in a mandated transparency report.

halfaker2020GroundingAcademicSave

Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189

https://arxiv.org/abs/1909.05189

Grounds: model org: wikipedia_ores

Topics: co-design

wikimediafoundationbGroundingVendorSave

Wikimedia Foundation. Machine learning models: Production language-agnostic revert risk and Production multilingual revert risk (Meta-Wiki model cards stating the deployed model's own bias against new users, temporary accounts and IP edits) https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

Grounds: model org: wikipedia_ores

levonian2024GroundingAcademicSave

Levonian, Z., Hagen, L., Li, L., Lilleboe, J., Wastvedt, S., Halfaker, A., & Terveen, L. (2024). ORES-Inspect: A technology probe for machine learning audits on enwiki. Wiki Workshop 2024 (arXiv:2406.08453) https://arxiv.org/abs/2406.08453

https://arxiv.org/abs/2406.08453

Grounds: model org: wikipedia_ores

EmpiricalEvery operational lever in this deployment is held by the governed community rather than by the model's operat…

Every operational lever in this deployment is held by the governed community rather than by the model's operator, and the record shows each of them being exercised. A bot may not edit at all before English Wikipedia's Bot Approvals Group has approved it and it has run a trial, and any administrator may block one that malfunctions. The rollback right that the fastest patrolling tool requires is granted and revoked by administrators, and that tool's own documentation states it 'is not intended for new Wikipedia users' and that misuse 'may result in revocation of rollback permissions or being blocked from editing'; the cross-wiki patrolling tool gates its global queue on global rollback, steward or global sysop rights, or at least 1,000 global edits and no active block, and its documentation records no scoring integration. Automoderator, the operator's own reverting agent, 'will not begin running until a local administrator turns it on', its caution level is written to MediaWiki:AutoModeratorConfig.json through Special:CommunityConfiguration — an ordinary watchlistable wiki page — and creating a false-positive reporting page is a required step of deployment. It is live on twelve Wikipedias (Indonesian, Turkish, Ukrainian, Vietnamese, Afrikaans, Bengali, Azerbaijani, Chinese, Spanish, Italian, Dutch and Albanian, first Turkish in June 2024) and not on English, whose project page records the team's own position that 'If English Wikipedia editors don't want to use Automoderator, that's fine!' When a volunteer wired a score directly to automatic reversion on Spanish Wikipedia, the community stopped it without the operator: PatruBOT was crowd-audited on ordinary wiki pages, judged to be erring too often, and had its account blocked by an administrator — an episode the ORES team recorded as 'entirely a community governed activity that required no intervention of our team or the Wikimedia Foundation staff', with a successor, SeroBOT, later resuming at a higher confidence threshold.

halfaker2020GroundingAcademicSave

Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189

https://arxiv.org/abs/1909.05189

Grounds: model org: wikipedia_ores

Topics: co-design

englishwikipediabGroundingVendorSave

English Wikipedia. Wikipedia:Bot policy and Wikipedia:Huggle; and Wikimedia Meta-Wiki, SWViewer (the community-granted and community-revocable permission to act quickly on a flagged edit) https://en.wikipedia.org/wiki/Wikipedia:Bot_policy

https://en.wikipedia.org/wiki/Wikipedia:Bot_policy

Grounds: model org: wikipedia_ores

wikimediafoundationaGroundingVendorSave

Wikimedia Foundation. Moderator Tools/Automoderator and Moderator Tools/Automoderator/Testing (MediaWiki.org; the community-switch, the watchlistable configuration, the required false-positive page, and the caution-level table from the builders' own pre-deployment testing on 22 spreadsheets over about 600 edits) https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator

https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator

Grounds: model org: wikipedia_ores

EmpiricalThe acting agent's control variable is a human-chosen error budget rather than a score, and both halves of the…

The acting agent's control variable is a human-chosen error budget rather than a score, and both halves of the trade it buys are published. ClueBot NG's documentation states that 'The threshold is not randomly chosen by a human, but is instead calculated to match a given false positive rate... A human selects a false positive rate, which is the percentage of constructive edits incorrectly classified as vandalism.' At the current 0.1 percent setting the bot catches approximately 40 percent of vandalism; at the previous 0.25 percent setting it caught approximately 55 percent, so roughly fifteen points of catch rate were given up to halve the wrongful-reversion rate. Chosen instead for total accuracy, it classifies over 90 percent of edits correctly. Post-processing filters on self-reverts, edit count and warning share cut actual false positives below the stated rate before any revert happens. These figures are the volunteer maintainers' own, computed on a held-out, human-reviewed slice of their own dataset: methodologically described and not independently audited. The agent's scale and standing are separately verifiable: registered 20 October 2010, 6,682,890 edits as of 28 August 2026, holding the bot, reviewer, rollbacker and autoconfirmed user groups, all community-granted, and stoppable by any administrator editing a run page to 'False'. The operator's own agent applies its own published carve-outs, never reverting administrators, global sysops, stewards or bots, self-reverts, reverts of its own actions, or new page creations.

englishwikipediaGroundingVendorSave

English Wikipedia. User:ClueBot NG/Documentation (volunteer maintainers' own published statistics on their own held-out data; the false-positive-budget mechanism, the 0.1 and 0.25 per cent settings, and the dataset-bias concession) https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation

https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation

Grounds: model org: wikipedia_ores

wikimediafoundationaGroundingVendorSave

Wikimedia Foundation. Moderator Tools/Automoderator and Moderator Tools/Automoderator/Testing (MediaWiki.org; the community-switch, the watchlistable configuration, the required false-positive page, and the caution-level table from the builders' own pre-deployment testing on 22 spreadsheets over about 600 edits) https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator

https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator

Grounds: model org: wikipedia_ores

EmpiricalThe deployed classifier's own bias against anonymous editors was measured independently and is large. Mykola T…

The deployed classifier's own bias against anonymous editors was measured independently and is large. Mykola Trokhymovych, Muniza Aslam, Ai-Jou Chou, Ricardo Baeza-Yates and Diego Saez-Trumper reported at KDD in 2023 a Disparate Impact Ratio of 20.02 for the deployed ORES model against a base-rate ratio of 7.93 in the same data, meaning it flagged anonymous editors well out of proportion even to their genuinely higher revert rate; the successor multilingual model reaches 9.54 with the same user features and 1.98 to 3.08 without them. The deployed model's area under the curve was 0.84 with precision at recall 0.75 of 0.22 on an unbalanced holdout, against 0.75 and 0.07 for a deliberately unfair baseline that simply reverts every anonymous edit. Model accuracy correlates negatively with a language edition's share of anonymous editors in both systems. The operator states the same limitation in its own documentation: the model card for the language-agnostic revert-risk model — the model its own reverting agent uses — says the model 'may exhibit bias against edits from new users, temporary accounts, or IP edits' and recommends the multilingual model instead for anonymous edits in the languages that model covers, while forbidding use of its predictions as ground truth for training other models and forbidding scoring a page's first revision.

trokhymovych2023GroundingAcademicSave

Trokhymovych, M., Aslam, M., Chou, A.-J., Baeza-Yates, R., & Saez-Trumper, D. (2023). Fair multilingual vandalism detection system for Wikipedia. KDD '23 (arXiv:2306.01650) https://arxiv.org/abs/2306.01650

https://arxiv.org/abs/2306.01650

Grounds: model org: wikipedia_ores

wikimediafoundationbGroundingVendorSave

Wikimedia Foundation. Machine learning models: Production language-agnostic revert risk and Production multilingual revert risk (Meta-Wiki model cards stating the deployed model's own bias against new users, temporary accounts and IP edits) https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

Grounds: model org: wikipedia_ores

EmpiricalPublishing the flag demonstrably changed the human decision, and the measured direction was toward more equal …

Publishing the flag demonstrably changed the human decision, and the measured direction was toward more equal treatment — a separate finding from the classifier's own bias, at a different layer, and the two must be carried together. Nathan TeBlunthuis, Benjamin Mako Hill and Aaron Halfaker ran a regression discontinuity across 23 Wikipedia language editions from January 2019 to March 2020, using the RCFilters threshold cutoffs as the discontinuity and revisions within 0.03 of a cutoff. At the 'maybe damaging' cutoff, being flagged raised revert probability from 13.5 to 19.2 percent for unregistered editors and from 4.6 to 14.3 percent for registered ones; at 'likely damaging', from 33.5 to 50.2 percent and from 15.5 to 44.5 percent respectively. Because flagging moved the under-scrutinised group more than the over-scrutinised one, aggregated across thresholds it INCREASED demographic parity between registered and unregistered editors. It also lowered the odds that a revert was itself controversial for unregistered editors, from 3.08 to 2.81 percent at 'likely damaging' and 3.33 to 2.92 percent at 'very likely damaging', which the authors read as evidence that flagging lowered the decision system's false-positive rate. The same authors state plainly in the same paper that 'ORES encodes biases against unregistered editors and - to a lesser extent - against editors without user pages'. The honest reading of the two findings together is a biased classifier whose published score reduced a larger human bias, and neither half stands alone.

teblunthuis2021GroundingAcademicSave

TeBlunthuis, N., Hill, B. M., & Halfaker, A. (2021). Effects of Algorithmic Flagging on Fairness: Quasi-experimental Evidence from Wikipedia. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1, Article 56 (arXiv:2006.03121) https://arxiv.org/abs/2006.03121

https://arxiv.org/abs/2006.03121

Grounds: model org: wikipedia_ores

Topics: algorithmic-fairness

trokhymovych2023GroundingAcademicSave

Trokhymovych, M., Aslam, M., Chou, A.-J., Baeza-Yates, R., & Saez-Trumper, D. (2023). Fair multilingual vandalism detection system for Wikipedia. KDD '23 (arXiv:2306.01650) https://arxiv.org/abs/2306.01650

https://arxiv.org/abs/2306.01650

Grounds: model org: wikipedia_ores

EmpiricalThe costs this deployment was designed against were measured on Wikipedia itself before the scoring service ex…

The costs this deployment was designed against were measured on Wikipedia itself before the scoring service existed, and are never a measured effect of it. Aaron Halfaker, R. Stuart Geiger, Jonathan T. Morgan and John Riedl reported in 2013 that the share of good-faith newcomers whose first-session edit was reverted rose from 6.1 percent in the first half of 2006 to 18.2 percent in the first half of 2007, that two-month survival of those newcomers fell from 25.6 percent to 11.7 percent within a year and did not recover, and that tool-mediated rejection of them rose from about 0 percent in 2006 to about 40 percent in 2010, with both rejection and tool-mediated rejection significant negative predictors of newcomer survival. The harm sat in the interaction as much as in the classification: reciprocation of a reverted newcomer's attempt to open a discussion averaged 7 percent for editors using Huggle, about 30 percent for Rollback, 53 percent for Twinkle and 56 to 67 percent for manual reverters, and 2,250 discussion attempts from 918 registered editors were addressed to an algorithmic editor that could not reply. The 2020 ORES paper concedes that after this research 'the often-hostile quality control processes that were designed over a decade ago remain largely unchanged'. Against that, the same score is also used prosocially: the 'Very likely good' filter, about 99 percent precise at over 90 percent recall, is documented as a way to find good-faith newcomers to thank, and a Wiki Education tool asks the article-quality model to re-score a student's draft with one more citation, header or image in order to recommend the most productive next edit.

halfaker2013GroundingAcademicSave

Halfaker, A., Geiger, R. S., Morgan, J. T., & Riedl, J. (2013). The Rise and Decline of an Open Collaboration System: How Wikipedia's Reaction to Popularity Is Causing Its Decline. American Behavioral Scientist, 57(5), 664-688 https://stuartgeiger.com/papers/abs-rise-and-decline-wikipedia.pdf

https://stuartgeiger.com/papers/abs-rise-and-decline-wikipedia.pdf

Grounds: model org: wikipedia_ores

halfaker2020GroundingAcademicSave

Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189

https://arxiv.org/abs/1909.05189

Grounds: model org: wikipedia_ores

Topics: co-design

EmpiricalAuditing this deployment is a tooled activity for the governed population rather than a privilege of the opera…

Auditing this deployment is a tooled activity for the governed population rather than a privilege of the operator, and its limit is a retention decision. ORES-Inspect, described by Zachary Levonian, Lauren Hagen, Lu Li, Jada Lilleboe, Solvejg Wastvedt, Aaron Halfaker and Loren Terveen at the 2024 Wiki Workshop, is an open-source Toolforge interface that lets any editor sample two disagreement quadrants — 'Unexpected Reverts', where the model called an edit fine and the community reverted it, and 'Unexpected Consensus', where the model called an edit damaging and the community left it standing — across the 35.6 million non-bot English Wikipedia edits of 2019, with the prediction as it was made at the time, and turn a single noticed misclassification into a quantified false-positive or false-negative rate for a chosen slice such as newcomers, LGBT-history pages or stubs. The design targets exactly the loop the deployment carries: reverts become training labels for the next model, so an over-flagging threshold can teach itself. The same paper records that historical predictions were retained only until the end of 2019, so an audit of what the model said at the moment of a past edit cannot run past that year even though the live scores have always been public.

levonian2024GroundingAcademicSave

Levonian, Z., Hagen, L., Li, L., Lilleboe, J., Wastvedt, S., Halfaker, A., & Terveen, L. (2024). ORES-Inspect: A technology probe for machine learning audits on enwiki. Wiki Workshop 2024 (arXiv:2406.08453) https://arxiv.org/abs/2406.08453

https://arxiv.org/abs/2406.08453

Grounds: model org: wikipedia_ores

halfaker2020GroundingAcademicSave

Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189

https://arxiv.org/abs/1909.05189

Grounds: model org: wikipedia_ores

Topics: co-design

EmpiricalThe label definition itself is contestable by the governed population here, and the training data is collected…

The label definition itself is contestable by the governed population here, and the training data is collected in the open. Tzu-Sheng Kuo, Aaron Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein and Haiyi Zhu reported at CHI 2024 on Wikibench, a system that puts AI evaluation-data curation through Wikipedia's ordinary talk-page and consensus machinery; the authors report that datasets curated this way 'can effectively capture community consensus, disagreement, and uncertainty' and that participants used it to refine label definitions, set data inclusion criteria and author data statements. The acting agent's own training labels come from a public Toolforge review interface whose maintainers state their aim plainly — 'We need volunteers to help review edits and classify them as either vandalism or constructive. We hope to eventually completely replace our current dataset with a random sampling of edits, reviewed and classified by volunteers' — with each edit in the trial slice behind their published statistics reviewed by at least two humans, and with the same documentation conceding that 'Our current dataset has some degree of bias, as well as some inaccuracies.' The older edit-quality models were trained on per-wiki volunteer labelling campaigns; the successor language-agnostic model was trained on published MediaWiki History and Wikitext History tables for January 2022 to January 2023 excluding bot edits on a 70/30 split, and the multilingual model on 8.6 million revisions from January to July 2022, sampled up to 300,000 per language, with a 17 percent unregistered-edit rate and an 8 percent revert rate in the training data.

kuo2024GroundingAcademicSave

Kuo, T.-S., Halfaker, A., Cheng, Z., Kim, J., Wu, M.-H., Wu, T., Holstein, K., & Zhu, H. (2024). Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia. CHI 2024 (arXiv:2402.14147) https://arxiv.org/abs/2402.14147

https://arxiv.org/abs/2402.14147

Grounds: model org: wikipedia_ores

englishwikipediaGroundingVendorSave

English Wikipedia. User:ClueBot NG/Documentation (volunteer maintainers' own published statistics on their own held-out data; the false-positive-budget mechanism, the 0.1 and 0.25 per cent settings, and the dataset-bias concession) https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation

https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation

Grounds: model org: wikipedia_ores

wikimediafoundationbGroundingVendorSave

Wikimedia Foundation. Machine learning models: Production language-agnostic revert risk and Production multilingual revert risk (Meta-Wiki model cards stating the deployed model's own bias against new users, temporary accounts and IP edits) https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

Grounds: model org: wikipedia_ores

halfaker2020GroundingAcademicSave

Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189

https://arxiv.org/abs/1909.05189

Grounds: model org: wikipedia_ores

Topics: co-design

EmpiricalLanguage coverage is the honest limit of this deployment's transparency, and the inequality runs the wrong way…

Language coverage is the honest limit of this deployment's transparency, and the inequality runs the wrong way. The better-calibrated multilingual revert-risk model covers 47 languages; the language-agnostic model runs on any of more than 250 language editions and is the model whose own card warns about bias against new users, temporary accounts and unregistered editors — and it is the model the operator's reverting agent actually uses. Across both systems, model accuracy correlates negatively with a language edition's share of anonymous editors, so the communities that lean most on anonymous contribution get the least accurate scoring. The estate also narrowed: roughly 110 classifiers in four families across 44 languages have given way to the revert-risk family. Precision figures are per-wiki and per-threshold and travel badly, which the operator's own help page demonstrates: on Polish Wikipedia the 'Likely have problems' filter captures 91 percent of problem edits against 34 percent for the corresponding English filter, and Polish Wikipedia therefore 'does not need - or have' the broader, noisier filter English Wikipedia relies on. English Wikipedia meanwhile keeps an agent trained on English Wikipedia alone, and the operator's own project page names running both as one of three choices open to that community.

wikimediafoundationbGroundingVendorSave

Wikimedia Foundation. Machine learning models: Production language-agnostic revert risk and Production multilingual revert risk (Meta-Wiki model cards stating the deployed model's own bias against new users, temporary accounts and IP edits) https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

Grounds: model org: wikipedia_ores

trokhymovych2023GroundingAcademicSave

Trokhymovych, M., Aslam, M., Chou, A.-J., Baeza-Yates, R., & Saez-Trumper, D. (2023). Fair multilingual vandalism detection system for Wikipedia. KDD '23 (arXiv:2306.01650) https://arxiv.org/abs/2306.01650

https://arxiv.org/abs/2306.01650

Grounds: model org: wikipedia_ores

wikimediafoundationaGroundingVendorSave

Wikimedia Foundation. Moderator Tools/Automoderator and Moderator Tools/Automoderator/Testing (MediaWiki.org; the community-switch, the watchlistable configuration, the required false-positive page, and the caution-level table from the builders' own pre-deployment testing on 22 spreadsheets over about 600 edits) https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator

https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator

Grounds: model org: wikipedia_ores

EmpiricalFour outages of the fastest automated tier in the first half of 2011 give this deployment a measured counterfa…

Four outages of the fastest automated tier in the first half of 2011 give this deployment a measured counterfactual, and what it measures is latency rather than coverage. R. Stuart Geiger and Aaron Halfaker, at WikiSym 2013, analysed ClueBot NG's downtime on 15 to 18 February, 13 to 17 March, 29 March to 7 April and 15 April to 1 May 2011. Comparing only Wednesdays and Thursdays in order to control for the weekly editing rhythm, median time-to-revert rose from 744 seconds with the bot running to 1,286 seconds with it down, and the geometric mean from 941 to 1,674 seconds. The proportion of edits that were reverts fell significantly during downtime (chi-square 115.9, p<0.001), but the proportion of revisions EVENTUALLY reverted did not differ (chi-square 0.64, p=0.43): the human tiers absorbed the work at a slower rate rather than losing it. The authors were careful not to read this as the bot being dispensable, asking instead what the workaround cost the editors who performed it. The same paper documents the reviewer tiers whose different clocks make that absorption possible: fully automated bots reverting within seconds, tool-assisted humans mostly within a minute, manual browser reverts between a minute and a day, and batch scripts on an idiosyncratic scatter. The measurement is excellent evidence for the shape of the effect and weak evidence for its present magnitude: it concerns one bot on one wiki at a scale and tooling mix that no longer obtain.

geiger2013GroundingAcademicSave

Geiger, R. S., & Halfaker, A. (2013). When the Levee Breaks: Without Bots, What Happens to Wikipedia's Quality Control Processes? WikiSym 2013 https://stuartgeiger.com/wikisym13-cluebot.pdf

https://stuartgeiger.com/wikisym13-cluebot.pdf

Grounds: model org: wikipedia_ores

EmpiricalThe transparency mechanism the peer-reviewed record identifies as the core of participatory machine learning h…

The transparency mechanism the peer-reviewed record identifies as the core of participatory machine learning here was removed in an infrastructure migration as an unused feature, and nobody decided against it. ORES exposed threshold optimisations in a machine-readable format so a wiki's tool developer could ask for the maximum filter rate at a stated recall for their own wiki and their own model version; the published worked example on English Wikipedia was a threshold of 0.32 giving a filter rate of 0.89, a false-positive rate of 0.087, precision of 0.23 and recall of 0.75. Probed live on 28 August 2026, that query returns: 'model_info query parameter is not supported by this endpoint anymore.' Chris Albon announced on the wikitech-l list on 3 August 2023 that the ORES API endpoint would move onto Lift Wing by 30 September 2023, that the Foundation wanted zero traffic on the old endpoint by January 2024, that 'The servers that run ORES are at the end of their planned lifespan and so to save cost we are going to shut them down in early 2024', and that 'The ores-legacy endpoint is not a 100% replacement for ores, we removed some very old and not used features.' The migration remains an open programme: a Phabricator task opened on 5 March 2026 records the deprecation guidance as scattered across three wiki pages and 'difficult to find and to maintain', and the MediaWiki modernization page carries its own warning that it 'contains outdated information that may not accurately reflect the current state of Wikimedia ML systems'. The published-score channel survived the migration; the published-fitness-statistics channel did not survive on this endpoint.

albon2023GroundingVendorSave

Albon, C., Wikimedia Foundation (2023, August 3). ORES To Lift Wing Migration (wikitech-l announcement), with the Wikitech ORES and Lift Wing pages and a live probe of both endpoints on 2026-08-28 showing the threshold-statistics query removed https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/EK65B7QCQHEG37C2ERPIUSP64OX3ZEUJ/

https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/EK65B7QCQHEG37C2ERPIUSP64OX3ZEUJ/

Grounds: model org: wikipedia_ores

halfaker2020GroundingAcademicSave

Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189

https://arxiv.org/abs/1909.05189

Grounds: model org: wikipedia_ores

Topics: co-design

EmpiricalThe correction channel for a wrongly reverted editor is an ordinary public wiki page rather than a ticket nobo…

The correction channel for a wrongly reverted editor is an ordinary public wiki page rather than a ticket nobody outside can see, and its openness is exactly what makes it unmeasured. The acting agent's documentation tells anyone reverted in error to redo the edit, remove the warning and report the false positive, and points at a page that also holds the full public list of reported false positives; that page was reachable when checked on 28 August 2026. The operator's own agent makes creating such a page a required step of deployment and links it from the talk-page message, from the page history and from the user's contributions beside the ordinary Undo and Thank actions, in a default message that reads: 'Because the model I use is not perfect, it sometimes reverts good edits. If you believe the change you made was constructive, please report it here.' The agent's team states an intention to investigate retraining on reported false positives. What does not exist is a disposition: because no operator adjudicates anything, a false-positive report is a wiki page rather than a case, so this deployment publishes no reversal rate on contested reverts and cannot supply the overturn statistic that a mandated platform transparency report supplies. Its correction channel is more open and less measured than a regulated one.

englishwikipediaGroundingVendorSave

English Wikipedia. User:ClueBot NG/Documentation (volunteer maintainers' own published statistics on their own held-out data; the false-positive-budget mechanism, the 0.1 and 0.25 per cent settings, and the dataset-bias concession) https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation

https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation

Grounds: model org: wikipedia_ores

englishwikipediaaGroundingVendorSave

English Wikipedia. User:ClueBot NG/FalsePositives and the ClueBot NG dataset review interface on Toolforge (a public list of reported errors and an open interface for labelling training data) https://en.wikipedia.org/wiki/User:ClueBot_NG/FalsePositives

https://en.wikipedia.org/wiki/User:ClueBot_NG/FalsePositives

Grounds: model org: wikipedia_ores

wikimediafoundationaGroundingVendorSave

Wikimedia Foundation. Moderator Tools/Automoderator and Moderator Tools/Automoderator/Testing (MediaWiki.org; the community-switch, the watchlistable configuration, the required false-positive page, and the caution-level table from the builders' own pre-deployment testing on 22 spreadsheets over about 600 edits) https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator

https://www.mediawiki.org/wiki/Moderator_Tools/Automoderator

Grounds: model org: wikipedia_ores

EmpiricalThe scale this deployment exists to address is published, and so is the labour arithmetic behind it. Queried f…

The scale this deployment exists to address is published, and so is the labour arithmetic behind it. Queried from the Wikimedia Analytics REST API on 28 August 2026, English Wikipedia alone took between 2.59 and 2.94 million non-bot edits to content pages per month across the twelve months to July 2026 — 2,839,865 in July 2026 — on the order of 85,000 to 95,000 human content edits a day on one of more than 250 language editions the models serve, against an operator-stated base rate of fewer than 5 problem edits in 100. The ORES paper's all-editions figure was about 290,000 edits a day, and its arithmetic is that reviewing that at an aggressive ten revisions a minute is about 483 volunteer labour hours daily, which a model filtering ninety percent of the stream reduces to about 48.3 — turning 240 volunteers at two hours a day into 24, and, for a small wiki, turning the task into one or two part-time volunteers. The service that does the filtering ran at 50 to 125 external requests a minute in steady state with bursts to 400 to 500 a second, precaching requests roughly an order of magnitude higher because a scoring job starts for nearly every edit, an approximately 80 percent cache hit rate, and most predictions computed in about a second. The team behind it 'never had more than 3 paid staff and 3 volunteers at any time, and no more than 2 requests typically in progress simultaneously', against roughly 66,000 monthly active English Wikipedia editors at the time.

halfaker2020GroundingAcademicSave

Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189

https://arxiv.org/abs/1909.05189

Grounds: model org: wikipedia_ores

Topics: co-design

EmpiricalThere is no litigation, no regulator, no court, no consent order and no statutory transparency mandate anywher…

There is no litigation, no regulator, no court, no consent order and no statutory transparency mandate anywhere in this deployment's record, and that absence is a structural finding rather than an absence of controversy. Litigation posture: None. Every accountability artefact here — the open scoring interface, the published operating points, the model cards stating their own biases, the public false-positive pages, the community audit tooling — exists because the operator and the self-governing volunteer communities chose it, and could be withdrawn the same way; the removal of the machine-readable threshold-statistics query in the 2023 to 2024 infrastructure migration is a small, dated instance of exactly that. Two further honesty notes belong with any description of the strength of this record. First, this is not a solved or harm-free deployment: its own operators concede the hostile quality-control processes documented in 2013 remain largely unchanged, the deployed model's card warns it may be biased against new users, temporary accounts and unregistered editors, and the acting agent's maintainers concede their dataset carries bias and inaccuracies. Second, the peer-reviewed evidence base leans on a small overlapping author group — Halfaker and Geiger appear on the ORES paper, the outage paper and the newcomer-decline paper, and Halfaker also co-authors the flagging-fairness paper and the audit probe, having been a Wikimedia Foundation employee for most of the period covered — with the KDD evaluation and the audit probe the clearest checks outside that lineage, and the KDD paper the one that measures the deployed model unfavourably.

halfaker2020GroundingAcademicSave

Halfaker, A., & Geiger, R. S. (2020). ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia. Proceedings of the ACM on Human-Computer Interaction 4, CSCW2, Article 148 (author preprint arXiv:1909.05189v3) https://arxiv.org/abs/1909.05189

https://arxiv.org/abs/1909.05189

Grounds: model org: wikipedia_ores

Topics: co-design

wikimediafoundationbGroundingVendorSave

Wikimedia Foundation. Machine learning models: Production language-agnostic revert risk and Production multilingual revert risk (Meta-Wiki model cards stating the deployed model's own bias against new users, temporary accounts and IP edits) https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

https://meta.wikimedia.org/wiki/Machine_learning_models/Production/Language-agnostic_revert_risk

Grounds: model org: wikipedia_ores

englishwikipediaGroundingVendorSave

English Wikipedia. User:ClueBot NG/Documentation (volunteer maintainers' own published statistics on their own held-out data; the false-positive-budget mechanism, the 0.1 and 0.25 per cent settings, and the dataset-bias concession) https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation

https://en.wikipedia.org/wiki/User:ClueBot_NG/Documentation

Grounds: model org: wikipedia_ores

trokhymovych2023GroundingAcademicSave

Trokhymovych, M., Aslam, M., Chou, A.-J., Baeza-Yates, R., & Saez-Trumper, D. (2023). Fair multilingual vandalism detection system for Wikipedia. KDD '23 (arXiv:2306.01650) https://arxiv.org/abs/2306.01650

https://arxiv.org/abs/2306.01650

Grounds: model org: wikipedia_ores

levonian2024GroundingAcademicSave

Levonian, Z., Hagen, L., Li, L., Lilleboe, J., Wastvedt, S., Halfaker, A., & Terveen, L. (2024). ORES-Inspect: A technology probe for machine learning audits on enwiki. Wiki Workshop 2024 (arXiv:2406.08453) https://arxiv.org/abs/2406.08453

https://arxiv.org/abs/2406.08453

Grounds: model org: wikipedia_ores

albon2023GroundingVendorSave

Albon, C., Wikimedia Foundation (2023, August 3). ORES To Lift Wing Migration (wikitech-l announcement), with the Wikitech ORES and Lift Wing pages and a live probe of both endpoints on 2026-08-28 showing the threshold-statistics query removed https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/EK65B7QCQHEG37C2ERPIUSP64OX3ZEUJ/

https://lists.wikimedia.org/hyperkitty/list/wikitech-l@lists.wikimedia.org/thread/EK65B7QCQHEG37C2ERPIUSP64OX3ZEUJ/

Grounds: model org: wikipedia_ores