Domain Atlas / Content moderation & editorial AI
The score is published and the service cannot act on it
Explore this deployment in the PAN Lab ↗
In the PAN Lab, the readouts of this case's model organization carry a shaded evidence band whose width follows the least-established class among the modeling inputs the readings rest on.
The least-established input behind this case's model organization's readings is an assumption, not a measurement. Evidence base: 2 assumed · 7 published baseline · 3 measured.
Wikipedia's edit-scoring service publishes a damage or revert-risk probability for essentially every edit as it is saved and has no path of its own by which it can act on one. Aaron Halfaker and R. Stuart Geiger, writing the system up for CSCW in 2020, describe it as built to decouple four activities normally performed by the same engineers: curating training data, building models, auditing predictions, and building the interfaces or bots that act on predictions. The operator did not gate access either, on its own stated reasoning that 'Given the open API, there is no barrier where we can selectively decide who can request a score from a classifier.' Acting on the score is done by separately governed agents: a volunteer-run bot on English Wikipedia, an operator-built agent local administrators switch on, tool-assisted patrollers working score-ranked queues, and ordinary editors with score-driven filters enabled in their own preferences. At the time of that paper the estate ran roughly 110 classifiers across 44 languages in four families; the ORES infrastructure has since been retired and what runs today is Lift Wing, with ores-legacy.wikimedia.org as a compatibility endpoint in front of it, serving a revert-risk family that replaced the older edit-quality models. Verified live on 28 August 2026: the legacy host answers, self-identifies as the 'ORES legacy service' and still returns scores, and the Lift Wing endpoint returns a revert-risk score for an arbitrary revision without any credential.[2]
What happened
A scoring service computes a damage or revert-risk probability for essentially every edit to every supported Wikipedia as the edit is saved, caches it, and publishes it over an interface that asks for no credential. Nothing in that service reverts, blocks, warns or reports anything. The design intent was explicit: Aaron Halfaker and R. Stuart Geiger, writing up the system for CSCW in 2020, describe it as built to decouple four activities normally done by the same engineers — curating training data, building models, auditing predictions, and building the interfaces or bots that act on predictions. The operator did not gate access either, and said why: "Given the open API, there is no barrier where we can selectively decide who can request a score from a classifier."
Acting on the score is a separate job under separate governance, and there are four kinds of actor doing it. A fully automated bot, ClueBot NG, has edited English Wikipedia since October 2010 and had made 6,682,890 edits as of 28 August 2026, holding the bot, reviewer, rollbacker and autoconfirmed user groups — all of them community-granted and community-revocable. Automoderator, built by the Wikimedia Foundation's Moderator Tools team, scores every main-namespace edit and reverts above a threshold; it is deployed on twelve Wikipedias, and not on English. Tool-assisted patrollers work score-ranked queues in Huggle and similar tools. And ordinary editors watch recent changes with the score-driven filters switched on in their own preferences.
The threshold is the actual policy in a damage-detection system, and here it is chosen by the community that will live with it. On English Wikipedia the operating points are published to patrollers in ordinary help text, with both of their costs stated: the strictest quality filter is "right more than 90% of the time" but "finds only about 10% of all the problem edits", while the broadest is "right only about 15% of the time" and catches "about 82% of problem edits", against a stated base rate of "fewer than 5 in 100". The same architecture yields different arrangements on different wikis, because the model and the threshold are both local: on Polish Wikipedia the "Likely have problems" filter captures 91 percent of problem edits against 34 percent for the corresponding English one, and Polish Wikipedia therefore "does not need - or have" the broad, noisy filter English Wikipedia relies on.
The volunteer bot's control variable is not a score at all but an error budget. Its documentation states that "The threshold is not randomly chosen by a human, but is instead calculated to match a given false positive rate... A human selects a false positive rate, which is the percentage of constructive edits incorrectly classified as vandalism." At the current setting of 0.1 percent the bot catches approximately 40 percent of vandalism; at the previous 0.25 percent setting it caught approximately 55. Fifteen points of catch rate were spent, deliberately and in public, to halve the wrongful-reversion rate. Both halves of that trade are published, which is rare enough in this domain to be worth saying plainly. These are the maintainers' own figures, computed on their own held-out, human-reviewed slice: methodologically described, and not independently audited.
The Foundation's own reverting agent was built so that it would not be an operator enforcement node. "After deployment, Automoderator will not begin running until a local administrator turns it on." Its configuration lives at a Community Configuration page that writes an ordinary, watchlistable wiki page, so a change to the operating point appears in the same feed as an edit to an article. Creating a false-positive reporting page is a required step of deployment. It never reverts administrators, global sysops, stewards or bots, never reverts self-reverts or reverts of its own actions, and never touches new page creations. And the choice of threshold is put to a community as a table of consequences rather than a slider with no units: scores above 0.99 at about 100 percent precision and roughly 152 English Wikipedia reverts a day, 0.985 at about 95 percent and 350, 0.98 at about 93 percent and 680, 0.975 at about 82 percent and 1,077. The builders withdrew the least cautious option themselves for falling below their own 90 percent accuracy target. That table is their pre-deployment testing, aggregated from 22 completed review spreadsheets covering over 600 edits across six projects, not measured production performance.
When a volunteer did wire a score straight to automatic reversion, the community stopped it and the model's operators were not involved. PatruBOT auto-reverted Spanish Wikipedia edits above a threshold its developer had set too low; Spanish Wikipedians crowd-sourced an audit of its errors on ordinary wiki pages, reached a consensus that it was making too many mistakes, and an administrator blocked the bot's account. The ORES team wrote that this "was entirely a community governed activity that required no intervention of our team or the Wikimedia Foundation staff", and recorded that a successor, SeroBOT, later resumed auto-reverting at a higher confidence threshold.
Two measurements of fairness exist here, they point in opposite directions, and the record only makes sense if both are held at once. Nathan TeBlunthuis, Benjamin Mako Hill and Aaron Halfaker ran a regression discontinuity across 23 language editions between January 2019 and March 2020, using the filter cutoffs as the discontinuity. Being flagged raised revert probability from 13.5 to 19.2 percent for unregistered editors and from 4.6 to 14.3 percent for registered ones at the "maybe damaging" cutoff, and from 33.5 to 50.2 and 15.5 to 44.5 percent respectively at "likely damaging" — so the flag moved the under-scrutinised group more than the over-scrutinised one and narrowed a pre-existing gap. It also lowered the odds that a revert of an unregistered editor was itself contested, from 3.08 to 2.81 percent at one cutoff and 3.33 to 2.92 at another. Those authors state in the same paper that "ORES encodes biases against unregistered editors and - to a lesser extent - against editors without user pages". Mykola Trokhymovych and colleagues, at the KDD conference in 2023, measured that classifier bias directly: a Disparate Impact Ratio of 20.02 for the deployed model against a base-rate ratio of 7.93 in the same data, with the successor multilingual model at 9.54 with the same user features and 1.98 to 3.08 without them. The deployed model's area under the curve was 0.84 with precision at recall 0.75 of 0.22; a deliberately unfair baseline that simply reverts every unregistered edit scores 0.75 and 0.07. A biased classifier whose published score reduced a larger human bias is the honest summary, and it is why where the threshold sits, and who sets it, is the thing that matters here.
The costs this arrangement is designed against were measured on Wikipedia itself before any of it existed. Aaron Halfaker, R. Stuart Geiger, Jonathan T. Morgan and John Riedl reported in 2013 that the share of good-faith newcomers whose first-session edit was reverted rose from 6.1 percent in the first half of 2006 to 18.2 percent a year later, that two-month survival of those newcomers fell from 25.6 to 11.7 percent and did not recover, and that tool-mediated rejection of them rose from roughly zero in 2006 to about 40 percent in 2010. The harm was in the interaction as much as the classification: reverted newcomers who tried to open a discussion were answered 7 percent of the time by editors using Huggle, about 30 percent for Rollback, 53 percent for Twinkle and 56 to 67 percent for manual reverters, and 2,250 discussion attempts were addressed to an algorithmic editor that could not reply. This is the harm the scoring service was built in response to, not a measured effect of it — and the ORES paper concedes that after that research "the often-hostile quality control processes that were designed over a decade ago remain largely unchanged".
Auditing here is a tooled activity for the governed population rather than a privilege of the operator. ORES-Inspect, described by Zachary Levonian and colleagues in 2024, lets any editor sample the two disagreement quadrants — edits the model called fine that the community reverted, and edits it called damaging that the community left standing — across 35.6 million non-bot English Wikipedia edits from 2019, with the prediction as it was made at the time, and turn one noticed error into a quantified rate for a slice of their choosing. Wikibench, a CHI 2024 field study by Tzu-Sheng Kuo and colleagues, put evaluation-data curation through Wikipedia's ordinary talk-page and consensus machinery and reported that participants refined label definitions, set inclusion criteria and authored data statements. The label definition itself is contestable, and the bot's maintainers concede their own dataset "has some degree of bias, as well as some inaccuracies" while running a public interface asking volunteers to help replace it.
The operator publishes model cards as a standing practice, written to the Mitchell et al. framework, on wiki pages that carry talk pages. The card for the language-agnostic revert-risk model — the model Automoderator actually uses — says it "may exhibit bias against edits from new users, temporary accounts, or IP edits" and recommends the multilingual model instead for anonymous edits in the languages that model covers. It also forbids using the model as ground truth for training other models, and forbids scoring a page's first revision. Coverage is the honest limit of the whole transparency story: the better-calibrated model — the one whose scores match the rate at which flagged edits really do turn out to be reverted — covers 47 languages against more than 250 editions, and across both systems model accuracy correlates negatively with a language edition's share of anonymous editors, so the communities that lean most on anonymous contribution get the least accurate scoring.
There is one measured counterfactual for what an automated damage-detection channel actually contributes, and it is worth the space. ClueBot NG went down four times in the first half of 2011 — 15 to 18 February, 13 to 17 March, 29 March to 7 April, and 15 April to 1 May. Comparing only Wednesdays and Thursdays to control for the weekly editing rhythm, Geiger and Halfaker found median time-to-revert rose from 744 seconds to 1,286 and the geometric mean from 941 to 1,674 while the bot was down. Significantly fewer reverts happened in those windows, chi-square 115.9, p<0.001 — but the proportion of revisions eventually reverted was statistically indistinguishable, chi-square 0.64, p=0.43. The human tiers absorbed the work more slowly. The authors declined to read that as the bot being dispensable, asking instead what the workaround cost the editors who performed it.
ORES no longer exists as infrastructure and has not since early 2024. Chris Albon announced on the wikitech-l list on 3 August 2023 that the API endpoint would move onto Lift Wing by 30 September 2023, that the Foundation wanted zero traffic on the old endpoint by January 2024, and that "The servers that run ORES are at the end of their planned lifespan and so to save cost we are going to shut them down in early 2024". What runs today is Lift Wing, a model-serving platform on Kubernetes, with ores-legacy.wikimedia.org as a compatibility shim in front of it. Verified live on 28 August 2026: the host answers, self-identifies as the "ORES legacy service", and still returns scores from the old model names; the Lift Wing endpoint returns revert-risk scores for an arbitrary revision without any credential. The published-score channel survived the migration intact.
One channel did not, and it is the one the peer-reviewed record identifies as the core of participatory machine learning here. ORES exposed threshold optimisations in a machine-readable format, so a wiki's tool developer could ask, in effect, for the maximum filter rate at recall of at least 0.75 for their wiki and their model version; the published worked example on English Wikipedia was a threshold of 0.32 giving a filter rate of 0.89, a false-positive rate of 0.087, precision of 0.23 and recall of 0.75. That query now returns: "model_info query parameter is not supported by this endpoint anymore." The migration announcement had warned that "The ores-legacy endpoint is not a 100% replacement for ores, we removed some very old and not used features." Nobody voted to end participatory threshold-setting. The mechanism that implemented it was removed as an unused feature when the servers reached end of life, and a Phabricator task opened on 5 March 2026 records that the deprecation documentation is still scattered across three pages and "difficult to find and to maintain". The historical audit trail is thinner than the live one for the same kind of reason: ORES-Inspect notes that historical predictions were retained only until the end of 2019, so an audit of what the model said at the moment of a past edit stops there.
The scale problem all of this exists to address has not gone away. English Wikipedia alone took between 2.59 and 2.94 million non-bot edits to content pages per month over the twelve months to July 2026, on the order of 85,000 to 95,000 a day on one language edition. The operator's own arithmetic on the all-editions figure is that reviewing roughly 290,000 edits a day at an aggressive ten a minute would be about 483 volunteer labour hours daily, which a model that filters ninety percent of the stream reduces to about 48. Against that stands a team that, on its own account, "never had more than 3 paid staff and 3 volunteers at any time, and no more than 2 requests typically in progress simultaneously".
Litigation posture: None. There is no court, no regulator, no consent order and no statutory transparency mandate anywhere in this record. That is the structural fact that makes this case worth reading beside every other deployment in the domain — and it cuts both ways, because every accountability artefact here exists because the operator and the communities chose it, and could be withdrawn the same way. The retirement of the threshold-statistics query is a small, dated, documented instance of exactly that.
The sociotechnical reading
Most cases in this atlas sit inside one organisation: a model, the people who act on it, the records it writes, and someone whose job is to check. This one has the same four parts and puts each of them under a different authority, and almost everything interesting follows from where those lines fall.
Start with what the automated element cannot do. The scoring service has no enforcement path. It computes a probability and writes it to a public record, and that is the end of its powers. The thing that acts on the score reads it out of the same public record an outside researcher reads. There is no private wire between the classifier and the actor, which means the interface between "what the machine thinks" and "what happens to an edit" is an artefact anybody can inspect, argue with, and copy. Read against the closed pipelines elsewhere in this domain — where the classifier, the threshold, the acting agent and the appeal channel all sit inside one operator and the public record is whatever that operator chose to publish — this is the same task with the joints exposed.
The second structural fact is that the operating point is chosen by the people who will live with it, three times over: once for the wiki, once in a tool configuration page anybody may edit and watch, and once by each patroller who sets a personal minimum. That is a governance arrangement rather than a feature, and it is what makes the published precision and recall load-bearing rather than decorative. Numbers that describe a choice you cannot make are trivia; numbers that describe a choice you must make are the choice. The bot maintainer's error budget is the sharpest version: pick a maximum false-positive rate and let the threshold be computed from it, so the question a human answers is "how much wrongful reversion is acceptable" rather than "what number feels right".
Third, the checking here runs sideways rather than downward. The people who are governed are the people who hold the levers: a bot may not edit until a local approvals group has cleared it and it has run a trial; any administrator may block one that misbehaves; the permission the fastest tools require is granted and revoked by the same community; the operator's own agent will not start until a local administrator turns it on. Four of the five governance hops in the underlying PAN model are checks between classes rather than influence between them. That is not a milder version of a regulator; it is a different shape entirely, and its failure modes are different too. It has no floor. Nothing compels a volunteer maintainer to act on a published finding, and nothing preserves an instrument that stops being convenient to run.
Fourth, the record loops back on itself in a way this atlas sees everywhere and can rarely watch. Reverts become the labels the next model trains on, so an operating point that surfaces one kind of edit teaches the next model to surface it again. Here the check on that loop exists and is public — sample the two quadrants where the model and the community disagreed, and read what the model said at the time against what people then did — and its limit is a retention decision rather than a design one: predictions were kept only through the end of 2019.
The two fairness measurements are the hardest thing in the file to hold steady, and the Lab network draws both. TeBlunthuis, Hill and Halfaker measured the effect of the FLAG on the human decision and found it raised revert probability far more for registered editors than for unregistered ones, narrowing a pre-existing profiling gap; Trokhymovych and colleagues measured the CLASSIFIER's own disparate impact at 20.02 against a base rate of 7.93. Neither refutes the other. They are measurements of different layers, and the useful lesson is that a biased instrument placed inside a well-arranged decision process can improve the process's fairness while remaining biased — which is an argument for arranging the process, not for excusing the instrument.
What this deployment cannot supply is the statistic every regulated platform in this domain does supply. Because no operator adjudicates anything, a false-positive report is a wiki page rather than a case with a disposition, and there is no published reversal rate on contested reverts. Its correction channel is more open and less measured than a mandated one: any reverted editor is told a machine did it, in a message that concedes the model sometimes undoes good edits, and can add their case to a public list that anybody can read — and nobody publishes what share of that list was upheld.
Finally, the pressure this network carries is not a scandal but an attrition. One infrastructure decision, taken to save cost at the end of a server's planned life, narrowed two things at once: the sensor, because the machine-readable fitness-statistics query that made threshold-setting participatory was dropped as an old, unused feature; and the model estate, because roughly 110 classifiers in four families across 44 languages gave way to a revert-risk family whose language-agnostic member runs anywhere and whose better-calibrated multilingual member covers 47 of more than 250 editions. Nobody decided against participatory machine learning. The thing that implemented it was removed while nobody was looking at that particular line item, and there was no regulator, court or reporting duty anywhere in the arrangement whose job it would have been to notice.
Read this board against the enforcement pipelines beside it and the question it answers is a specific one: which harms in those networks come from the classifier, and which come from the governance wrapped around it. This deployment has a measurably biased classifier and it publishes the measurement. What it does not have is a closed pipeline.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.