Domain Atlas / Content moderation & editorial AI
The CyberTipline: triage under a rule against looking
Explore this deployment in the PAN Lab ↗
In the PAN Lab, the readouts of this case's model organization carry a shaded evidence band whose width follows the least-established class among the modeling inputs the readings rest on.
The least-established input behind this case's model organization's readings is an assumption, not a measurement. Evidence base: 7 assumed · 7 measured.
The CyberTipline is the single congressionally authorised reporting mechanism for online child sexual exploitation in the United States, built by the National Center for Missing & Exploited Children in March 1998, when it received 2,772 reports in its first calendar year. The congressionally mandated transparency report to the appropriations committees gives the recent series: 36,210,368 reports in calendar 2023, 20,512,803 in 2024 and 21,351,493 in 2025, the 2025 arrivals carrying 61,833,177 files. Of the 2025 total, 21,181,300 came from electronic service providers and 170,193 from members of the public, a 99.2 to 0.8 per cent split, and the public channel carried more than 5,700 reports directly from the person depicted. More than 2,000 providers are registered, just over 300 submitted any report in 2025, and five accounted for more than 75 per cent. The automated element at the centre is not a classifier: it resolves where a report belongs and matches it against entities already in the record across fields such as electronic mail addresses and network addresses, with analyst review on the matching, and no accuracy figure for it is published by anyone. The matching is exact-match; fuzzy matching that would catch a suspended account's near-identical new handle is not implemented, and material attached to a report is not automatically scanned for matches. The clearinghouse's own resolution failures are published as a series: reports whose location could not be determined ran 1,368,404 (3.8 per cent) in 2023, 1,957,640 (9.5 per cent) in 2024 and 3,044,434 (14.3 per cent) in 2025, and reports whose location cannot be resolved are made available to United States federal law enforcement by default, so a failure of resolution is itself a routing rule.[4]
What happened
Start with the statute, because every operational fact in this case is downstream of three sentences in it.
18 U.S.C. § 2258A(a) requires a provider to report an apparent violation to the CyberTipline "as soon as reasonably possible after obtaining actual knowledge". § 2258A(b) says the report "may, at the sole discretion of the provider, include" the identity of the person involved, the historical reference showing when the material was uploaded, the geographic location including the IP address, the visual depictions themselves, and the complete communication. § 2258A(f) says nothing in the section requires a provider to monitor any user or to "affirmatively search, screen, or scan" for violations. And § 2258A(c) says NCMEC "shall make available each report" to the relevant law enforcement agencies.
Read those together and the design is unusual and deliberate. Reporting is compulsory. Detecting is not. Everything that would make a report usable is optional. And the recipient may not filter. The governance object that results is a queue.
NCMEC built the CyberTipline in March 1998 and received 2,772 reports in its first calendar year. The congressionally mandated transparency report to the appropriations committees gives the recent series in one document: 36,210,368 reports in calendar 2023, 20,512,803 in 2024, 21,351,493 in 2025. The 2025 arrivals carried 61,833,177 files. Of those, 21,181,300 came from electronic service providers and 170,193 from members of the public — 99.2 per cent to 0.8 per cent — and the small channel is the only one with a human author, carrying more than 5,700 reports directly from the person depicted, up over 100 per cent on the year before.
The senders are concentrated and volatile. More than 2,000 providers are registered; just over 300 submitted anything in 2025; five accounted for more than 75 per cent. Between 2024 and 2025 one company's volume rose more than thirty-six-fold, from 30,759 to 1,105,405, while the largest sender's roughly halved, from 8,590,357 to 4,907,710. The arrival rate at a fixed-capacity clearinghouse is set entirely by product decisions taken elsewhere, and the operator's own tables say so: "there are no legal requirements for proactive efforts to detect this content or what information an ESP must include in a CyberTipline report. As a result, both the volume and content of reports can vary greatly."
What arrives is also substantially the same thing arriving again. Congress asked for the measurement and the 2025 report supplies it: of 29,408,181 images submitted, 19,091,252 were unique by exact hash and only 13,994,568 were distinct under visual-similarity matching; of 26,324,863 videos, 15,144,788 and 7,304,334. Roughly 35 per cent of the file volume is exact or near duplicate of something already in the store. Hundreds of reports may concern one person: the Belgian Federal Police reported receiving over 500 distinct CyberTipline reports about a single offender in five months.
Then comes the part that makes this case unlike any other in the atlas. The clearinghouse frequently may not look at what it is triaging.
In United States v. Ackerman (10th Cir. 2016) the court held that NCMEC qualifies as a governmental entity in light of its authorising statutes and the functions Congress gave it, and in the alternative acted as a government agent, so its opening and viewing of reported files was a warrantless search. In United States v. Wilson (9th Cir. 2021) the court held the private-search exception does not cover files the platform never opened. NCMEC disagrees with the first holding — its staff described the decision as painful and as reflecting confusion about NCMEC having an investigative role "when it's merely a middleman" — describes itself as a private non-profit, prints that disclaimer at the foot of every report it sends and in its law-enforcement tooling, and complies anyway. It opens only the files a platform employee checked as viewed, for reports bound for United States law enforcement.
The practice reduces to one indication on someone else's form, added to the reporting form at the start of 2014. A 2026 appellate record shows it working file by file. In United States v. Lowers (4th Cir. 2026), a platform's hashing flagged 156 files uploaded to one account; a reviewer at the platform opened 31 of them and confirmed apparent CSAM; the report identified which roughly 20 per cent had been viewed and which 80 per cent had not. The opinion states it plainly: "An employee at NCMEC received that CyberTip and opened and viewed the same 31 images as the Google Reviewer. The NCMEC employee did not open any of the remaining 125 unreviewed files."
The same rule cuts both ways in the same organisation on the same day. The no-opening practice applies only to reports bound for United States law enforcement; NCMEC is able to open a file with the indication absent where the report will go abroad. And 77.1 per cent of 2025 reports resolved outside the United States. The same file is examinable or not by the same analyst depending on where the report is going.
The knock-on effect is the study's own causal chain. Platforms over-report to reduce their own risk, including widely circulated items shared without malice. If the platform did not record a prior view, NCMEC cannot open the file to see what it is. Not every platform uses the potential-meme indication either. So the report is forwarded as potentially actionable without context, and a receiving officer may obtain a warrant only to find a meme. Had either indication been present, NCMEC could have labelled the report informational.
That label is the clearinghouse's only lever over the downstream load, and it is defined by the absence of information rather than by the seriousness of the conduct. The operator's own words: "An informational report contains severely limited information in which there is no apparent child sexual exploitation nexus; or so little information was provided by the reporting party that it is impossible to identify a location to refer the report to; or contains frequently seen child sexual exploitation or abuse material that has been shared in a non-malicious context, such as for inappropriate comedic effect or moral outrage or concern for the child depicted." US law enforcement typically reads "informational" as meaning a report can be set aside. Not every report that could be set aside carries it.
The clearinghouse's ability to say where a report belongs is also degrading, and the operator publishes the series. Reports whose location could not be determined ran 1,368,404 (3.8 per cent) in 2023, 1,957,640 (9.5 per cent) in 2024 and 3,044,434 (14.3 per cent) in 2025. Unresolvable reports are made available to United States federal law enforcement by default, so a failure of jurisdiction resolution is itself a routing rule. The Lowers record shows the other failure shape end to end: files uploaded 20 September 2019, reported 23 September, forwarded to one Virginia county on 29 October, left there for half a year, a subpoena on 16 April 2020 showing the address was in a different city, the file closed on 13 May, and the receiving city's detective applying for a warrant on 27 May 2020. Eight months from upload to warrant, one wrong jurisdiction, and — under the 90-day preservation rule then in force — a preservation window that had lapsed twice.
Downstream, capacity is appropriated rather than scaled. In fiscal 2025 the Office of Juvenile Justice and Delinquency Prevention funded the 61 ICAC task forces, a network of more than 6,200 federal, state, local and Tribal agencies, at $33,976,146 across a competition with 61 expected awards and a published award maximum of $1,042,765. In 2025 those task forces conducted nearly 347,000 investigations leading to more than 17,000 arrests and trained about 73,000 professionals. NCMEC made 1,932,435 reports available to them the same year — roughly double the 2023 figure of 908,762, while total arrivals fell 41 per cent. That is a routing shift, not a change in what is happening.
Field-study respondents describe the load in queue terms and they do not agree with each other. One officer: "You have a stack [of CyberTipline reports] on your desk and you have to be ok with not getting to it all today. There is a kid in there, it's really quite horrible." A single task force detective may carry 2,000 reports a year. One local officer said 90 per cent of what reached him was garbage — striking because his task force had already filtered it — while an officer in a better-resourced department said very few were truly unactionable. The study's own reading is that actionability is partly a function of the receiving agency's resources. Asked whether more detectives would fix it, one officer said: "I could have ten of me, but I need a team of people who could help me execute search warrants, interview everyone, forensically process [devices]... It's a lot of work for just one tip."
The channel that would settle any of this exists and is barely used. NCMEC built a structured feedback system: case status (conviction, arrest, ongoing investigation, referred, closed), whether a child victim was identified on arrest, ten named closure reasons including no crime committed, false report and unfounded, and a direct question on whether the information NCMEC supplied was useful. It states that agencies "are not generally required by law to provide feedback on CyberTipline reports, and NCMEC has no authority to require such feedback be submitted" and that "most agencies provide little or no feedback." In 2025 task force units returned 549,584 feedback instances against 1,932,435 reports; federal agencies, which received 3,435,257, returned 7,085; local agencies returned 156. So the system has run for a quarter century without an empirical prioritisation rule, and the study is explicit that it is unknown what share of reports, if fully investigated, would reveal hands-on abuse.
Quality is a per-sender property the sender does not pay for, and since April 2026 it is published. Responding to a 16 March 2026 oversight letter from the Senate Judiciary Chairman, NCMEC supplied 2025 figures for eight companies accounting for over 17 million reports, 81 per cent of the total, and the committee released them: one artificial intelligence service's more than 1.1 million reports contained zero per cent actionable information because the service was designed not to collect user or content data; over 80 per cent of one messaging platform's more than 752,000 reports were deemed inactionable for insufficient information; over 90 per cent of another sender's more than 135,000 were originally inactionable; one platform supplied location information in 4 per cent of its 2025 reports against 35 per cent in 2024; another routinely submitted unrelated content; a fifth's omissions of location or account information rendered reports inactionable. Those are NCMEC's answers as characterised in a committee majority release. They are oversight findings about data completeness, not enforcement findings, and the companies were pressed for responses rather than found to have violated anything.
Volume moved sharply in the other direction first, and the parties disagree about why. CY 2024 brought 20,512,803 reports against CY 2023's 36,210,368, the largest drop in the programme's history. NCMEC's chief legal officer said her first question was whether a company had stopped reporting or gone out of business, and that "there wasn't anything like that"; NCMEC's analytics attributed the fall almost entirely to default end-to-end encryption on one platform's Facebook and Messenger surfaces, with that platform still supplying over 67 per cent of the total and reporting 6.9 million fewer incidents than in 2023. That platform's own account is different in emphasis: it partnered with NCMEC on a bundling feature grouping duplicate viral or meme content into a single report, which it says contributed significantly to the drop and let NCMEC and law enforcement manage and prioritise more easily. NCMEC says unbundling the reports to count every incident still leaves a 7 million-report gap between the years. Two other companies claimed consolidation, and an NCMEC spokesperson said any such changes were "not via the official feature in the CyberTipline reporting pipeline". All four accounts belong together or none of them do.
Congress read the decline as non-compliance. On 30 April 2025 the author of the REPORT Act opened an inquiry with four companies over what her office called a sharp decline in reports since the Act's passage. The Act itself, passed a fortnight after the field study published, had already changed the constraint set: it added minor sex trafficking and enticement to the duty (the categories it added then grew by roughly an order of magnitude in eighteen months — online enticement from 186,819 reports in 2023 to 1,413,347 in 2025), extended preservation from 90 days to one year, raised failure-to-report penalties, required preservation consistent with the national cybersecurity framework, and extended NCMEC's limited liability to contracted vendors, which is the provision that makes commercial cloud hosting of this data workable at all.
That last provision answers a constraint the study had documented. As of 2024 NCMEC could not house CyberTipline data on commercial cloud services, because while NCMEC held limited legal liability for hosting the material other entities did not — which blocked scaled classification work and machine translation for foreign recipients. Two of the study's specific technical recommendations were unshipped capacity rather than new ideas: a commissioned interface matching report IP addresses against peer-to-peer file-sharing data, which would let an officer tell a single-item report with a sharing history apart from one without, was completed in autumn 2020 and, as of 2024, had not been integrated; and an offer of cloud translation capacity was not taken up, partly for engineering resource and partly over the risk of inaccurate translations.
NCMEC's own answer arrived in 2025-2026: a declared $10 million, three-year CyberTipline Modernization Initiative to build "a faster, more resilient, and more scalable platform" that will "reduce processing times, identify urgent cases sooner, support law enforcement more effectively", supported by three cloud and analytics companies with further corporate investors. It is an announced programme. No published measurement of its effect exists, and this file describes it as announced rather than delivered.
Two constraints are worth stating together at the end because they are the same constraint. The organisation will not write down the guidance it gives senders: staff told the researchers that if there were a written document, "defense attorneys would characterize this in criminal cases as NCMEC is advising companies what to report." And nothing else anywhere sets a standard either. NCMEC concedes it lacks authority to make platforms change their reporting; it has no authority to require feedback from any agency; and the statute still says every useful field is included at the sole discretion of the provider. Congress has raised the penalty for failing to report. It has not written a content standard.
The stressor the study predicted has arrived as volume in the meantime. NCMEC recorded more than 400,000 CY 2025 reports with a generative-AI nexus on its own count, more than 182,000 involving offenders possessing, generating or attempting to generate such material, and more than 158,000 files so categorised; the figure released through Senate Judiciary oversight for the same year was 1.5 million reports with such a connection, including over 12,000 reports of the material found in AI training data. Two counts on two bases, both stated. The study had warned that NCMEC's first million-report day, caused by one widely circulated item, was survivable only because of automated clustering, and that a million genuinely distinct generated images would not be.
The sociotechnical reading
Most governance cases in this atlas are about a decision system that got a decision wrong. This one is about a queue that cannot be shortened, whose contents are composed by parties with no duty to compose them well, and whose operator is legally forbidden to discard anything and constitutionally discouraged from looking.
Begin with the incentive geometry, because it is unusually clean. A platform faces exactly one price signal in this arrangement: a failure-to-report penalty, now between $600,000 and $1,000,000 per violation and scaled by user base. It faces no penalty for reporting something that turns out to be nothing, no duty to search, and no requirement to include any particular field. Reporting broadly is therefore individually rational, and the practitioner reading of the statute says so out loud — the "apparent" violation standard means providers should report even where age or content is ambiguous. The cost of that rationality lands on a clearinghouse that may not filter and on 61 task forces whose funding is set by appropriation. This is a congestion externality with a statutory prohibition on charging for it.
The second thing to see is that the clearinghouse's central constraint is not technical. There is no classifier at the decision point. The automated element resolves jurisdiction and matches entities across fields, and its documented failure is a report sent to the wrong place or to nowhere. The thing that decides whether a report can be triaged well is a field on the sender's form, and the reason that field matters is a line of Fourth Amendment cases. That is a governance problem no amount of model improvement reaches.
The third is the shape of the epistemic asymmetry, which runs in a direction that keeps surprising people. The organisation that receives the evidence may not examine most of it. The organisation that generated it — the platform — examined as much or as little of it as it chose, and the appellate record notes what nobody knows about that: no evidence of how the platform trains its reviewers, how accurate they are, or how accurate its hashing is in practice. So the one human judgement the whole downstream chain rests on is made by a party outside the deployment, at a standard nobody has measured, and recorded in a single indication.
The fourth is the loop that would have made all of this legible and does not close. NCMEC built a well-designed structured feedback schema — case status, victim identified, ten closure reasons, a direct usefulness question — and it is voluntary at both ends. Federal agencies returned 7,085 instances on 3.4 million reports. Because that channel is empty, nobody can answer the question the system exists to answer: which reports were worth the hours. The study says it plainly, and it is the most important sentence in the record: it is unknown what share of reports, if fully investigated, would reveal hands-on abuse. A quarter century of operation has produced no empirical prioritisation rule, and that is not because anyone failed to build the instrument. The instrument is built. Nothing obliges anyone to use it, and the operator states it has no authority to require that they do.
The fifth is that the only channel that has ever changed this deployment's behaviour is a court, and it changed the wrong variable from the operator's point of view. Ackerman and Wilson did not make the triage better; they narrowed what the triaging organisation may see, and the organisation complied while disagreeing. The oversight channel, which does hold instruments that reach the input, has used them to raise penalties for failing to report and to publish per-sender quality figures — real work, and in April 2026 it produced measurement nothing inside the deployment could have produced. What it has still not done is write a standard for what a report must contain. Every part of this record points at that one change, and no instrument anywhere has been pointed at it.
The sixth is about numbers, and it is why this file states denominators so insistently. The headline count is the most-quoted figure in this policy area and it is not a count of children. Hundreds of reports may concern one person; a third of the file volume is duplicate; a large share is older material recirculating where the child is already known and safe; the operator's own 2022 figure was that 49 per cent of reports were actionable. And the inference does not run backwards either — every interviewee with a view believed the underlying threat is understated. One respondent put the cost of the confusion exactly: "We aren't doing a good enough job of selling the threat... The number gets trotted out to justify everything, and then people wonder why they don't get resources."
Finally, the system's only working safety valve is upstream withdrawal, and it is not a governance instrument at all. When one platform turned on default encryption, 15.7 million reports stopped arriving. That reduced the load without improving anything, and the operator counts it as pure loss. A pipeline whose only relief comes from its inputs disappearing is a pipeline nobody is steering.
Two boundaries hold in this file and they are not decoration. Served people are not modelled: the children in the material, the people a platform reported on information it alone chose the contents of, and the people whose report went to the wrong agency are all outside the operator network the Lab board draws, and no outcome for any of them is computed from anything on it. And the registers stay apart — the statute, the mandated report, operator statement, respondent perception and committee correspondence are five different weights of evidence, and this file never lets one borrow the authority of another.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.