Domain Atlas / Content moderation & editorial AI
Google CSAM detection and total account closure
Explore this deployment in the PAN Lab ↗
In the PAN Lab, the readouts of this case's model organization carry a shaded evidence band whose width follows the least-established class among the modeling inputs the readings rest on.
The least-established input behind this case's model organization's readings comes from a published baseline, not this deployment's own record. Evidence base: 10 published baseline · 2 measured.
In February 2021 a father in San Francisco photographed his toddler son's swollen, painful groin because an advice nurse asked for images ahead of an emergency telehealth consultation during pandemic-era remote care; the doctor used the photographs to diagnose the infection and prescribed antibiotics, which cleared it up. The images auto-uploaded from an Android phone to Google Photos, and two days later his entire Google Account was disabled for 'harmful content' that was 'a severe violation of the company's policies and might be illegal'. Google reported the material to the National Center for Missing & Exploited Children's CyberTipline, and San Francisco police served search warrants on Google and on his internet service provider within a week of the photographs, seeking 'everything in Mark's Google account: his internet searches, his location history, his messages and any document, photo and video'. He learned of it in December 2021, when an envelope arrived containing the warrants and a letter telling him he had been investigated. THE POLICE CLEARED HIM: the investigator, with access to everything Google held, concluded that 'the incident did not meet the elements of a crime and that no crime occurred'. A near-identical case ran in parallel in Houston, where a father photographed his toddler's genital infection at a pediatrician's request, the images auto-backed up and were sent to his wife over a Google messaging service, and his decade-old paid account was locked while he was in the middle of buying a house; he too was cleared, quickly, after showing a detective his correspondence with the pediatrician. Both men appealed with the exculpatory material in hand: 'A few days after Mark filed the appeal, Google responded that it would not reinstate the account, with no further explanation.' Google publicly stood by the decisions and has never conceded error in either case. Asked directly in December 2025, the reporter who broke the story said neither parent had recovered his account, though one had been able to retrieve some account data that was turned over to police; that answer is relayed second-hand by the writer who asked her, and no first-party statement of the outcome exists.[5]
What happened
In February 2021 a father in San Francisco photographed his toddler son's swollen, painful groin. He did it because an advice nurse asked for images ahead of an emergency telehealth consultation, in a period when pandemic-era practice had pushed routine paediatric care onto video. The doctor used the photographs to diagnose the infection and prescribed antibiotics, which cleared it up. The photographs were medically requested and medically used.
They also auto-uploaded from an Android phone to Google Photos. Two days later the father's phone made a notification noise: his account had been disabled because of "harmful content" that was "a severe violation of the company's policies and might be illegal". What went was not a photograph. It was the whole Google Account — Gmail going back more than a decade, Google Drive documents, Google Photos including the entire photographic record of his son's first years, contacts for friends and former colleagues, and his Google Fi phone service, which meant obtaining service and a new number from another carrier.
Google reported the material to the CyberTipline operated by the National Center for Missing & Exploited Children, as 18 U.S.C. section 2258A requires it to do once it has actual knowledge of an apparent violation. The report routed to police. Within a week of the photographs being taken, San Francisco police had served search warrants on Google and on the family's internet service provider, seeking, in the New York Times's account, "everything in Mark's Google account: his internet searches, his location history, his messages and any document, photo and video". He learned none of this at the time. He learned it in December 2021, when an envelope arrived from the police department containing copies of the warrants and a letter telling him he had been investigated.
THE POLICE CLEARED HIM. The investigator, with access to everything Google held, concluded that "the incident did not meet the elements of a crime and that no crime occurred". No charge was ever brought.
A near-identical case ran in parallel. A father in Houston photographed his toddler's genital infection at a pediatrician's request; the images auto-backed up to Google Photos and were sent to his wife over a Google messaging service. His decade-old paid account was locked while he was in the middle of buying a house, disrupting the transaction. He too was cleared, quickly, after showing a detective his correspondence with the pediatrician.
Both men appealed with the exculpatory material in hand. Neither got anywhere. The Times reported the first appeal's outcome in one sentence: "A few days after Mark filed the appeal, Google responded that it would not reinstate the account, with no further explanation." Google's spokesperson at the time said the company follows US law in defining what constitutes this material and uses "a combination of hash matching technology and artificial intelligence to identify it and remove it from our platforms". The company publicly stood by the decisions. It has never conceded error in either case.
## What actually happened in the pipeline, and why nothing in it malfunctioned
Google describes its detection as two technologies used in combination and augmented by human review. Hash matching compares uploads against verified sets of hashes of previously confirmed material, drawn from the Internet Watch Foundation, NCMEC and content Google itself confirms, each hash independently verified before deployment. Separately, machine-learning classifiers trained on confirmed material "flag new content that is very similar to patterns of previously confirmed CSAM" and sort it into a queue for specialist human confirmation. Reviewers with backgrounds in law enforcement, child advocacy and social work then confirm the item against the federal definition before anything happens.
The detection path in both cases was the second one, and the distinction matters. The photographs were newly created, so no hash of them could have existed anywhere; the flag had to come from the classifier for never-before-seen material. The Times reported exactly that: the images were flagged by the artificial intelligence, and "a human content moderator for Google would have reviewed the photos after they were flagged by the artificial intelligence to confirm they met the federal definition of child sexual abuse material". Two trade outlets attributed the flag to Microsoft PhotoDNA hash matching, one of them in its own headline slug. That attribution is wrong for these photographs and is not repeated here.
So run the sequence again with nothing broken in it. The classifier surfaced content resembling previously confirmed material, which is precisely what it was built to do. A trained specialist looked at the image and agreed it met the federal definition. The item genuinely was a photograph of a child's genitals. The one variable that would have changed the answer — a clinician asked for it — lived outside every input the system had. Google told the Guardian that staff who review this material are trained by medical experts to look for rashes or other issues, that the reviewers are not themselves medical experts, and that medical experts were not consulted when reviewing each case. The expertise needed to read a clinical image was in the loop as training and never as a consult.
## Why no published number covers this
Google publishes measured error counts, and they cover the other channel. In its regulated filings under Regulation (EU) 2021/1232 it reported 18 content items incorrectly flagged by hash matching in 2023 and 10 in 2024, every one of them caught by human review during detection, none removed, none reported externally, and no account access lost. The 2025 filing reports 1,604 items automatically flagged as known material, 335 subject to human review, and a reported error rate of zero.
For the classifiers, Google's position is that because they only sort and prioritise content for human confirmation, "there is no risk of false positives by reason of this technology alone". Under that definition the medical-photo cases are, by construction, not counted as detection errors in any published figure. That is a fact about the scope of the measurement, not an accusation that any published number is false. It is also why an operator can truthfully report a handful of hash-matching mis-flags in a year in which its classifier-driven pipeline suspended accounts at a rate its own round number puts near 270,000.
The scope of those filings narrows further. They cover Google's messaging and mail services in the European Union. They exclude Google Photos and YouTube — the two services in which every documented closure in this file happened.
## What an appeal does, in the operator's own words
Across three consecutive European reporting years, Google Ireland's filings give: 297 accounts appealed and 10 reinstated in 2023; 216 appealed and 19 reinstated in 2024; 254 complaints lodged and 7 accounts restored in 2025, with 6 initial content verdicts overturned on review and the files made available to the users for download. Judicial complaints in all three years: zero. (The 2025 figures use a new Commission standard form whose denominator is complaints rather than appealing accounts, so read the three as a trend and not as a series.)
Beside those counts sits the sentence that explains them, repeated almost verbatim each year: a reinstatement "was not due to an error in detection or a content-level false positive, but rather a reinstatement based on contextual information identified during the appeal process, which indicated that the content was correctly identified but did not appear to be possessed or shared with intent to harm, abuse, or exploit children."
The appeal re-reads intent. It does not re-read the image. And a police finding that no crime occurred is a conclusion about the person, produced by an actor outside the platform, arriving at a channel whose stated authority reaches the inference and not the verdict — which is why two men holding the strongest available counter-evidence found the form had no slot for it.
The redress budget is also finite and published. Google's help text states that "For some policy violations, Google will review up to 2 appeals", that data download is unavailable for certain violations "including but not limited to: Valid legal requests, Account hijacking, Egregious content violations", and that if an appeal is not approved "your entire Google Account will remain unavailable... your account will be permanently disabled and considered for deletion."
## The channel that carried the truth, and could not carry it back
Google's 2025 filing describes its outbound leg in its own word: "While Google's reports to the NCMEC CyberTipline are one-way reporting, the information sharing and collaboration with NCMEC and NGOs provide the necessary feedback loop to continuously improve Google's detection technology."
Stanford's April 2024 study of the same ecosystem, built on dozens of interviews with platforms, NCMEC and law enforcement, found the same thing independently: platforms "rarely get either" outcome information or report-quality feedback; NCMEC built a structured law-enforcement outcome field into the report flow and law enforcement rarely fills it in. One platform respondent described where that leaves a company: without feedback "you are stuck in a system where turning over anything is better than trying to think through how to do this well."
The study proposed the exact mechanism these two cases needed. NCMEC should "publish a negative hash set of images that have been reported as CSAM but have been verified to not be violative... This would allow platforms to stop reports (and automated processes such as account termination) on known legal content." NCMEC's April 2024 response appreciated the analysis, disputed nothing specifically, and said it would explore the recommendations. No such published negative hash set was located as of this file's verification date.
## The channel that worked has no standing at all
On 28 October 2022, two months after the reporting, Google's VP of Trust and Safety Operations published the company's account of the pipeline and said Google was "actively working on ways to increase transparency" about suspension reasons and to improve the appeals experience. On 30 December 2022 the New York Times reported the shipped result — a more specific reason and a path to supply context — and recorded a mother in Colorado recovering her account after four months, following a Times inquiry. That piece also summarised the two fathers in one line: "The police determined that the fathers had committed no crime, but the company still deleted their accounts."
In December 2023 the Times documented a third shape. A mother in Australia lost her whole Google Account after her seven-year-old uploaded a video to YouTube. The upload was flagged within minutes. Her repeated appeals were denied, including on a paid support channel. The account was restored one day after a Times reporter asked about it. Google's statement was that "we understand that the violative content was not uploaded maliciously", and the company "had no response for how to escalate a denial of an appeal beyond emailing a Times reporter."
The policy change is real and it is partial. It altered what a person is told and what they may submit. It did not alter the sanction, the two-appeal cap, the one-way referral, or the fact that an exculpatory finding has nowhere to land. In February 2026 trade press reported a fresh wave of Google Photos false-positive account bans, with one appeal rejected in about ten minutes against a stated review window of up to two days; that report aggregates user posts with no operator comment, and is carried here only for the claim that the pattern persists.
## The pressure runs the other way at the same time
On 4 March 2026 a United States senator opened an investigation into Google for failing to remove child sexual abuse material and assist survivors, demanding by 18 March the internal detection and removal policies, victim removal-request response times since January 2020, annual CyberTipline reports broken out by product, every case where content was not removed within 48 hours, Trust and Safety staffing levels and budgets, and any internal decision limiting deployment of detection technology. Those are a senator's allegations and demands, not adjudicated findings, and no Google response was located.
They belong in this file for one reason: they price the asymmetry every reinstatement decision sits on. Under-detection draws a congressional document demand and a statutory penalty schedule running to a million dollars. A wrongly destroyed account draws a form with room for up to two appeals.
## Where the litigation stands
Neither documented father sued, and no class action or regulatory enforcement over these facts was located. That absence is not vindication. The two decisions that do exist both went Google's way on other questions. Baker v. Google LLC (D.D.C., 26 July 2024) dismissed a self-represented plaintiff's contract, fraud and due-process challenge to a CSAM-based account termination at the pleading stage, holding among other things that "Defendant Google is a private business, not a state actor". State v. Rauch Sharak (Wis., 24 February 2026) held unanimously that Google "acted as a private actor — not as an instrument or agent of the government — when it scanned Rauch Sharak's files and an employee opened and viewed files flagged as CSAM", reasoning from section 2258A(f)(3) and from section 230(c) being "entirely passive". That case involves a convicted defendant, has no connection whatever to the medical-photo cases, and is cited here only for the legal architecture.
Asked directly in December 2025 whether either father ever recovered his account, the reporter who broke the story said neither had, though one was able to retrieve some of the account data that had been turned over to police. That answer reaches this file second-hand, through the writer who asked her. No operator statement, court record or first-party account confirms it.
The sociotechnical reading
The instructive thing about this deployment is that its failure is not a defect anywhere. Read the pipeline as a sequence of components and every one of them performed to specification. The classifier surfaced never-before-seen content resembling confirmed material; that is its function. The specialist reviewer applied the federal definition to an image of a child's genitals and confirmed it; that is the check, and it fired. The statutory report went out as soon as reasonably possible after actual knowledge; that is the law. The account action followed the confirmation; that is the policy. Call this a true-positive-shaped false positive: the system was right about every question it was capable of asking, and the question that mattered — did a clinician ask for this — was outside every input it had.
That shape has three consequences worth reading off the network.
The first is a measurement consequence. A pipeline whose designed check on a machine judgement is a human confirmation can, quite honestly, report that the machine carries no false-positive risk of its own. Google says exactly that, and its published error counts cover the hash-matching channel, where the mis-flags are countable and were all caught during detection. The channel that produced these cases is the one with no published number anywhere. So the deployment's most consequential error mode is, by the definition its own reporting uses, invisible in its own reporting. Nothing here is falsified; a category is simply drawn so that the failure does not land in it.
The second is a sanction consequence. The unit of enforcement is not the item but the person's identity infrastructure. One confirmation replicates across mail, documents, photographs, contacts, telephone service and every third-party account reachable only through that sign-in, and the record contains no partial remedy: the operator's own help text runs from suspension to permanent disablement and consideration for deletion. That is a governance design in which the error cost is enormous, indivisible, and borne entirely by a party with no seat in the process.
The third is the correction consequence, and it is the reason this case belongs in a network model rather than an anecdote. Correction requires a channel, and every channel here is pointed the wrong way. The referral leaves and does not return: the operator's own filing calls it one-way, and independent field work found platforms rarely receive outcome or quality feedback and that the clearinghouse's structured outcome field is rarely completed. The appeal exists and re-reads the wrong thing: the operator says three years running that a reinstatement follows from context about intent rather than from an error in detection or a content-level false positive. The one channel with a demonstrated success rate — a reporter asking — has no formal standing at all, and the operator was recorded having no answer for how to escalate a denial beyond it. The party that reached the correct verdict, an investigator holding warrant returns on an entire account, sits on the far side of the one-way leg; and the constitutional doctrine that keeps the scan outside the Fourth Amendment, by insisting the platform is a private actor and not a government agent, is the same doctrine that makes a platform wary of taking direction, or evidence, from law enforcement. The exonerating fact is not merely unheard. It is unhearable by construction.
Two amplifiers sit on top of that. A confirmation does not stay a judgement: material the operator confirms joins the reference sets its comparison runs against, and the reporting records that the clearinghouse can add newly confirmed material to its own database of known images, so an error is promoted to a matchable fact for platforms this operator does not run. And the decision boundary itself is exported — the same classifier layer is distributed to other platforms through the operator's partner toolkit — so the error mode is replicated across an industry rather than contained within one company. A field study recommended the obvious counterweight, a published set of material reported as abusive but verified not to be, precisely so platforms could stop reports and automated account terminations on known legal content. It had not shipped.
Finally, read the gradient the whole thing sits on, because it explains why a reasonable operator would build it this way. A federal statute prices a failure to report at up to a million dollars for the largest providers while expressly disclaiming any duty to search or scan. Section 230 immunises the restriction decision and grants no corresponding protection for getting a reinstatement wrong. And in the same year the medical-photo shape was recurring, a United States senator opened an investigation into the operator for removing too slowly, demanding staffing, budgets and any decision limiting detection deployment. Under-detection is expensive, legible and politically salient. Over-detection is cheap, invisible in the error metrics, and absorbed by one person at a time. A platform respondent's summary of what that gradient produces was the plainest sentence in the record: without feedback, turning over anything is better than trying to think through how to do this well.
What a governance model can do with this is narrow and worth stating precisely. It cannot compute a harm rate: no denominator exists, the volume figures come from different periods and counting bases, and a trade analysis noted at the time that most people affected by this shape will never speak publicly. What it can do is show which pathways carry a correction and which do not, and let a reader see that the instruments that would help sit before the sanction rather than after it — a halt in the two days between an automatic backup and a whole-account closure, a reconciliation between a promoted entry and what an investigation concluded, a measurement of the channel nobody measures, and a scheduled duty to argue the other side before a confirmation becomes an account.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.