Chat control proposes one law for two problems
In the German state of North Rhine-Westphalia (NRW) in 2024, prosecutors specialised in child sexual abuse material (CSAM) declined to open around 2,300 investigations after concluding that the people referred to them had done nothing wrong, according to Markus Hartmann, the senior prosecutor who heads ZAC NRW, the state’s central cybercrime unit. Accounts on several social media platforms had been hijacked and used to post abuse material; automated detection flagged the results; the referrals arrived; and each one had to be read by a prosecutor before it could be dropped.
Thirteen prosecutors handle this work at ZAC NRW, completing roughly 11,600 investigations against identified suspects and 3,200 against unidentified suspects in 2025 – well over a thousand cases each. They operate at capacity, and EU policymakers are considering handing them the output of a general scanning obligation on top.
The EU Child Sexual Abuse Regulation returns to trilogue in September under an Irish Presidency that has historically sided with the Council’s pro-scanning bloc. Five rounds have failed over one question: whether providers can be made to search private communications without suspicion. Underneath it sits a drafting problem that no compromise on that question will fix.
Two axes: who and what to scan
Any scanning mandate varies along two independent dimensions.
The first is who is scanned – a targeted population, or everyone. This governs the base rate, and it is why suspicionless scanning generates a large volume of false positives. This implicates an inference problem, and it is mitigable: better models, higher thresholds, and more review capacity all move it in the right direction.
The second is what is being looked for – CSAM already classified, or novel material. Critics have long flagged this distinction; the Chat Control 1.0 versus 2.0 shorthand delineates it, and Parliament’s position would confine detection to hash-matching against known material. But the case has always been argued from privacy and proportionality: unknown-material detection is less accurate, more intrusive, more dangerous, and therefore to be constrained more tightly within a single instrument.
This axis governs whether an error rate can be well-defined at all. It embodies a specification problem that nothing mitigates; there is no quantity there to improve. We should not be legislating the conditions under which to do something whose accuracy cannot, even in principle, be established beforehand.
Two problems: one with an answer sheet, and one without
Matching content against a database of previously identified material is a lookup. Analysts review an image, confirm its classification as abuse material, and compute a short numerical fingerprint (hash) which goes on a shared list. The U.S. National Center for Missing and Exploited Children (NCMEC) alone has shared over nine million hashes with technology providers and partner organisations. A scanner computes the fingerprint of each image passing through it and checks for a match. That makes errors measurable: the technology is tested against known material.
Identifying material no one has seen before is a different task. There is no answer sheet to test the technology against. By definition, none can be supplied.
A limit – and a blast from the crypto royalty past
Statistics has understood the shape of this since I. J. Good's 1953 work with Alan Turing on unseen species. They showed that a sample can reveal approximately how much probability mass lies in unobserved categories, but not identify those categories. Even that estimate assumes a fixed distribution – a generous assumption, given that criminals evade detection.
Scientists famously end papers by calling for more research. Here we need research to stop. Research fills gaps in what we happen not to know. A statute cannot enable scientists to identify something that cannot be identified.
What the numbers we do have are measuring
The favourable case is not clean, either. Internal research at ZAC NRW on classifiers for material already seized under reasonable suspicion found false positive rates of five to seven per cent. These are research systems, not tools in operational use. They are tuned to minimise false negatives, because missing something is the greater risk when examining seized evidence; a higher false positive rate is the accepted price. Where to make that trade-off is a choice; that there is one is not. Where signal and noise distributions overlap, every threshold buys fewer misses at the cost of more false alarms.
Such classifiers therefore flag irrelevant material, too: not only a baby in a nappy, but flowers. Image classifiers learn whatever regularities in the training data happen to predict the label, rather than the semantic category a lawyer has in mind – shortcut learning, in the literature. The flowers are not a bug we will engineer away. They are a product of the structure of the world and the laws of mathematics. And that is a purpose-built tool under the most favourable conditions available, on a targeted population, with a warrant already in hand.
Moreover, existing independent evaluations test against surrogate datasets, because the operational CSAM database on which hash-matching tools like PhotoDNA run is inaccessible to independent researchers. The widely quoted industry figure of one false positive in a trillion has no published methodology behind it, and Steinebach notes it may be partly evaluation and partly mathematical derivation.
What none of the texts separate
The Commission’s proposal does distinguish between target classes. It sets different conditions for each, caps how long an order may run – longer for known material, shorter for solicitation – and tightens reporting as the target grows more sensitive.
Every one of those escalations tracks how intrusive the search is. Not one tracks whether its error rate can be established. The reliability standard is a single sentence applied identically to all three targets: technologies must be “sufficiently reliable, in that they limit to the maximum extent possible the rate of errors regarding the detection” (Art. 10(3)(d)). Recital 28 verifies this by regular assessment of false negative and false positive rates on anonymised representative data samples – and a representative sample of material defined by never having been identified is the one object that cannot be assembled.
Parliament narrowed the order by suspicion – who may be searched, on whose authority. The Council deleted the detection chapter outright and added a recital stating that nothing in the Regulation imposes a detection obligation. Three redesigns, all on the first axis.
Splitting the instrument
The fix is a drafting fix, and it is available now.
For detection against an enumerated list, keep something close to the current architecture and add what is already achievable: performance measured on unfiltered traffic, reported at intervals, published. Providers can meet that obligation.
For open-class detection, make an order conditional on something the current text does not require – a validation design, registered before deployment, specifying how ground truth for novel material will be obtained, by whom, against what baseline, and on what schedule. Not results, which do not yet exist. A design. If the validation problem remains unsolved, that is the answer.
There is precedent for insisting. The National Academies’ 2003 “lie detection” report concluded that polygraph proponents could not establish net non-harm, because the validation problem was unsolved. Twenty-three years later, Europe is asked to mandate detection technology with the same unsolved problem.
The Council’s November 2025 position deleted the detection chapter but required the Commission to assess, within three years, the necessity and feasibility of detection obligations, including an evidence-based assessment of the reliability and accuracy of available technologies. The questions remain open – postponed to an evaluation that, for the unknown-material half, cannot be performed.
This proposal will satisfy nobody, which is weak evidence in its favour. The privacy camp will dislike that it concedes detection against a known list could be made to work. The child-protection camp will dislike being told that the thing it most wants – to find abuse in progress by scanning everyone for images no one has catalogued – cannot be evaluated.
Until the Council responds to Parliament’s encryption amendments, the text can still be made to distinguish between a detection program we know how to check and one we do not. Two problems. Two instruments. The alternative is a law that cannot tell them apart, enforcing a standard only one of them can meet.