# Laundered statistic

Part of [The AI Tells Index](https://feedsquad.com/ai-tells). A tell signals low effort. It does not identify an author. Skilled writers produce every shape listed here, some of them daily, and automated detectors misread those writers for it at rates measured above 60 percent on non-native English prose. Nothing in this index proves that a machine wrote anything. Read an entry as one piece of evidence to weigh against the false-positive notes printed beside it.

## Facts

- Id: laundered-statistic
- Category: Semantic (semantic)
- Subcategory: statistics
- Also known as: benchmark stripped of its scope, headline drift
- Status: Active. Currently signals low-effort writing.
- Severity: high
- Evidence grade: corroborated (named independently by multiple credible secondary sources)
- Languages: en
- Added: 2026-08-15
- Updated: 2026-08-15
- Page: https://feedsquad.com/ai-tells/laundered-statistic

## Description

A real number from real work, republished without the scope that made it true. The study measured one population with one instrument over one period. The sentence quoting it names none of that, so the number arrives larger than it was.

## Why it reads as machine-written

Summarisation drops qualifiers, and every hop drops a few more, so a benchmark result becomes a general fact in about three republications. The 2024 paper introducing generative engine optimization reports visibility gains of up to 40 percent for its own methods on the benchmark its authors built, measured against the engines of 2023 and 2024. It circulates without any of that. A 2025 measurement of new English-language articles crossing the halfway mark for AI generation circulates as a claim about the internet, and its own detector reports a 4.2 percent false-positive rate and a 0.6 percent false-negative rate and did not evaluate text where a human and a tool worked together. Those clauses have to travel with the number.

## Detection

Type: judge (a rubric for a lightweight model judge)

Rubric for a lightweight model judge:

```text
Task: compare each imported statistic with the scope of the work it came from.

Step 1. For every statistic attributed to a study, report or benchmark, find the original.
Step 2. Record the original's population, sample size, date range, instrument and stated limits, including any error rate.
Step 3. Record the population the citing sentence implies.
Step 4. Compare, and mark each of these separately: population widened; date dropped; benchmark result stated as a general effect; error rate dropped from a detector-derived figure; a subset presented as the whole; a maximum presented as a typical value.

Escape hatches: the sentence carries the scope; the original itself claims generality and defends it; the figure is a rounded version of a widely replicated result; the citing text is a headline with the scope in the body immediately below.

Decision. FLAG on any of the six marks in Step 4. Report the original scope and the published scope side by side so the reader can see the distance.

Output: original scope, published scope, marks, verdict.
```

## Examples

Constructed specimens. Written for this index. Never quoted from anyone's posts.

### Specimen 1

Before, exhibiting the tell:

> Optimizing for generative engines raises visibility by 40 percent, and more than half the internet is now AI-generated.

After, repaired:

> A 2024 paper reports visibility gains of up to 40 percent for its own methods, measured on the benchmark its authors built against the search engines of 2023 and 2024. Separately, a 2025 sample of new English-language articles put the AI share past half; the detector behind that number reports a 4.2 percent false-positive rate and did not evaluate work where a person and a tool collaborated. Both are worth citing. Neither is a fact about the internet.

Note: The figures here are the published ones and the scope clauses are what the sources actually state. This entry is the one this index is most likely to breach itself, which is why its own numbers were checked against the sources before it shipped.

## False positives

Who legitimately writes this way.

Good-faith summarisation oversimplifies, and it has to: a sentence cannot carry a methods section. Subeditors cut qualifiers for length, headlines are written by people who did not write the piece, and press offices publish the widest defensible version of their own findings. Researchers themselves generalise in interviews in ways their papers do not. The pattern this entry describes is systematic scope-stripping across a body of work rather than one loose sentence, and the repair is usually a clause rather than a retraction. Anyone applying it should apply it here first, since a directory that publishes numbers is the easiest place in the world to commit this.

## Model attribution

No vendor can be named. Scope-stripping is a property of chains of republication, and the documented examples travelled through human marketing writing.

## Sources

1. GEO, generative engine optimization benchmark (arXiv:2311.09735)
   https://arxiv.org/abs/2311.09735
   (tier: primary-doc; accessed 2026-08-14)
2. Graphite, more articles are now created by AI than humans
   https://graphite.io/five-percent/more-articles-are-now-created-by-ai-than-humans
   (tier: press; accessed 2026-08-14)

## Status history

Ids are permanent. A retired tell keeps its id and its page.

- 2026-08-15, Active: Two worked examples with reachable primaries: a benchmark result circulating as a general effect, and a measurement of new articles circulating as a measurement of the web. Both originals state their scope plainly, which is what makes the drift visible.

## License

CC BY 4.0. https://creativecommons.org/licenses/by/4.0/

Attribution: The AI Tells Index, feedsquad.com/ai-tells

Reuse the data, including commercially. Keep the attribution line.

---

Dataset version 1.0.0. Schema version 1.
Part of [FeedSquad](https://feedsquad.com). Built by [Herman Foundry](https://hermanfoundry.com) from Levi, Finnish Lapland.