2026 precision-hedge cluster
- Id
- precision-hedge-cluster-2026
- Status
- Active
- Severity
- medium
- Detection
- statistical
- Evidence grade
- community-observed
- Languages
- en
- Added
- 2026-08-15
- Updated
- 2026-08-17
Currently signals low-effort writing.
What it is
The current frontier vocabulary profile, which is nothing like the 2023 one. We extracted the CC0 LexA English science split for GPT-5.2 on 2026-08-14 and ranked words by their model-to-human occurrence ratio. Delve, tapestry and vibrant do not appear in the split at all, and meticulous appears once on each side. What sits at the top is precision and hedge vocabulary: comparatively at about 25 times the human rate in that split, modest at about 10, minimize at about 9, attributable at about 8, typically at about 8. That ranking is ours. No third party has published this profile, which is why the grade is community-observed and why this entry counts markers rather than banning any of them.
Why it reads as machine-made
The 2023 flourishes were trained down and something replaced them. What replaced them reads as careful rather than ornate, which makes it harder to notice and much more dangerous to flag. A directory built on the old word list is a museum piece, so this entry exists mainly to date-stamp the problem.
Specimens
Results were comparatively modest and largely attributable to seasonal variation, which we typically observe in this period. Performance across segments was not markedly different once the data had been standardized, and the residual variation is difficult to quantify with any confidence. Steps have been taken to minimize the effect going forward, though it should be acknowledged that the underlying drivers are only partially understood at this stage. Management remains cautiously optimistic, subject to the usual caveats around comparability. The picture remains broadly stable.
Signups fell 12 percent in July. They fell 14 percent last July. We think it is the summer holiday and we will know in September, when the comparison stops being a guess.
Figures in this repair are invented for the specimen. The repair swaps the hedge for the two numbers the hedge was standing in front of.
How it is detected
- Metric
- precision-hedge-markers-per-1000-words
- Threshold
- 4
- Direction
- above
- Threshold basis
- Ours, and provisional. The word ranking comes from our own extraction of the CC0 LexA English science split for GPT-5.2 on 2026-08-14; nobody has published a per-document threshold for this vocabulary, because nobody has published the profile at all. Four distinct markers per 1,000 words is a FeedSquad review trigger chosen so that one hedged sentence cannot fire it. The false-positive risk here is the worst in the directory, since this is also how careful quantitative writing sounds. Treat the number as a prompt to read the text, and recalibrate it against a corpus in the genre before treating it as anything else.
Who writes this way legitimately
Epidemiologists, statisticians, clinical trialists and anyone writing a results section produce this register by obligation, and the register is a virtue there. Attributable and stratify are technical terms with defined meanings in that context. Regulatory and pharmacovigilance writing is hedged because overclaiming is a compliance failure. This makes the cluster the most dangerous one in the directory to point at technical prose, and it should not be pointed at technical prose. The intended use is marketing copy, social posts and general commentary, where a stack of hedges usually means the writer had no measurement. A reviewer should check whether the numbers the hedges qualify appear anywhere; hedging around a stated figure is careful writing, and hedging around nothing is filler.
Model attribution
Extracted from the LexA science split for GPT-5.2 specifically. Treat it as a dated snapshot of one model in one register, not as a property of frontier models in general, and expect it to move.
Status history
| Date | Status | Rationale |
|---|---|---|
| 2026-08-15 | Active | Active, and new. The extraction on 2026-08-14 puts precision and hedge vocabulary at the top of the frontier science split while the 2023 markers are absent from it. Juzek documents post-2022 lexical uptake continuing across languages, and Geng and Trotta explain the substitution mechanism: publicised markers decay and unpublicised ones keep rising. |
Sources
- 01LexA-Index, CC0 per-language overuse datasetprimary-docaccessed 2026-08-14
- 02Juzek, LLM lexical uptake across 34 languages, preprint (arXiv:2605.25358)primary-docaccessed 2026-08-14
- 03Geng and Trotta, coevolution of human and LLM writing, Findings of ACL 2025peer-reviewedaccessed 2026-08-14
- 04FeedSquad extraction of the LexA GPT-5.2 English science splitfeedsquad-observedobserved 2026-08-14
CC0 LexA-Index release, English science register, GPT-5.2 split: 66,865 word rows with model and human counts, ranked by log-probability-ratio. Delve, delves, tapestry and vibrant absent from the split; meticulous present once on each side. Ranking performed by us and not published elsewhere.
CC BY 4.0 / The AI Tells Index, feedsquad.com/ai-tells