Multilingual lexical uptake
- Id
- multilingual-lexical-uptake
- Status
- Active
- Severity
- medium
- Detection
- statistical
- Evidence grade
- corroborated
- Languages
- en, af, ar, bg, cs, de, el, es, et, fa, fi, fr, hi, hr, id, is, it, ja, kk, ko, ky, lt, lv, mr, nl, pl, pt, ro, ru, sr, ta, tr, uk, zh
- Added
- 2026-08-15
- Updated
- 2026-08-17
Currently signals low-effort writing.
What it is
The post-2022 lexical shift is not an English problem. Juzek tested 34 languages and found post-ChatGPT lexical uptake in 26 of them, with a mean prevalence increase of 15.1 percent. The companion CC0 dataset covers those languages in a news register. The English word lists that circulate do not transfer: a Finnish or German text does not become suspect because it contains a translation of delve, and no per-language marker list is published that we could responsibly reprint. This entry exists to record the finding and to stop the English list being applied where it does not belong.
Why it reads as machine-made
The same pull produces the same shape in any language: a verb of depth, a noun of significance, and no date anywhere in the paragraph. What changes across languages is the vocabulary, not the emptiness.
Specimens
Tämä artikkeli syventyy etätyön monimutkaiseen kokonaisuuteen ja korostaa selkeän viestinnän merkitystä. Luottamus rakentuu hitaasti ja katoaa nopeasti, minkä useimmat tiimit oppivat vasta jälkikäteen. Toimivat käytännöt syntyvät harvoin kerralla, vaan ne muotoutuvat matkan varrella. Lopulta ratkaisee se, miten arjessa oikeasti toimitaan, ei se, mitä ohjeeseen on kirjoitettu.
Siirryimme etätyöhön maaliskuussa. Palaverit vähenivät neljästä kahteen viikossa, ja kirjoitettu muistio korvasi loput. Muistio on nyt ensimmäinen asia, jonka uusi työntekijä lukee.
The specimen is Finnish. Figures in this repair are invented for the specimen. The before text makes the same move as the English version: a verb of depth and a noun of significance where a month and a number would do.
Dieser Beitrag beleuchtet die vielschichtigen Herausforderungen der Fernarbeit und unterstreicht die Bedeutung klarer Kommunikation. Vertrauen entsteht langsam und geht schnell verloren, was die meisten Teams erst im Nachhinein bemerken. Gute Gewohnheiten lassen sich nicht verordnen, sie wachsen mit der Zeit. Am Ende zählt, wie im Alltag tatsächlich gearbeitet wird, und nicht, was in einem Leitfaden steht.
Wir haben im März auf Fernarbeit umgestellt. Aus vier Meetings pro Woche wurden zwei, den Rest übernimmt ein schriftliches Protokoll.
The specimen is German. Figures in this repair are invented for the specimen.
How it is detected
- Metric
- distinct-per-language-uptake-markers-per-1000-words
- Threshold
- 3
- Direction
- above
- Threshold basis
- Carried over from the English density trigger, and weaker here. Juzek reports 26 of 34 languages showing post-2022 uptake with a mean prevalence increase of 15.1 percent, all at corpus level, and publishes no per-document cutoff and no per-language marker list. The metric is named separately from the English one because the marker list differs by language and the two numbers are not comparable. Three distinct markers per 1,000 words is therefore a FeedSquad review trigger with no per-language calibration behind it at all. Anyone applying this outside English must build the marker list from a measured corpus in that language first. Running the English list through a translator produces nothing worth reading.
Who writes this way legitimately
Translators and staff at multilingual institutions write a formal register in every language they publish in, and that register is a large part of what the measurement picks up. European Union and United Nations texts are drafted to be translatable, which flattens exactly the features a marker list keys on. Second-language writers in any target language produce the most predictable register available to them, which is the same mechanism behind the 61.3 percent average false-positive rate Liang and colleagues measured across seven detectors on non-native English essays. Languages with rich morphology break naive word matching outright, so a Finnish marker count that ignores inflection measures tokenisation rather than writing. A reviewer working outside English should treat this entry as a warning about method, not as a detector.
Model attribution
Corpus-level and cross-model. The LexA news data was generated with a small 2025-era model and the science data with several families, so no single vendor sits behind the finding.
Status history
| Date | Status | Rationale |
|---|---|---|
| 2026-08-15 | Active | Active. Juzek measures uptake in 26 of 34 languages and the companion dataset covers them; Liang and colleagues find the same shift in consumer complaints, corporate communications and United Nations press releases at population scale. The source is a preprint committed to EMNLP 2026 and is labelled as such on the entry. |
Sources
- 01Juzek, LLM lexical uptake across 34 languages, preprint (arXiv:2605.25358)primary-docaccessed 2026-08-14
- 02LexA-Index, CC0 per-language overuse datasetprimary-docaccessed 2026-08-14
- 03Kobak et al., Delving into LLM-assisted writing in biomedical publications, Science Advances 11(27)peer-reviewedaccessed 2026-08-14
- 04Liang et al., Widespread LLM adoption in society, Patterns (arXiv:2502.09747)peer-reviewedaccessed 2026-08-14
- 05
CC BY 4.0 / The AI Tells Index, feedsquad.com/ai-tells