Methodology
How entries get in, how they get graded, how they get retired, and how to argue with any of it. Written by Ville Ylläsjärvi, who maintains the index.
What a tell is
A tell is a pattern that correlates with low-effort writing. That is the whole claim. Correlation measured across a corpus says very little about one document and nothing at all about who typed it. Every entry here has a human population that produces the shape deliberately, which is why the false-positive field is mandatory and why the build gate rejects a short one. Use the index to improve a draft or to explain why a draft reads flat. Do not use it to accuse a person.
The evidence standard
Every entry carries its sources and one evidence grade. Sources are tiered by what they are: peer-reviewed study, primary documentation, vendor material, press, or community observation. Grades run from peer-reviewed, where a published study measures the pattern, down through primary-doc, corroborated and community-observed, and end at feedsquad-observed, which means our own corpus and no external attestation. The schema refuses to parse an entry graded above feedsquad-observed unless it cites at least one external source. The rule behind all of it is short. No source, no claim. An unattributed observation gets the lowest grade and says so on its own page rather than borrowing authority from a citation that does not support it.
How an entry is added
A candidate needs a source, a constructed specimen, a repair, a category, a severity and a detection rule. Then the build gate tries to break it. The pattern has to compile, it has to match its own specimen, the repaired version has to stop matching, and no other rule in the dataset may fire on the same text. Two rules firing on one specimen means the two entries were always one entry, so one of them gets merged away. The id is assigned at that point and fixed forever, and three ids are reserved because static routes claim them first: changelog, methodology and feed. Nothing is hand-maintained downstream: the changelog, the feed and the API all read the entries themselves.
Burned, contested, and why retired signal stays
Burned means the signal has collapsed. Either the pattern was publicised so widely that human writers now trip it constantly, or the model vendors trained the habit out. Contested means credible people dispute that the pattern signals anything at all. Both keep their place in the dataset with their rules de-armed, because a signal that stopped working is still information, and because downstream tools pin entry ids. Deleting an entry would break those tools in silence. Retirement here means a status change plus a dated history event. Ids are permanent and never get reused.
Thresholds
A statistical entry may carry a threshold only alongside a stated basis for it, and the schema enforces the pair. Most of these numbers are review triggers rather than published cutoffs. A number with no stated basis is exactly the fabricated precision this index catalogues, so one does not ship. Where a study measures at population scale and publishes no per-document figure, the entry says that plainly and labels its own working default as a default. The lint engine that ships with the index computes two metrics: em dashes per 1,000 words, and the coefficient of variation of sentence length. Every other statistical entry names a metric the engine does not compute, which makes it a review trigger a caller has to implement before it measures anything. The engine lists those entries back to the caller rather than dropping them, so an empty result never passes for a clean statistical pass. Recalibrate against your own corpus before treating any threshold here as a verdict.
Who writes this way legitimately
False positives get their own required field because they are the product. Liang and colleagues measured an average false-positive rate of 61.3 percent across seven detectors on human-written TOEFL essays. What those detectors were reacting to was plain vocabulary and even sentence length, which in practice means they were reacting to the writer having learned English second. Academic register, technical documentation, translated business English, legal drafting and plain formal writing all generate shapes listed in this index. So does any writer who has read a great deal of the same material the models read. Naming those populations by hand, per entry, is most of the work.
Corrections
Email hello@feedsquad.com with the entry id, what is wrong with it, and a source. Disagreement about status counts: if a tell has burned and we have not caught up, that is the report worth sending. Accepted changes land as a dated status event with a rationale, which the changelog then publishes. The record of having been wrong stays visible on the entry.
Versioning
The dataset carries a semantic version, served in the API envelope and pinnable by any consumer. Major means the envelope or a field contract broke, which we do not expect, since ids are permanent and retirement happens through status. Minor covers entries added and statuses changed. Patch covers copy edits that touch no id, no status and no detection rule. Ids never change. Pin one and rely on it.
How to cite
Reuse is free under CC BY 4.0, including commercial reuse, provided the attribution points at the index itself. Copy the line below and set the accessed date to the day you read it.
FeedSquad. The AI Tells Index. feedsquad.com/ai-tells. Accessed YYYY-MM-DD. Licensed CC BY 4.0.