On this page
01 / The finding
Most famous tells fail; the strong ones are about shape, not vocabulary
Ninety-eight phrases people quote as AI giveaways were measured on two independent AI/human pairings. Twenty-one survive. Thirteen point the wrong way — they are HUMAN markers — and the classics collapse on 2026 models.
13.2% of AI documents (506 of 3,839) keep their paragraphs within a CV of 0.2, against 0.8% of structured human ones (26 of 3,190). The strongest document-shape tell measured so far, and it holds on the hard negatives.
Humans put lists in 34% of sections against AI's 18%, and heavy bullet use fires on 15.7% of human documents against 4.9% of AI. On 2026 models the absence of lists leans machine, not their presence.
A static folk list would flag human writing and miss 2026 output.
02 / The lexicon
Ninety-eight famous phrases, two independent corpora, twenty-one survivors
Measured at a rule that no longer ships
- Corpus
- Two independent pairings: cycle-2 (5,655 AI / 9,859 human, register-balanced with equal weight per register) and generated-2026 (4,016 AI from 21 current models / 4,144 human web documents). A phrase ships only when elevated at least 2× on BOTH.
- Retired flag point
- 0.9855 / 0.9763
- Detector
- tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d (cycle 2, superseded 1 September 2026)
- Runtime
- Word-boundary token matching after normalising apostrophes and dashes — the same matcher the site runs.
- Measured
- 30 August 2026
- Also
- Rates are documents per 1,000 containing the phrase. A raw mined n-gram table was also built and DECLINED: its top rows were corpus-construction artefacts (source pairing, prompt leakage, British topic vocabulary), not authorship style.
The rule that ships today is margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562). The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.
| Phrase | AI /1,000 | Human /1,000 | Ratio |
|---|---|---|---|
| “robust” | 60.3 | 16.2 | 3.7× |
| “whether you're” | 43.8 | 8.7 | 5.0× |
| “seamless” | 22.4 | 5.8 | 3.8× |
| “here's the thing” | 2.2 | 0.2 | 6.5× |
| “paving the way” | 1.7 | 0.2 | 5.2× |
| “tapestry” | 0.75 | 0 | 7.2× — rare, small sample |
DOCUMENT-TELLS-2026-08-31.md §3b, known-phrases.json. AI side 4,016 documents, human side 4,144; the checker's phrase tell quotes these same rates with denominators.
03 / The anti-tells
Thirteen "AI phrases" are human markers, and the classics have aged out
The most shareable finding in the study is the failure list. Thirteen phrases from the folk lexicon fire more often in human writing on both pairings, and several 2023-era classics survive only in older-model text:
| Phrase | Measured verdict on 2026 output |
|---|---|
| “when it comes to” | 0.2× — a HUMAN marker |
| “a variety of” | 0.1× — human |
| “a plethora of” | 0.0× — human |
| “in this article” | 0.0× — human |
| “myriad” | 0.1× — human |
| “delve into” | 0.2× on 2026 models — dead |
| “in the realm of” | 0.1× on 2026 models — dead |
| “it's important to note” | 0.1× on 2026 models — dead |
| “moreover” | 0.6× on 2026 models |
| “additionally” | 0.4× on 2026 models |
Anyone still deciding authorship by spotting "delve" is reading the 2023 internet. The habit these phrases came from is real; the models moved, the lists did not — which is why every phrase this tool quotes carries its measured rates on current models, and why the list is re-measured rather than inherited.
The gate also applies to us. Seven candidates supplied from the founder's own live reading were put through the identical dual-corpus gate on 31 August 2026. One passed and ships — "at its core", 2.1× on the 2026 pairing (13 of 4,016 AI documents against 6 of 4,144 human) and 2.8× register-balanced. Six failed and do not ship, including his own strongest candidate: "in short" reads 0.9× register-balanced and 1.2× on 2026 models (45 v 39 documents) — at population level it is not a tell, however it reads in a single draft — and "simply put" (0.2×) and "in essence" (0.3×) run backwards on 2026 output.
04 / The shapes — and the flip
Half the shape verdicts flipped when the human baseline became fair, and both readings are published
The first pass measured the shape tells against the existing human corpus and refuted several — including the owner-favoured section-shape uniformity, which fired more on human docs. That refutation was itself an artefact: web scraping had stripped the structure from most human documents, and the few that kept it were disproportionately rigid, short pages. A new baseline was banked the same day — 3,529 licence-recorded human documents with headings, paragraphs and lists preserved, 2,513 parsing to three or more sections against 292 before — and the verdicts were re-taken on it:
| Tell | Against the stripped baseline | Against the structured baseline |
|---|---|---|
| Section-shape uniformity (mode share ≥ 0.8) | Refuted — human 13.0% > AI 8.5% | A real tell: human 2.9% (73/2,513) v AI 8.5% — ~2.9× |
| Composite scaffold (shape + sentence uniformity) | 2.8× (1.3× on hard negatives) | ~6.6×: AI 4.8% (111/2,332) v human 0.72% (18/2,513); hard negatives 0.78% |
| Sentence-length CV ≤ 0.3 | 2.4× (24.1% v 10.2%) | 3.6×: AI 31.0% (1,191/3,838) v human 8.6% (273/3,186) |
| Words-per-paragraph CV ≤ 0.2 | — | 16×: AI 13.2% (506/3,839) v human 0.8% (26/3,190) |
| Section lengths within 15% of median (≥ 90%) | — | ~14×: AI 10.5% (229/2,190) v human 0.76% (16/2,112) |
| Bullet-list rhythm | Unmeasurable — human lists lost upstream | An ANTI-tell: humans list more (34% of sections v 18%) |
| Formulaic closer ("Final Thoughts") | 2–2.7×, low coverage | Weakens to ~1.8× (2.4% v 1.3%) — demoted to colour |
| Keyphrase echo (SEO-style repetition) | Declined, ≈1.5× | An ANTI-tell: 9.4% of structured human docs v 5.1% of AI |
The corpus lesson matters more than any row. A shape tell measured against structure-stripped humans measures the scraper, not the writing — in either direction. The re-measurement flipped a refutation into a tell and two folk beliefs into anti-tells, and the study publishes both passes so the flip is visible rather than silently corrected. Also deliberately not shipped: sections-per-article, which separates in the table but is partly a corpus-length artefact.
05 / What ships
The evidence layer quotes only the survivors, with both rates every time
The checker's "Why it reads this way" card is built from this study and nothing else: the 21 surviving phrases, sentence-length rhythm (31.0% v 8.6%), the composite scaffold (4.8% v 0.72%), words-per-paragraph evenness (13.2% v 0.8%), section-length uniformity (10.5% v 0.76%) — plus the under-repetition and cadence tells from their own measurements. The weakened closer survives only as a footnote line, labelled too common in human writing to count.
Every one of those rates fires on some human writing — the strongest at under 1%, the broadest at 14% — which is why the card's register is fixed: patterns that illustrate the model's reading, stated with both rates and their denominators, never a verdict. The tells have no vote in the score, and the separation is enforced in code.
06 / Limits
What this page does not prove
In-distribution AI. The AI side is this programme's own generated corpora — its own prompts, 21 models, British-English briefs. Rates on other people's prompts and models will differ, and British topic vocabulary contaminates any naive phrase mining, which is why the mined table was declined outright.
The structured human corpus is professional, edited web writing. GOV.UK, developer documentation, editorial blogs, Wikinews — sources with recordable licences and pre-2022 provenance. Casual blogs and commercial listicles with clear licences remain unobtainable; FAQ and heavy-heading subsets stand in as hard negatives and behave consistently, but they are stand-ins. Register and source are also coupled — each register comes mostly from one source family — so register effects and source effects cannot be fully separated, and every headline claim was checked to hold direction in each register with n > 100.
Counts, not verdicts. No figure here involves the detector. These are corpus tendencies with denominators; people write every pattern on this page, and the interface says so wherever one is quoted.