On this page
01 / The finding
Honest phrase evidence is small, interval-shaped, and one judgement call from wrong
30.1× on the mining half — because our AI corpus was generated in British English and the human half is largely American. 29.9% of AI documents contain "per cent" against 1.4% of human ones. A spelling detector wearing an AI detector's label, caught before shipping and published here.
Three-, four- and five-word phrases appearing in at least five documents on both sides. That is why nothing longer than three words is published: 21 five-word phrases is not a table. A commercial panel shows longer phrases because it indexes millions of documents; ours holds thousands and says so.
Ranked on the mining half, reported on the held-out half: log-ratio correlation 0.73, median held-out ratio 10.3×, and not one phrase reversed. The effect is real. What it is not is precise — intervals commonly span a factor of five.
02 / The near miss
The first table was a British-English detector, and it is recorded because it nearly shipped
Built on raw text, the three strongest "AI phrases" in the corpus were "per cent of" (30.1×), "per cent in" (28.3×) and "cent of the" (28.0×) — the same artefact three times. Our AI half writes "per cent" and "-our" spellings (47.2% of AI documents against 11.6% of human); our human half is largely American. The fix normalises spelling to one convention on both sides before anything is counted, and the "per cent" family disappears from the ranking entirely. The same failure is visible in the competitor's own published panel, where a domain phrase from the user's document is ranked at 307× — quotable, specific, and measuring the topic rather than the author.
03 / The rules
Document frequency, a count floor on both sides, and a register control
Measured at a rule that no longer ships
- Corpus
- Mining half 467 AI / 2,315 human; held-out half 455 AI / 2,321 human, assigned by id hash. Phrases are ranked and register-controlled on the mining half and REPORTED on the held-out half — selecting on the numbers you then publish is how a table ends up describing its own noise.
- Retired flag point
- 0.9855 / 0.9763
- Detector
- tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d (cycle 2, superseded 1 September 2026)
- Runtime
- Pure counting after spelling normalisation. Document frequency, not raw frequency: a phrase used eleven times in one verbose white paper is one document's worth of evidence.
- Measured
- 30 August 2026
- Also
- A phrase must appear in at least five documents on each side of the mining half, must lean AI inside every register with enough documents to test (19 of the top 40 do; the 21 that fail do so for want of power, none reversed), and its held-out interval must exclude 1.
The rule that ships today is margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562). The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.
04 / Replication
The finding under the table: machine prose argues by explicit contrast
Most of the surviving phrases are contrastive or hedging frames — "rather than a", "is not simply", "not the same", "the gap between", "the limits of". Set beside the repetition finding — machine writing shares only 2.1% of content words between neighbouring sentences against 6.3% for human prose — the two are one description of how the prose is built: machine prose argues by explicit contrast where people repeat and accumulate. The house writing guidelines banned "it's not just X, it's Y" on taste; 922 documents arrived at the same judgement by counting.
05 / The table
The eighteen rows, exactly as the checker serves them
This table is read from the same JSON file the checker fetches — phrase-ratios-v1, measured 2026-08-30 — not retyped. Every ratio is the held-out 95% interval; every row carries its document counts, and the weakest rows sit on nine or ten documents, visibly.
| Phrase | Held-out ratio, 95% interval | Registers tested | AI docs | Human docs |
|---|---|---|---|---|
the literature is | 7.1–55.5× | 3 | 17 | 4 |
the evidence base | 19.9–174.8× | 4 | 40 | 3 |
the result is | 10.6–65.4× | 4 | 28 | 5 |
rather than an | 5.2–19.1× | 3 | 26 | 13 |
not the same | 7.8–50.5× | 6 | 21 | 5 |
but as a | 6–35.2× | 5 | 18 | 6 |
is not simply | 11.9–161.6× | 3 | 21 | 2 |
rather than as | 10.2–75.4× | 4 | 24 | 4 |
question of whether | 4.1–14.5× | 3 | 23 | 15 |
the gap between | 5.2–18.1× | 7 | 27 | 14 |
a fraction of | 2.6–15.5× | 3 | 10 | 8 |
a mixture of | 3.5–23.4× | 4 | 11 | 6 |
evidence from the | 3.9–29.3× | 3 | 11 | 5 |
the first is | 3.9–13.5× | 5 | 23 | 16 |
rather than a | 8.5–17× | 7 | 100 | 42 |
should therefore be | 5.6–59.2× | 3 | 12 | 3 |
the limits of | 5.1–23.6× | 3 | 20 | 9 |
ways that are | 6.2–49.9× | 3 | 15 | 4 |
06 / The one that got through
"the bank of" passed every automatic filter, and was cut by judgement
One phrase — "the bank of", 2.5–16.7× held out, on 9 AI and 7 human documents — cleared the count floor, the register control and the held-out interval test, and is still not a machine tell. Our AI half writes about banking across several registers, so a control that compares a phrase's lean between registers cannot separate subject from style when the subject is everywhere. It was excluded on credibility, not statistics: a reader pasting a banking article who sees their own subject marked as an AI phrase will conclude the tool is not thinking, and on that row they would be right.
The exclusion is not a quiet edit. It lives in the table-building script as a named entry with its reason, and is emitted into the shipped JSON under excluded_by_judgement, so it is reproducible and cannot be mistaken for tuning the table until it read well. What it means for the other rows is also stated: at least one topical artefact reached the final filter, others may sit below the point where they were obvious, and the counts are printed against every row so a reader can see exactly how thin each one is.
07 / Limits
What travels with the panel, wherever it renders
A phrase in your draft is not evidence about your draft. These are tendencies across thousands of documents; people write every one of them, and the panel leads with that sentence before any number.
Small corpus, wide intervals, long-form only. 922 AI documents against a commercial service's millions is why the table stops at three words and eighteen rows. Short marketing, SEO and social copy are not represented and cannot be — every sample this programme owns of those registers sits inside the training set. 268 of the 922 AI documents touch a cycle-2 split; no classifier is involved in a phrase count, but the documents are not guaranteed independent of the model either, and the caveat ships.