On this page
01 / The finding
The most-cited signal in AI detection is within a coin flip of chance
Sentence-length burstiness scores an AUROC of 0.521 on 5,935 register-and-length-matched document pairs. Chance is 0.500.
AUROC over 5,935 matched pairs. A threshold strict enough to wrongly flag one human document in a hundred catches 2.5% of machine documents.
Em dashes per 1,000 words, on 670 fresh long-form pairs. It catches about one machine document in five, and four in five go past.
Vocabulary variety in a 100-word window. Machine prose repeats itself less than human prose, which is the opposite of the popular belief.
Most of the other folk signals do no better. What does separate the two populations is a property almost nobody names, and it runs against the received wisdom: machine prose repeats itself less than human prose does. A person introduces a term and keeps using it. A model introduces a term and reaches for a synonym.
02 / Reading the figures
How to read every number on this page
These are statistics over a corpus of documents. They are not detector outputs, so no threshold, no model file and no runtime applies to any figure below, and none is printed. Where a detection rate appears it is single-feature detection at a fixed false-positive budget: the threshold on that one measurement is set so that 1% of human documents are wrongly flagged, and the rate reported is the share of machine documents that clears it. That budget is stated with every such number.
- AUROC
- The probability that a randomly chosen machine document scores above a randomly chosen human one on that single measurement. 0.500 is a coin. 1.000 is perfect separation. It is threshold-free, which is why it flatters a signal that no usable threshold can be set on.
- Detection at a 1% false-positive budget
- What the same signal is worth once a threshold has to be chosen. Set that one measurement’s threshold so it wrongly flags 1% of human documents, then count how many machine documents it catches. A published tool has to operate at a budget like this, and it is where most of these signals collapse.
- Cliff’s δ
- Effect size, from −1 to +1. It carries a sign, which matters here: several of these signals are real and point the opposite way to the belief attached to them.
Measurement conditions
- Corpus
- 25,723 documents: 10,890 machine, 14,833 human, from six labelled pools, de-duplicated on a normalised text hash. Minimum document length 60 words.
- Runtime
- No model and no runtime. Every figure is a statistic computed over text.
- Measured
- 30 August 2026
- Also
- Two matched comparisons run throughout: 5,935 machine documents paired 1:1 with human documents on register family and log word count, and 670 pairs drawn only from the long-form corpus.
The corpus was assembled from six labelled pools on this project, unified onto one label vocabulary and de-duplicated on a normalised text hash, which removed 6,679 duplicates. The machine and human halves do not share a register mix, so any pooled statistic would partly measure register rather than authorship. Both matched comparisons are reported wherever they disagree.
This is distributional evidence. Every number describes what is typical of thousands of documents on each side. None of it says anything about a single piece of writing.
03 / Burstiness
Burstiness, in full
| Measurement | Matched set, 5,935 pairs | Fresh long-form, 670 pairs |
|---|---|---|
| AUROC | 0.521 | 0.547 |
| Cliff’s δ | −0.043 | not reported |
| Detection at a 1% human false-positive budget | 2.5% | 0.1% |
Sentence-length coefficient of variation, which is the same quantity computed a second way, returns the same three figures. Sentence-length standard deviation, sometimes described as perplexity’s partner, reaches AUROC 0.633 on the matched set and 1.9% detection at the same budget.
Burstiness was a reasonable read of 2022 output. Early instruction-tuned models did produce evenly metered sentences. Against a 2026 corpus the property has gone, and the metric built on it has gone with it. There is a second, stranger result about burstiness at the foot of this page.
| Heuristic | AUROC, matched set | AUROC, fresh long-form | Detection at a 1% budget |
|---|---|---|---|
| Hedging language | 0.500 | 0.516 | 1.4% |
| “AI phrases” | 0.511 | 0.533 | 1.9% |
| Closing with “in conclusion” | 0.514 | 0.508 | 0.0% |
| Rule of three | 0.515 | 0.579 | 1.1% |
| Discourse markers | 0.517 | 0.570 | 1.8% |
| Burstiness of sentence length | 0.521 | 0.547 | 2.5% |
| Sentence-length CV | 0.521 | 0.547 | 2.5% |
| Spectral flatness of rhythm | 0.529 | 0.571 | 1.0% |
| UID, between-sentence variance | 0.533 | 0.510 | 1.2% |
| Intensifiers | 0.534 | 0.541 | 0.0% |
| Compression ratio | 0.561 | 0.734 | 2.3% |
| “AI vocabulary” | 0.578 | 0.589 | 6.6% |
| Curly quotes | 0.591 | 0.669 | 0.0% |
| Em dashes per 1,000 words | 0.595 | 0.772 | 7.2% |
| Passive voice | 0.601 | 0.716 | 0.0% |
| UID, adjacent-word surprisal | 0.619 | 0.683 | 1.6% |
| Sentence-length SD | 0.633 | 0.602 | 1.9% |
| Any markdown present | 0.660 | 0.773 | 0.0% |
| MATTR, for reference | not in the matched set | 0.911 | 47.3% fresh |
signal-science/tables/famous-heuristics.md, whole table; the MATTR reference row from SIGNAL-SCIENCE.md §2. Feature statistics over a corpus: no detector, no model file, no runtime and no operating point applies to any value here.
04 / Vocabulary
The vocabulary tells are the second casualty
The “AI words” list — delve, leverage, robust, tapestry, seamless — reaches AUROC 0.578 on the matched set and 6.6% detection at a 1% budget; on fresh long-form 0.589 and 5.5%. In the per-source sweep described further down it clears 0.65 on 5 of 33 pairings. Machine text does use these words somewhat more often. The gap is nowhere near wide enough to accuse a writer with.
Six further heuristics sit at or beside chance on the matched set, with detection at a 1% budget in brackets: “AI phrases” such as in today’s and dive into 0.511 (1.9%); discourse markers such as moreover and furthermore 0.517 (1.8%); the rule-of-three list 0.515 (1.1%); closing with in conclusion or ultimately 0.514 (0.0%); hedging language 0.500 (1.4%); intensifiers 0.534 (0.0%). Uniform information density, measured as between-sentence variance, reaches 0.533 (1.2%), and spectral flatness of the sentence-length rhythm 0.529 (1.0%).
05 / Signals with the wrong sign
Two signals that run backwards
Passive voice reaches AUROC 0.601 on the matched set with a Cliff’s δ of −0.202 and 0.0% detection at a 1% budget. The sign is the interesting part: machine prose uses less passive voice than human prose, so an editor hunting passives is hunting in the wrong direction.
Curly quotes and apostrophes reach 0.591, δ −0.182, again 0.0% detection at a 1% budget, and again people carry more of them. Published human prose has usually passed through a typesetting system that converts straight quotes. That is a fact about where a document was produced, not about who wrote it.
06 / The em dash
The em dash is real, and much weaker than its reputation
On 670 fresh long-form pairs this gives AUROC 0.772 and 20.3% detection at a 1% false-positive budget — the strongest of the folk signals.
The total absence in the human half is a property of published long-form rather than of human writing generally. On the wider matched set the same feature falls to 0.595 and 7.2%.
Catching about one machine document in five means four in five go past. The information is in presence against absence rather than in density, it varies heavily by provider, and it can be removed by find-and-replace.
07 / Formatting
Formatting detects the interface, not the author
Any markdown present reaches AUROC 0.773 on fresh long-form and 0.660 on the matched set, and markdown headings alone reach 36.1% detection at a 1% budget across the whole corpus, the highest single-feature figure in the study. It is also the one signal here we are confident is not about authorship at all. An article pasted out of a chat window carries hashes, asterisks and bullets; the same article pasted out of a CMS does not; a human-written README carries all three. What the feature separates is the interface a document travelled through.
A transparent 24-feature classifier forbidden from using any formatting feature beat the otherwise identical model that was allowed them, taking detection at a 1% budget from 62.5% to 72.1%.
08 / Perplexity
Perplexity, and the family built on it
The zero-shot detection family assumes machine text is what a language model finds predictable. Nine published methods were reimplemented and run on 600 machine and 600 human fresh long-form documents, with GPT-2 small — 124M parameters, int8 with an fp16 head, 512-token cap — as the observer model.
| Method | AUROC | Detection at a 1% budget | Detection at a 5% budget |
|---|---|---|---|
| Surprisal kurtosis (DivEye-inspired) | 0.766 | 10.3% | 30.7% |
| Surprisal skew | 0.763 | 0.0% | 27.7% |
| Surprisal autocorrelation | 0.757 | 4.5% | 20.7% |
| GLTR, share of tokens in the observer’s top 100 | 0.735 | 0.0% | 18.0% |
| Mean log rank | 0.728 | 0.0% | 19.0% |
| Log perplexity | 0.715 | 0.0% | 15.8% |
| Fast-DetectGPT curvature | 0.545 | 14.8% | 21.2% |
600 machine and 600 human fresh long-form documents. Observer model GPT-2 small, 124M parameters, int8 with an fp16 head, 512-token cap. Sources: SIGNAL-SCIENCE.md §5.1 and tables/open-source-baselines.md.
Log perplexity detects nothing at a budget a published tool could operate at. The distributions overlap so heavily that the top human percentile sits above almost the whole machine distribution.
The founding assumption is also inverted. Median log perplexity is 3.68 for machine documents against 3.31 for human ones, with the direction holding in every register measured, on the same 600-a-side sample. Current models write with a vocabulary and phrasing that a 2019 observer finds less expected than human web prose.
Two rows that need protecting from misreading
Fast-DetectGPT’s own paper reports around 0.93 using far larger scoring models. The 0.545 here is a floor for the method under a browser-deployable observer, not a refutation of the published result. Binoculars was not implemented at all, because it requires two different models and only one is available offline here; the degenerate one-model proxy in the source table says nothing about it and its published results stand.
DivEye’s central claim — that the diversity of the surprisal sequence beats its mean — is confirmed on this corpus, with the three diversity moments at 0.766, 0.763 and 0.757 against 0.715 for the mean.
09 / The property that works
What actually separates the two populations
On the 670 fresh long-form pairs, eight of the top ten signals measure a single property.
| Rank | Signal | AUROC | Detection at a 1% budget | Median machine | Median human |
|---|---|---|---|---|---|
| 1 | Content-word overlap between neighbouring sentences | 0.912 | 23.1% | 0.021 | 0.063 |
| 2 | Vocabulary variety in a 100-word window (MATTR) | 0.911 | 47.3% | 0.776 | 0.694 |
| 3 | Type-token ratio, first 400 words | 0.876 | 17.0% | 0.595 | 0.503 |
| 4 | Distinct word triples / word triples | 0.839 | 16.4% | 0.979 | 0.938 |
| 6 | Distinct word pairs / word pairs | 0.830 | 19.7% | 0.878 | 0.797 |
| 7 | Word-unigram entropy | 0.788 | 11.8% | 8.28 | 7.78 |
| 8 | Share of vocabulary used exactly once | 0.782 | 3.9% | 0.675 | 0.618 |
670 machine and 670 human fresh long-form documents. Ranks 5, 9 and 10 measure the same property and are omitted for length; the full table is in SIGNAL-SCIENCE.md §2.
| Population | Documents | Median overlap | AUROC | Detection at a 1% budget |
|---|---|---|---|---|
| Machine long-form | 670 | 0.021 | 0.912 | 23.1% |
| Human long-form | 670 | 0.063 |
SIGNAL-SCIENCE.md §2, top-ten table. 670 machine and 670 human fresh long-form documents. AUROC 0.912; detection at a 1% human false-positive budget 23.1%. A feature statistic: no detector, no operating point.
A person introduces a term and keeps using it. A model introduces a term and reaches for a synonym, a pronoun, a rephrasing. That produces wider vocabulary per unit length, less overlap between neighbouring sentences, fewer repeated phrases and higher entropy, which are six instruments reading one behaviour.
MATTR, the strongest of them, catches 47.3% of machine long-form at a 1% human false-positive budget. It is the best single readable signal in the study, and it still misses more than half.
10 / Robustness
Does it survive the source?
The obvious objection is that the human long-form here is drawn from term-repeating technical sources, so the study may have discovered that its humans are scientists. Each human source was therefore tested separately against only the machine documents sharing its registers, and each provider likewise: 33 pairings.
Show the numbers this was drawn from
| Signal | Median AUROC | Worst | Worst pairing | Pairings above 0.65 |
|---|---|---|---|---|
| Vocabulary variety (MATTR) | 0.791 | 0.562 | Meta models | 29 / 33 |
| Adjacent-sentence cohesion | 0.777 | 0.561 | Meta models | 31 / 33 |
| Yule’s K | 0.700 | 0.530 | not named in source | 23 / 33 |
| Repeated word-triple rate | 0.698 | 0.527 | not named in source | 23 / 33 |
| Any markdown present | 0.667 | 0.519 | Nvidia models | 18 / 33 |
| Em dashes per 1,000 words | 0.637 | 0.501 | Nvidia models | 12 / 33 |
| Burstiness | 0.597 | 0.500 | OpenAI models | 11 / 33 |
| “AI vocabulary” | 0.542 | 0.501 | internet-archive texts | 5 / 33 |
SIGNAL-SCIENCE.md §2.3 and signal-science/tables/robustness.md. Each pairing tests one human source, or one provider's models, against only the documents sharing its registers. Per-pairing values are in the source table; the median and the worst case are plotted here because those are the published summary statistics.
The under-repetition signal holds against SEC 10-K filings (0.99), GOV.UK (0.90), PERSUADE student essays (0.89), Mongabay (0.87), Global Voices and Europe PMC (0.83), Common Crawl news (0.82) and C4 web text (0.78).
Its two weak points are recorded rather than buried. Meta’s models show it weakly, with MATTR at 0.562 and adjacent-sentence cohesion at 0.561, and Nvidia’s only moderately, at 0.633 and 0.694. By register it is weakest on creative writing (0.668) and social posts (0.658), and those are also the thinnest parts of the corpus at 251 and 2,132 machine documents against 260 and 1,119 human ones. A universal law is not what was measured.
11 / The reversal
Burstiness does predict something, and it is not what anyone wanted
Correlating the deployed model’s raw margin against each interpretable feature within the human population only — 4,184 documents, so the result is not just restating the label — burstiness comes out at Spearman ρ = −0.264, alongside discourse markers at +0.368 and mean word length at +0.359.
The metric that fails to find machines does help predict which people get accused.
12 / Limits
What this page does not establish
The human half of this corpus is not a random sample of human writing. It is what was obtainable under a clear licence: open-access science, government publications, filings, licensed news, web text, student essays. Modern trade non-fiction, professional journalism behind a paywall and personal writing of every kind are under-represented.
Register labels are machine-assigned and unvalidated for a substantial part of the corpus, and the register families were normalised by this project from vocabularies that differ between the upstream pools. Both are judgements, and the matched comparisons inherit their error.
The syntactic features are heuristics. Passive voice and proper nouns are approximated with regular expressions and are named _approx throughout; no part-of-speech tagger was used. The surprisal features in the main battery use a corpus unigram and bigram background model rather than a neural one, and the GPT-2 measurements in the perplexity section are on a 600-a-side subsample rather than the full corpus.
It is drawn from this project’s 5,558-document long-form corpus, and 268 of that corpus’s 922 machine documents — 29.1% — are also present in the cycle-2 training dataset, 168 of them in the train split. The human half is effectively clean at 11 of 4,636. For the feature statistics on this page that overlap is immaterial, because computing a type-token ratio or an em dash count involves no model and no training. It is not immaterial for the one figure here that compares a transparent classifier against the neural one, and the phrase “documents no model has seen” should not be repeated about this corpus without that qualification.
Source: corpus-reconciliation-2026-08-29/analysis.txt §2.
Nothing here is a shelf-stable law. The vocabulary tells that worked in 2024 are worthless now, and the under-repetition signal will be trained against once it is published. The methodology and the feature definitions are published so the measurement can be repeated rather than believed.
Provenance
| Figure | File | Section |
|---|---|---|
| Corpus 25,723 = 10,890 + 14,833; 6,679 duplicates removed; 5,935 matched pairs; 670 fresh pairs | signal-science/SIGNAL-SCIENCE.md | §1 |
| Burstiness 0.521, δ −0.043, 2.5%; fresh 0.547 / 0.1%; all folk-heuristic rows | signal-science/tables/famous-heuristics.md | whole table |
| Em dash 3.50 against 0 per 1,000 words, 0.772, 20.3% | SIGNAL-SCIENCE.md | §2.1 |
| Top-ten signals; 2.1% against 6.3% adjacent-sentence overlap; MATTR 0.911 / 47.3% | SIGNAL-SCIENCE.md | §2 |
| Markdown 36.1% at a 1% budget; prose-only scorecard 72.1% against 62.5% on 793 / 4,179 | SIGNAL-SCIENCE.md | §2.4, §3 |
| 33-pairing robustness table and per-source AUROC | SIGNAL-SCIENCE.md, tables/robustness.md | §2.3 |
| Zero-shot baselines, 600 + 600, GPT-2 small int8 observer | SIGNAL-SCIENCE.md, tables/open-source-baselines.md | §5.1 |
| Log perplexity 3.68 against 3.31 | SIGNAL-SCIENCE.md | §5.2 |
| Burstiness ρ −0.264 within 4,184 human documents | SIGNAL-SCIENCE.md | §4.3 |
| Limitations, human-corpus composition, machine-assigned labels | SIGNAL-SCIENCE.md | §7 |
| 268/922 and 168 train-split overlap; human 11/4,636 | corpus-reconciliation-2026-08-29/analysis.txt | §2 |
Every file named above ships with the open measurement repository, so any figure on this page can be re-run rather than taken on trust.