Get in Touch
AI content toolsfrom Opace

Measurement paper · 25,723 documents

The folk signals, measured

Burstiness is the most repeated way of spotting machine-written text. On 5,935 register-and-length-matched document pairs it scores AUROC 0.521, against 0.500 for a coin. Nineteen popular tells were measured the same way, and one property nobody names beats all of them.

On this page
  1. The finding
  2. How to read every number
  3. Burstiness
  4. The vocabulary tells
  5. Two signals that run backwards
  6. The em dash
  7. Formatting detects the interface
  8. Perplexity and its family
  9. What actually separates them
  10. Does it survive the source?
  11. The reversal at the end
  12. What this does not establish
Published
30 August 2026
Measured
Re-verified against source 30 August 2026
Method
Feature statistics · no detector, no flag point
Corpus
25,723 documents from six labelled pools

01 / The finding

The most-cited signal in AI detection is within a coin flip of chance

Sentence-length burstiness scores an AUROC of 0.521 on 5,935 register-and-length-matched document pairs. Chance is 0.500.

Cliff’s δ is −0.043: negligible, and pointing the wrong way. Machine prose is marginally more variable in sentence length than human prose, not less.
Burstiness, matched set0.521

AUROC over 5,935 matched pairs. A threshold strict enough to wrongly flag one human document in a hundred catches 2.5% of machine documents.

The strongest folk signal0.772

Em dashes per 1,000 words, on 670 fresh long-form pairs. It catches about one machine document in five, and four in five go past.

What actually separates them0.911

Vocabulary variety in a 100-word window. Machine prose repeats itself less than human prose, which is the opposite of the popular belief.

Most of the other folk signals do no better. What does separate the two populations is a property almost nobody names, and it runs against the received wisdom: machine prose repeats itself less than human prose does. A person introduces a term and keeps using it. A model introduces a term and reaches for a synonym.

02 / Reading the figures

How to read every number on this page

These are statistics over a corpus of documents. They are not detector outputs, so no threshold, no model file and no runtime applies to any figure below, and none is printed. Where a detection rate appears it is single-feature detection at a fixed false-positive budget: the threshold on that one measurement is set so that 1% of human documents are wrongly flagged, and the rate reported is the share of machine documents that clears it. That budget is stated with every such number.

AUROC
The probability that a randomly chosen machine document scores above a randomly chosen human one on that single measurement. 0.500 is a coin. 1.000 is perfect separation. It is threshold-free, which is why it flatters a signal that no usable threshold can be set on.
Detection at a 1% false-positive budget
What the same signal is worth once a threshold has to be chosen. Set that one measurement’s threshold so it wrongly flags 1% of human documents, then count how many machine documents it catches. A published tool has to operate at a budget like this, and it is where most of these signals collapse.
Cliff’s δ
Effect size, from −1 to +1. It carries a sign, which matters here: several of these signals are real and point the opposite way to the belief attached to them.

Measurement conditions

Corpus
25,723 documents: 10,890 machine, 14,833 human, from six labelled pools, de-duplicated on a normalised text hash. Minimum document length 60 words.
Runtime
No model and no runtime. Every figure is a statistic computed over text.
Measured
30 August 2026
Also
Two matched comparisons run throughout: 5,935 machine documents paired 1:1 with human documents on register family and log word count, and 670 pairs drawn only from the long-form corpus.

The corpus was assembled from six labelled pools on this project, unified onto one label vocabulary and de-duplicated on a normalised text hash, which removed 6,679 duplicates. The machine and human halves do not share a register mix, so any pooled statistic would partly measure register rather than authorship. Both matched comparisons are reported wherever they disagree.

This is distributional evidence. Every number describes what is typical of thousands of documents on each side. None of it says anything about a single piece of writing.

None of these measurements is a detector, and none should be used as one.

03 / Burstiness

Burstiness, in full

Burstiness of sentence length, measured on the matched set and on fresh long-form
MeasurementMatched set, 5,935 pairsFresh long-form, 670 pairs
AUROC0.5210.547
Cliff’s δ−0.043not reported
Detection at a 1% human false-positive budget2.5%0.1%

Sentence-length coefficient of variation, which is the same quantity computed a second way, returns the same three figures. Sentence-length standard deviation, sometimes described as perplexity’s partner, reaches AUROC 0.633 on the matched set and 1.9% detection at the same budget.

Burstiness was a reasonable read of 2022 output. Early instruction-tuned models did produce evenly metered sentences. Against a 2026 corpus the property has gone, and the metric built on it has gone with it. There is a second, stranger result about burstiness at the foot of this page.

Figure 1 Nineteen folk signals, ranked by how well they separate the two populations Bars are the matched set: 5,935 machine documents paired 1:1 with human documents on register family and log word count. Dots are the 670 fresh long-form pairs. The right-hand column is what each signal is worth once a threshold has to be set, at a 1% human false-positive budget. Nine of the nineteen sit below 0.54.
AUROC of nineteen popular AI-writing heuristics on the matched set and on fresh long-form 0.45 0.55 0.65 0.75 0.85 0.95 Detection @1% Hedging language 0.500 1.4% “AI phrases” 0.511 1.9% Closing with “in conclusion” 0.514 0.0% Rule of three 0.515 1.1% Discourse markers 0.517 1.8% Burstiness of sentence length 0.521 2.5% Sentence-length CV 0.521 2.5% Spectral flatness of rhythm 0.529 1.0% UID, between-sentence variance 0.533 1.2% Intensifiers 0.534 0.0% Compression ratio 0.561 2.3% “AI vocabulary” 0.578 6.6% Curly quotes 0.591 0.0% Em dashes per 1,000 words 0.595 7.2% Passive voice 0.601 0.0% UID, adjacent-word surprisal 0.619 1.6% Sentence-length SD 0.633 1.9% Any markdown present 0.660 0.0% MATTR, for reference not in the matched set 0.911 47.3% fresh chance, 0.500 Value axis starts at 0.45, not zero
Matched set, 5,935 pairs (bar) Fresh long-form, 670 pairs (dot)
AUROC and detection at a 1% false-positive budget for nineteen heuristics
HeuristicAUROC, matched setAUROC, fresh long-formDetection at a 1% budget
Hedging language 0.500 0.516 1.4%
“AI phrases” 0.511 0.533 1.9%
Closing with “in conclusion” 0.514 0.508 0.0%
Rule of three 0.515 0.579 1.1%
Discourse markers 0.517 0.570 1.8%
Burstiness of sentence length 0.521 0.547 2.5%
Sentence-length CV 0.521 0.547 2.5%
Spectral flatness of rhythm 0.529 0.571 1.0%
UID, between-sentence variance 0.533 0.510 1.2%
Intensifiers 0.534 0.541 0.0%
Compression ratio 0.561 0.734 2.3%
“AI vocabulary” 0.578 0.589 6.6%
Curly quotes 0.591 0.669 0.0%
Em dashes per 1,000 words 0.595 0.772 7.2%
Passive voice 0.601 0.716 0.0%
UID, adjacent-word surprisal 0.619 0.683 1.6%
Sentence-length SD 0.633 0.602 1.9%
Any markdown present 0.660 0.773 0.0%
MATTR, for reference not in the matched set 0.911 47.3% fresh

signal-science/tables/famous-heuristics.md, whole table; the MATTR reference row from SIGNAL-SCIENCE.md §2. Feature statistics over a corpus: no detector, no model file, no runtime and no operating point applies to any value here.

04 / Vocabulary

The vocabulary tells are the second casualty

The “AI words” list — delve, leverage, robust, tapestry, seamless — reaches AUROC 0.578 on the matched set and 6.6% detection at a 1% budget; on fresh long-form 0.589 and 5.5%. In the per-source sweep described further down it clears 0.65 on 5 of 33 pairings. Machine text does use these words somewhat more often. The gap is nowhere near wide enough to accuse a writer with.

Six further heuristics sit at or beside chance on the matched set, with detection at a 1% budget in brackets: “AI phrases” such as in today’s and dive into 0.511 (1.9%); discourse markers such as moreover and furthermore 0.517 (1.8%); the rule-of-three list 0.515 (1.1%); closing with in conclusion or ultimately 0.514 (0.0%); hedging language 0.500 (1.4%); intensifiers 0.534 (0.0%). Uniform information density, measured as between-sentence variance, reaches 0.533 (1.2%), and spectral flatness of the sentence-length rhythm 0.529 (1.0%).

05 / Signals with the wrong sign

Two signals that run backwards

Passive voice reaches AUROC 0.601 on the matched set with a Cliff’s δ of −0.202 and 0.0% detection at a 1% budget. The sign is the interesting part: machine prose uses less passive voice than human prose, so an editor hunting passives is hunting in the wrong direction.

Curly quotes and apostrophes reach 0.591, δ −0.182, again 0.0% detection at a 1% budget, and again people carry more of them. Published human prose has usually passed through a typesetting system that converts straight quotes. That is a fact about where a document was produced, not about who wrote it.

06 / The em dash

The em dash is real, and much weaker than its reputation

Median machine long-form document 3.50 em dashes per 1,000 words

On 670 fresh long-form pairs this gives AUROC 0.772 and 20.3% detection at a 1% false-positive budget — the strongest of the folk signals.

Median human long-form document Zero

The total absence in the human half is a property of published long-form rather than of human writing generally. On the wider matched set the same feature falls to 0.595 and 7.2%.

Catching about one machine document in five means four in five go past. The information is in presence against absence rather than in density, it varies heavily by provider, and it can be removed by find-and-replace.

07 / Formatting

Formatting detects the interface, not the author

Any markdown present reaches AUROC 0.773 on fresh long-form and 0.660 on the matched set, and markdown headings alone reach 36.1% detection at a 1% budget across the whole corpus, the highest single-feature figure in the study. It is also the one signal here we are confident is not about authorship at all. An article pasted out of a chat window carries hashes, asterisks and bullets; the same article pasted out of a CMS does not; a human-written README carries all three. What the feature separates is the interface a document travelled through.

A transparent 24-feature classifier forbidden from using any formatting feature beat the otherwise identical model that was allowed them, taking detection at a 1% budget from 62.5% to 72.1%.

793 machine and 4,179 human documents, none seen by either model. Withholding the strongest-looking signal made the model better. Source: SIGNAL-SCIENCE.md §3.

08 / Perplexity

Perplexity, and the family built on it

The zero-shot detection family assumes machine text is what a language model finds predictable. Nine published methods were reimplemented and run on 600 machine and 600 human fresh long-form documents, with GPT-2 small — 124M parameters, int8 with an fp16 head, 512-token cap — as the observer model.

Zero-shot detection methods reimplemented on a browser-deployable observer model
MethodAUROCDetection at a 1% budgetDetection at a 5% budget
Surprisal kurtosis (DivEye-inspired)0.76610.3%30.7%
Surprisal skew0.7630.0%27.7%
Surprisal autocorrelation0.7574.5%20.7%
GLTR, share of tokens in the observer’s top 1000.7350.0%18.0%
Mean log rank0.7280.0%19.0%
Log perplexity0.7150.0%15.8%
Fast-DetectGPT curvature0.54514.8%21.2%

600 machine and 600 human fresh long-form documents. Observer model GPT-2 small, 124M parameters, int8 with an fp16 head, 512-token cap. Sources: SIGNAL-SCIENCE.md §5.1 and tables/open-source-baselines.md.

Log perplexity detects nothing at a budget a published tool could operate at. The distributions overlap so heavily that the top human percentile sits above almost the whole machine distribution.

The founding assumption is also inverted. Median log perplexity is 3.68 for machine documents against 3.31 for human ones, with the direction holding in every register measured, on the same 600-a-side sample. Current models write with a vocabulary and phrasing that a 2019 observer finds less expected than human web prose.

Two rows that need protecting from misreading

Fast-DetectGPT’s own paper reports around 0.93 using far larger scoring models. The 0.545 here is a floor for the method under a browser-deployable observer, not a refutation of the published result. Binoculars was not implemented at all, because it requires two different models and only one is available offline here; the degenerate one-model proxy in the source table says nothing about it and its published results stand.

DivEye’s central claim — that the diversity of the surprisal sequence beats its mean — is confirmed on this corpus, with the three diversity moments at 0.766, 0.763 and 0.757 against 0.715 for the mean.

09 / The property that works

What actually separates the two populations

On the 670 fresh long-form pairs, eight of the top ten signals measure a single property.

The strongest single-feature signals on 670 fresh long-form pairs
RankSignalAUROCDetection at a 1% budgetMedian machineMedian human
1Content-word overlap between neighbouring sentences0.91223.1%0.0210.063
2Vocabulary variety in a 100-word window (MATTR)0.91147.3%0.7760.694
3Type-token ratio, first 400 words0.87617.0%0.5950.503
4Distinct word triples / word triples0.83916.4%0.9790.938
6Distinct word pairs / word pairs0.83019.7%0.8780.797
7Word-unigram entropy0.78811.8%8.287.78
8Share of vocabulary used exactly once0.7823.9%0.6750.618

670 machine and 670 human fresh long-form documents. Ranks 5, 9 and 10 measure the same property and are omitted for length; the full table is in SIGNAL-SCIENCE.md §2.

Figure 2 Adjacent-sentence content-word overlap: the median machine document against the median human one The share of a sentence's content words that also appear in the sentence beside it. The median human document shares three times as much as the median machine document. Only the two medians are plotted, because only the summary statistics are published; a distribution shape has not been invented to fill the space.
Median share of content words shared with the neighbouring sentence, machine against human 0.00 0.02 0.04 0.06 0.08 Machine long-form 670 documents 0.021 Human long-form 670 documents 0.063
Median adjacent-sentence content-word overlap by population
PopulationDocumentsMedian overlapAUROCDetection at a 1% budget
Machine long-form6700.0210.91223.1%
Human long-form6700.063

SIGNAL-SCIENCE.md §2, top-ten table. 670 machine and 670 human fresh long-form documents. AUROC 0.912; detection at a 1% human false-positive budget 23.1%. A feature statistic: no detector, no operating point.

A person introduces a term and keeps using it. A model introduces a term and reaches for a synonym, a pronoun, a rephrasing. That produces wider vocabulary per unit length, less overlap between neighbouring sentences, fewer repeated phrases and higher entropy, which are six instruments reading one behaviour.

MATTR, the strongest of them, catches 47.3% of machine long-form at a 1% human false-positive budget. It is the best single readable signal in the study, and it still misses more than half.

10 / Robustness

Does it survive the source?

The obvious objection is that the human long-form here is drawn from term-repeating technical sources, so the study may have discovered that its humans are scientists. Each human source was therefore tested separately against only the machine documents sharing its registers, and each provider likewise: 33 pairings.

Figure 3 Thirty-three source-and-provider pairings, from the median to the worst case Each row runs from that signal's worst pairing to its median pairing across 33 separate comparisons, so a signal that only works on one kind of human writing shows as a long line reaching down towards chance. The right-hand column counts the pairings clearing 0.65.
Median and worst-case AUROC across 33 source-and-provider pairings 0.45 0.60 0.75 0.90 1.00 Above 0.65 Vocabulary variety (MATTR) worst pairing: Meta models 0.562 0.791 29 / 33 Adjacent-sentence cohesion worst pairing: Meta models 0.561 0.777 31 / 33 Yule’s K worst pairing: not named in source 0.530 0.700 23 / 33 Repeated word-triple rate worst pairing: not named in source 0.527 0.698 23 / 33 Any markdown present worst pairing: Nvidia models 0.519 0.667 18 / 33 Em dashes per 1,000 words worst pairing: Nvidia models 0.501 0.637 12 / 33 Burstiness worst pairing: OpenAI models 0.500 0.597 11 / 33 “AI vocabulary” worst pairing: internet-archive texts 0.501 0.542 5 / 33 chance 0.65 Value axis starts at 0.45, not zero
Median of 33 pairings Worst single pairing
Show the numbers this was drawn from
Median and worst-case AUROC by signal across 33 source-and-provider pairings
SignalMedian AUROCWorstWorst pairingPairings above 0.65
Vocabulary variety (MATTR)0.7910.562Meta models29 / 33
Adjacent-sentence cohesion0.7770.561Meta models31 / 33
Yule’s K0.7000.530not named in source23 / 33
Repeated word-triple rate0.6980.527not named in source23 / 33
Any markdown present0.6670.519Nvidia models18 / 33
Em dashes per 1,000 words0.6370.501Nvidia models12 / 33
Burstiness0.5970.500OpenAI models11 / 33
“AI vocabulary”0.5420.501internet-archive texts5 / 33

SIGNAL-SCIENCE.md §2.3 and signal-science/tables/robustness.md. Each pairing tests one human source, or one provider's models, against only the documents sharing its registers. Per-pairing values are in the source table; the median and the worst case are plotted here because those are the published summary statistics.

The under-repetition signal holds against SEC 10-K filings (0.99), GOV.UK (0.90), PERSUADE student essays (0.89), Mongabay (0.87), Global Voices and Europe PMC (0.83), Common Crawl news (0.82) and C4 web text (0.78).

Its two weak points are recorded rather than buried. Meta’s models show it weakly, with MATTR at 0.562 and adjacent-sentence cohesion at 0.561, and Nvidia’s only moderately, at 0.633 and 0.694. By register it is weakest on creative writing (0.668) and social posts (0.658), and those are also the thinnest parts of the corpus at 251 and 2,132 machine documents against 260 and 1,119 human ones. A universal law is not what was measured.

11 / The reversal

Burstiness does predict something, and it is not what anyone wanted

Correlating the deployed model’s raw margin against each interpretable feature within the human population only — 4,184 documents, so the result is not just restating the label — burstiness comes out at Spearman ρ = −0.264, alongside discourse markers at +0.368 and mean word length at +0.359.

The metric that fails to find machines does help predict which people get accused.

Read that column as a description of which human writers are at risk of being wrongly flagged: formal, long-worded, evenly paced prose that signposts itself. The more evenly a person paces their sentences, the higher they score. Source: SIGNAL-SCIENCE.md §4.3, within 4,184 human documents.

12 / Limits

What this page does not establish

The human half of this corpus is not a random sample of human writing. It is what was obtainable under a clear licence: open-access science, government publications, filings, licensed news, web text, student essays. Modern trade non-fiction, professional journalism behind a paywall and personal writing of every kind are under-represented.

Register labels are machine-assigned and unvalidated for a substantial part of the corpus, and the register families were normalised by this project from vocabularies that differ between the upstream pools. Both are judgements, and the matched comparisons inherit their error.

The syntactic features are heuristics. Passive voice and proper nouns are approximated with regular expressions and are named _approx throughout; no part-of-speech tagger was used. The surprisal features in the main battery use a corpus unigram and bigram background model rather than a neural one, and the GPT-2 measurements in the perplexity section are on a 600-a-side subsample rather than the full corpus.

The fresh long-form set is not fully held out.

It is drawn from this project’s 5,558-document long-form corpus, and 268 of that corpus’s 922 machine documents — 29.1% — are also present in the cycle-2 training dataset, 168 of them in the train split. The human half is effectively clean at 11 of 4,636. For the feature statistics on this page that overlap is immaterial, because computing a type-token ratio or an em dash count involves no model and no training. It is not immaterial for the one figure here that compares a transparent classifier against the neural one, and the phrase “documents no model has seen” should not be repeated about this corpus without that qualification.

Source: corpus-reconciliation-2026-08-29/analysis.txt §2.

Nothing here is a shelf-stable law. The vocabulary tells that worked in 2024 are worthless now, and the under-repetition signal will be trained against once it is published. The methodology and the feature definitions are published so the measurement can be repeated rather than believed.

Provenance

Source file and section for every figure on this page
FigureFileSection
Corpus 25,723 = 10,890 + 14,833; 6,679 duplicates removed; 5,935 matched pairs; 670 fresh pairssignal-science/SIGNAL-SCIENCE.md§1
Burstiness 0.521, δ −0.043, 2.5%; fresh 0.547 / 0.1%; all folk-heuristic rowssignal-science/tables/famous-heuristics.mdwhole table
Em dash 3.50 against 0 per 1,000 words, 0.772, 20.3%SIGNAL-SCIENCE.md§2.1
Top-ten signals; 2.1% against 6.3% adjacent-sentence overlap; MATTR 0.911 / 47.3%SIGNAL-SCIENCE.md§2
Markdown 36.1% at a 1% budget; prose-only scorecard 72.1% against 62.5% on 793 / 4,179SIGNAL-SCIENCE.md§2.4, §3
33-pairing robustness table and per-source AUROCSIGNAL-SCIENCE.md, tables/robustness.md§2.3
Zero-shot baselines, 600 + 600, GPT-2 small int8 observerSIGNAL-SCIENCE.md, tables/open-source-baselines.md§5.1
Log perplexity 3.68 against 3.31SIGNAL-SCIENCE.md§5.2
Burstiness ρ −0.264 within 4,184 human documentsSIGNAL-SCIENCE.md§4.3
Limitations, human-corpus composition, machine-assigned labelsSIGNAL-SCIENCE.md§7
268/922 and 168 train-split overlap; human 11/4,636corpus-reconciliation-2026-08-29/analysis.txt§2

Every file named above ships with the open measurement repository, so any figure on this page can be re-run rather than taken on trust.

Read on

The signals are one half. What the shipped model actually keys on is the other.