On this page
01 / The finding
The cliff is below 200 words, and the headline rate is a long-document rate
How much text you paste in changes the answer more than which model wrote it, what register it is in, or anything else we have measured.
29 of 172 machine documents, 95% CI 12.0–23.2. Five machine passages in six go past at this length.
883 of 922 machine documents, 95% CI 94.3–96.9. That corpus has a median machine document of 1,612 words.
From 16.9% at 100–199 words to 84.6% at 300–399. Nothing else measured on this project moves the answer that far.
The headline figure is a long-document figure and it should be read as one. Somebody checking an 800-word blog post is not operating at 95.8%, and the interface should say so where the text is pasted rather than in a footnote further down the page.
02 / Conditions
The measurement conditions, which apply to every figure on this page
Measurement conditions
- Corpus
- Two corpora, never pooled into one cell. Short-form: 816 machine passages generated 29 August 2026 from openai/gpt-5.6-sol and openai/gpt-5.6-luna across four target lengths and three prompt styles, with 4,368 human passages from nine sources. Long-form: the 5,558-document set of 28 August 2026, being 922 machine documents across 13 models and 4,636 human documents from eight sources.
- Operating point
- 0.9855 / 0.9763 · contract segments-v3 · T = 0.8324
- Detector
tier3-cycle2-e5small-fp32.onnx, SHA-256e313ab00de1fffd2…4d2788d- Runtime
- Python onnxruntime 1.29.0, CPU, fp32
- Measured
- 30 August 2026
- Also
- segments-v3 token-bounded segmentation. The reported document probability is the maximum across sections.
The detector is the fp32 parent of the shipped browser artefact and the file the EU inference server runs. Every figure below came through Python onnxruntime on CPU, which is the reference-server scoring path and not a browser measurement. Where an older figure is quoted for contrast, its own operating point is named beside it.
The two corpora are never pooled, because they were generated differently, by different models, in different registers. Bands were chosen from the corpora’s own word-count distributions rather than from round numbers, and where both corpora cover a band each keeps its own cell. Any cell holding fewer than 30 documents prints its count and no rate, which is the floor the source file applies throughout: a 21 of 23 is not “91.3%” in any sense that should travel.
03 / The curve
The curve, and the floor under it
| Words | Corpus | Machine flagged | Detection rate [95% CI] |
|---|---|---|---|
| under 100 | short-form | 6/27 | n below 30, no rate quoted |
| 100–199 | short-form | 29/172 | 16.9% [12.0–23.2] |
| 200–299 | short-form | 16/19 | n below 30, no rate quoted |
| 300–399 | short-form | 193/228 | 84.6% [79.4–88.8] |
| 400–599 | short-form | 161/199 | 80.9% [74.9–85.8] |
| 400–599 | long-form | 3/3 | n below 30, no rate quoted |
| 600–849 | short-form | 161/171 | 94.2% [89.6–96.8] |
| 600–849 | long-form | 46/52 | 88.5% [77.0–94.6] |
| 850–1,199 | long-form | 175/193 | 90.7% [85.7–94.0] |
| 1,200–1,699 | long-form | 259/265 | 97.7% [95.1–99.0] |
| 1,700–2,399 | long-form | 271/279 | 97.1% [94.4–98.5] |
| 2,400–3,499 | long-form | 129/130 | 99.2% [95.8–99.9] |
| 3,500–4,999 | long-form | no machine documents | — |
| 5,000 and above | long-form | no machine documents | — |
| all lengths | short-form | 566/816 | 69.4% [66.1–72.4] |
| all lengths | long-form | 883/922 | 95.8% [94.3–96.9] |
DETECTION-BY-LENGTH-AND-MODEL.md Table 1. The shaded region is labelled “unmeasured”: no machine text of this length exists in any corpus on this project. The longest machine document anywhere on the project is 3,061 words, which sits inside the 2,400–3,499 band, so the unmeasured range begins part-way through the last measured column.
| Words | Corpus | Human wrongly flagged | False-positive rate [95% CI] |
|---|---|---|---|
| under 100 | short-form | 0/388 | 0.0% [0.0–1.0] |
| 100–199 | short-form | 0/732 | 0.0% [0.0–0.5] |
| 200–299 | short-form | 3/822 | 0.4% [0.1–1.1] |
| 300–399 | short-form | 9/1,170 | 0.8% [0.4–1.5] |
| 400–599 | short-form | 10/1,167 | 0.9% [0.5–1.6] |
| 400–599 | long-form | 3/826 | 0.4% [0.1–1.1] |
| 600–849 | short-form | 0/86 | 0.0% [0.0–4.3] |
| 600–849 | long-form | 2/894 | 0.2% [0.1–0.8] |
| 850–1,199 | long-form | 30/1,050 | 2.9% [2.0–4.0] |
| 1,200–1,699 | long-form | 3/1,237 | 0.2% [0.1–0.7] |
| 1,700–2,399 | long-form | 6/461 | 1.3% [0.6–2.8] |
| 2,400–3,499 | long-form | 0/134 | 0.0% [0.0–2.8] |
| 3,500–4,999 | long-form | 0/21 | n below 30, no rate quoted |
| 5,000 and above | long-form | 1/13 | n below 30, no rate quoted |
| all lengths | short-form | 22/4,368 | 0.5% [0.3–0.8] |
| all lengths | long-form | 45/4,636 | 1.0% [0.7–1.3] |
DETECTION-BY-LENGTH-AND-MODEL.md Table 1. † The 850–1,199 band: 22 of these 30 are human fiction; excluding fiction, 8/809 = 0.99%.
Three things in those two panels deserve separating out
The cliff is below 200 words, and it is steep. From 16.9% at 100–199 words to 84.6% at 300–399 is 68 percentage points across a gap of a hundred words. These bands are cut by achieved word count rather than by the length a generator was asked for, and that matters: binning by target length puts passages that came back at 200 words and more into a “100-word” band, which lifts the sub-200 figure above what the text at that length actually scores.
Above 300 words the trend is a gradient with noise on it, not a staircase. Short-form reads 84.6%, then 80.9%, then 94.2%. Long-form reads 88.5%, 90.7%, 97.7%, 97.1%, 99.2%. Drawing a smooth curve through those points would claim a precision the denominators do not support, which is why Figure 1 draws columns with intervals and leaves them unjoined.
Human false positives do not rise with length in any straight line. The one band that looks bad is 850–1,199 words at 2.9%, 30 of 1,050, three times the long-form corpus average. Twenty-two of those 30 documents are human fiction. Human fiction has a median length of 1,190 words, so 241 of the corpus’s 260 human stories land in that single band. Take fiction out and the band reads 8 of 809, 0.99%, in line with the corpus. Fiction inside the band reads 22 of 241, 9.1%, which is the known fiction weakness appearing as a length band because of where fiction happens to sit on the axis.
It is evidence that this detector is weak on fiction, measured somewhere the length axis makes it visible. Anyone reading Figure 2 as a length effect will draw the wrong conclusion and mistrust the wrong band.
04 / The headline
The headline describes long documents
The 95.8% figure comes from a corpus whose median machine document is 1,612 words. At 600–849 words the same detector on the same corpus reads 88.5%, 46 of 52. At 850–1,199 words it reads 90.7%, 175 of 193. Those are the lengths most people actually paste in.
The median machine document in the 922-document long-form corpus. Half of what produced the 95.8% figure is longer than a typical long-read article.
46 of 52 long-form machine documents in the 600–849 band, 95% CI 77.0–94.6. The interval is wide because the band holds 52 documents, and that width is part of the answer.
05 / The confound
Length confounds the model comparison, and the confound runs the wrong way
Per-model detection, at the same operating point, on the same 922 machine documents:
| Model | Flagged | Detection rate [95% CI] | Median document length |
|---|---|---|---|
openai/gpt-5.6-luna | 121/121 | 100.0% [96.9–100.0] | 2,263 words |
qwen/qwen3.8-max | 109/110 | 99.1% [95.0–99.8] | not published |
deepseek/deepseek-v4-pro-0813 | 128/131 | 97.7% [93.5–99.2] | 2,144 words |
google/gemini-3.7-flash | 118/121 | 97.5% [93.0–99.2] | not published |
x-ai/grok-4.6 | 113/121 | 93.4% [87.5–96.6] | 1,176 words |
meta-llama/llama-4-maverick | 80/101 | 79.2% [70.3–86.0] | 905 words |
all models | 883/922 | 95.8% [94.3–96.9] | 1,612 words |
Source: DETECTION-BY-LENGTH-AND-MODEL.md Table 2. Further rows in that table hold fewer than 30 documents and print counts only; they are omitted here for the same reason the source file suppresses their rates.
The two lowest-scoring models wrote the shortest documents. Restrict the corpus to documents of 1,200 words or more and it reads 659 of 674, 97.8%; below that, 224 of 248, 90.3%. The model ordering in that table is therefore partly a length ordering, and reading it as a ranking of how detectable each model is would be reading the length axis by another name.
Separating the two would need a length-balanced generation run. A per-model-by-length cross-tabulation was deliberately not produced: 922 documents across 13 models and 8 bands would average nine documents per cell, well under the floor this work applies everywhere else.
06 / Certainty
The detector does not become less certain, it loses a tail
All thirteen models sit at a median machine probability between 98.7% and 99.0%, a spread of three tenths of a percentage point across the whole field, against a primary flag point of 98.55%. meta-llama/llama-4-maverick has the lowest detection rate at 79.2% and the lowest average at 98.4%, and its median is 98.7%, still above the flag point. Its misses are a minority tail rather than a shifted distribution.
| Model | Documents | Average | Median | Detection rate |
|---|---|---|---|---|
meta-llama/llama-4-maverick | 101 | 98.4% | 98.7% | 79.2% |
anthropic/claude-opus-5 | 23 | 98.5% | 99.0% | n below 30 |
x-ai/grok-4.6 | 121 | 98.8% | 99.0% | 93.4% |
google/gemini-3.7-flash | 121 | 98.9% | 99.0% | 97.5% |
deepseek/deepseek-v4-pro-0813 | 131 | 98.9% | 99.0% | 97.7% |
openai/gpt-5.6-luna | 121 | 99.0% | 99.0% | 100.0% |
qwen/qwen3.8-max | 110 | 99.0% | 99.0% | 99.1% |
DETECTION-BY-LENGTH-AND-MODEL.md Table 2. anthropic/claude-opus-5 holds 23 documents, below the 30-document floor, so it carries no published detection rate and is drawn for its distribution shape only. Where a model's median and average are equal only one marker is visible and only one figure is printed.
anthropic/claude-opus-5 shows the split most clearly: average 98.5%, median 99.0%, on 23 documents. Two documents drag the average below the flag point while the median sits above it. That is why both are printed, and why the machine and human populations are counted and averaged separately everywhere in this work, never as one number across both.
A related piece of arithmetic explains a column that looks alarming and is not. The average machine probability of human documents climbs steadily with length, from 17.8% under 100 words to 90.5% above 5,000 words, while the false-positive rate does not. The reported probability is the maximum across sections, and a longer document offers more sections for that maximum to be drawn from.
07 / Limits
What has not been measured
Nothing above 3,061 words. That is the longest machine document in any corpus on this project. The 3,500–4,999 and 5,000-and-above bands contain human documents only, 21 and 13 of them, both under the 30-document floor. No detection rate should be inferred there. The server refuses documents over 4,000 words and offers the browser route instead, so the unmeasured range begins inside what the server will accept and continues past it.
Nothing in the browser. Every figure on this page is fp32 through Python onnxruntime, and no browser figure appears in any cell or any sentence of it. The browser’s own full-corpus segmented length curve has never been measured. The two runtimes agree closely in the decision region, which is reassuring and is not the same as having measured the thing.
Nothing about edited text. Every machine document here is fully model-generated. Mixed documents, rewrites and human drafts a model tidied are a different problem with different numbers, and this corpus contains none of them.
A band with no rate in it is not a band with a low rate in it.
08 / Independence
The corpus is not fully held out, and by how much
Measured against cycle2-train/dataset.jsonl on normalised SHA-256, 268 of the 922 machine documents — 29.1% — appear in the cycle-2 dataset, 168 of them in the train split. The human half is effectively clean: 11 of 4,636, being 5 train, 3 calibration and 3 test.
The effect was measured rather than assumed. The subset cut was taken at a threshold of 0.984, which is the retired single-threshold rule and not the pair that ships, and the subset figures below are quoted at that retired point because that is where the cut exists.
Measured at a rule that no longer ships
- Corpus
- The 922 machine documents of the long-form corpus, split by whether each appears in the cycle-2 dataset on a normalised SHA-256 hash.
- Retired flag point
- 0.984
- Detector
tier3-cycle2-e5small-fp32.onnx, SHA-256e313ab00de1fffd2…4d2788d- Runtime
- Python onnxruntime 1.29.0, CPU, fp32
- Measured
- 30 August 2026
- Also
- Every subset figure below is at the retired single-threshold rule. None has been re-cut at the shipped pair.
The rule that ships today is 0.9855 / 0.9763. The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.
| Subset | n | Detected at the retired 0.984 rule |
|---|---|---|
| Never in the cycle-2 dataset | 654 | 620 = 94.80% (retired 0.984 rule) |
| In the cycle-2 dataset, any split | 268 | 257 = 95.90% (retired 0.984 rule) |
| Of which the train split | 168 | 163 = 97.02% (retired 0.984 rule) |
Source: corpus-reconciliation-2026-08-29/analysis.txt §2. These subset figures have not been re-cut at the shipped pair, and relabelling them as if they had would be the same error this page exists to correct.
Roughly 1.1 points separate the seen and unseen subsets at that threshold, worth about 0.3 points on the corpus headline. It changes no conclusion on this page. Where an argument rests entirely on unseen data, the 654-document independent subset is the honest denominator to use.
09 / Provenance
Where every figure on this page came from
| Figure | File | Section |
|---|---|---|
| Every cell of the length curve; 883/922; 45/4,636; median 1,612 words; the 850–1,199 fiction decomposition | docs/measurements/DETECTION-BY-LENGTH-AND-MODEL.md | Table 1 and “What Table 1 says” |
| Per-model rows, median document lengths, 659/674 and 224/248, the median-against-average split | docs/measurements/DETECTION-BY-LENGTH-AND-MODEL.md | Table 2 and “What Table 2 says” |
| No machine text above 3,061 words; no browser measurement; nothing about edited text | docs/measurements/DETECTION-BY-LENGTH-AND-MODEL.md | “What is not here” |
| Detector SHA, operating point, T = 0.8324, segments-v3, onnxruntime 1.29.0 | docs/measurements/DETECTION-BY-LENGTH-AND-MODEL.md | Table 1 header |
| 268/922, 168 in the train split, 11/4,636 human, and the two subset figures at the retired 0.984 rule | corpus-reconciliation-2026-08-29/analysis.txt | §2 |
| The shipped pair definition and the two-runtime fit | docs/programme/HANDOVER.md | §4.2, §4.4 |
The measurement script behind Tables 1 and 2 reproduces six previously published figures exactly before it emits a single cell, and by_register_and_length.py exits without printing anything if any of them fails to reproduce. Two of those gates are the ones this page publishes, 883/922 and 45/4,636 at the shipped pair. The others are reproductions at retired operating points, and they are not printed here, because a figure at a rule that no longer ships has no business appearing under a current heading.
Every file named above ships with the open measurement repository, so any figure on this page can be re-run rather than taken on trust.