Get in Touch
AI content toolsfrom Opace

Measurement paper · 922 long-form and 816 short-form machine documents

Detection is mostly a question of length

How much text you paste in changes the answer more than which model wrote it. Between 100 and 199 words the detector flags 29 of 172 machine documents, 16.9%. On a long-form corpus whose median document runs to 1,612 words it flags 883 of 922, 95.8%. Both are the same detector at the same operating point.

On this page
  1. The finding
  2. The measurement conditions
  3. The curve, and the floor under it
  4. The headline describes long documents
  5. Length confounds the model comparison
  6. A lost tail, not a lost signal
  7. What has not been measured
  8. The corpus is not fully held out
  9. Provenance
Published
30 August 2026
Measured
30 August 2026, re-verified against source the same day
Method
segments-v3 · shipped minimum-evidence pair · fp32 reference server
Corpus
816 short-form and 922 long-form machine documents; 4,368 and 4,636 human

01 / The finding

The cliff is below 200 words, and the headline rate is a long-document rate

How much text you paste in changes the answer more than which model wrote it, what register it is in, or anything else we have measured.

Between 100 and 199 words the detector flags 29 of 172 machine documents. Between 300 and 399 it flags 193 of 228. Same detector, same operating point, a hundred words apart.
100 to 199 words16.9%

29 of 172 machine documents, 95% CI 12.0–23.2. Five machine passages in six go past at this length.

Whole long-form corpus95.8%

883 of 922 machine documents, 95% CI 94.3–96.9. That corpus has a median machine document of 1,612 words.

Across one hundred words68 points

From 16.9% at 100–199 words to 84.6% at 300–399. Nothing else measured on this project moves the answer that far.

The headline figure is a long-document figure and it should be read as one. Somebody checking an 800-word blog post is not operating at 95.8%, and the interface should say so where the text is pasted rather than in a footnote further down the page.

02 / Conditions

The measurement conditions, which apply to every figure on this page

Measurement conditions

Corpus
Two corpora, never pooled into one cell. Short-form: 816 machine passages generated 29 August 2026 from openai/gpt-5.6-sol and openai/gpt-5.6-luna across four target lengths and three prompt styles, with 4,368 human passages from nine sources. Long-form: the 5,558-document set of 28 August 2026, being 922 machine documents across 13 models and 4,636 human documents from eight sources.
Operating point
0.9855 / 0.9763 · contract segments-v3 · T = 0.8324
Detector
tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d
Runtime
Python onnxruntime 1.29.0, CPU, fp32
Measured
30 August 2026
Also
segments-v3 token-bounded segmentation. The reported document probability is the maximum across sections.

The detector is the fp32 parent of the shipped browser artefact and the file the EU inference server runs. Every figure below came through Python onnxruntime on CPU, which is the reference-server scoring path and not a browser measurement. Where an older figure is quoted for contrast, its own operating point is named beside it.

The two corpora are never pooled, because they were generated differently, by different models, in different registers. Bands were chosen from the corpora’s own word-count distributions rather than from round numbers, and where both corpora cover a band each keeps its own cell. Any cell holding fewer than 30 documents prints its count and no rate, which is the floor the source file applies throughout: a 21 of 23 is not “91.3%” in any sense that should travel.

03 / The curve

The curve, and the floor under it

Figure 1 Machine detection by achieved word count, with each corpus kept separate Detection rate on a 0–100% axis, with Wilson 95% intervals as whiskers and the machine denominator printed under every column. Two corpora, drawn in two colours and never joined into one line: they overlap between 400 and 849 words and were generated differently, so a line across them would draw a shape nobody measured. Columns are labelled with the first word count in their band and the exact ranges are in the table below. A band holding fewer than 30 machine documents draws an open outline with its raw count and no rate. Shipped minimum-evidence pair, fp32 reference server, measured 30 August 2026.
Machine detection rate by word-count band, short-form and long-form corpora kept separate unmeasured 0% 20% 40% 60% 80% 100% detection rate corpus headline, 883/922 n below 30 <100 short 6/27 16.9% 100 short 29/172 n below 30 200 short 16/19 84.6% 300 short 193/228 80.9% 400 short 161/199 n below 30 400 long 3/3 94.2% 600 short 161/171 88.5% 600 long 46/52 90.7% 850 long 175/193 97.7% 1,200 long 259/265 97.1% 1,700 long 271/279 99.2% 2,400 long 129/130 no docs 3,500 long 0 machine no docs 5,000 long 0 machine
Short-form corpus, 816 machine passages Long-form corpus, 922 machine documents Dashed line: the long-form corpus headline, 883/922 = 95.8% Shaded: no machine text of this length exists in any corpus on this project The corpus median machine document is 1,612 words, which falls inside the 1,200–1,699 band
Machine detection by word-count band and corpus, with Wilson 95% intervals
WordsCorpusMachine flaggedDetection rate [95% CI]
under 100 short-form 6/27 n below 30, no rate quoted
100–199 short-form 29/172 16.9% [12.0–23.2]
200–299 short-form 16/19 n below 30, no rate quoted
300–399 short-form 193/228 84.6% [79.4–88.8]
400–599 short-form 161/199 80.9% [74.9–85.8]
400–599 long-form 3/3 n below 30, no rate quoted
600–849 short-form 161/171 94.2% [89.6–96.8]
600–849 long-form 46/52 88.5% [77.0–94.6]
850–1,199 long-form 175/193 90.7% [85.7–94.0]
1,200–1,699 long-form 259/265 97.7% [95.1–99.0]
1,700–2,399 long-form 271/279 97.1% [94.4–98.5]
2,400–3,499 long-form 129/130 99.2% [95.8–99.9]
3,500–4,999 long-form no machine documents
5,000 and above long-form no machine documents
all lengthsshort-form566/81669.4% [66.1–72.4]
all lengthslong-form883/92295.8% [94.3–96.9]

DETECTION-BY-LENGTH-AND-MODEL.md Table 1. The shaded region is labelled “unmeasured”: no machine text of this length exists in any corpus on this project. The longest machine document anywhere on the project is 3,061 words, which sits inside the 2,400–3,499 band, so the unmeasured range begins part-way through the last measured column.

Figure 2 Human documents wrongly flagged, on the same band axis The same fourteen columns as Figure 1, for human documents. Note the axis: this panel runs 0–5%, not 0–100%, so a column that looks tall here is a rate that would be invisible on the panel above. Wilson 95% intervals as whiskers, human denominators under every column. The two longest bands hold 21 and 13 human documents, both below the 30-document floor, so their counts are printed and no rate is. Shipped minimum-evidence pair, fp32 reference server, measured 30 August 2026.
Human false-positive rate by word-count band, on a 0 to 5 per cent axis 0% 1% 2% 3% 4% 5% false positives, 0–5% axis long-form corpus average, 45/4,636 0.0% <100 short 0/388 0.0% 100 short 0/732 0.4% 200 short 3/822 0.8% 300 short 9/1,170 0.9% 400 short 10/1,167 0.4% 400 long 3/826 0.0% 600 short 0/86 0.2% 600 long 2/894 2.9% † 850 long 30/1,050 0.2% 1,200 long 3/1,237 1.3% 1,700 long 6/461 0.0% 2,400 long 0/134 n below 30 3,500 long 0/21 n below 30 5,000 long 1/13
Short-form corpus, 4,368 human passages Long-form corpus, 4,636 human documents Axis maximum 5%, not 100% † 850–1,199 words: 22 of these 30 are human fiction; excluding fiction, 8/809 = 0.99%
Human false positives by word-count band and corpus, with Wilson 95% intervals
WordsCorpusHuman wrongly flaggedFalse-positive rate [95% CI]
under 100 short-form 0/388 0.0% [0.0–1.0]
100–199 short-form 0/732 0.0% [0.0–0.5]
200–299 short-form 3/822 0.4% [0.1–1.1]
300–399 short-form 9/1,170 0.8% [0.4–1.5]
400–599 short-form 10/1,167 0.9% [0.5–1.6]
400–599 long-form 3/826 0.4% [0.1–1.1]
600–849 short-form 0/86 0.0% [0.0–4.3]
600–849 long-form 2/894 0.2% [0.1–0.8]
850–1,199 long-form 30/1,050 2.9% [2.0–4.0]
1,200–1,699 long-form 3/1,237 0.2% [0.1–0.7]
1,700–2,399 long-form 6/461 1.3% [0.6–2.8]
2,400–3,499 long-form 0/134 0.0% [0.0–2.8]
3,500–4,999 long-form 0/21 n below 30, no rate quoted
5,000 and above long-form 1/13 n below 30, no rate quoted
all lengthsshort-form22/4,3680.5% [0.3–0.8]
all lengthslong-form45/4,6361.0% [0.7–1.3]

DETECTION-BY-LENGTH-AND-MODEL.md Table 1. † The 850–1,199 band: 22 of these 30 are human fiction; excluding fiction, 8/809 = 0.99%.

Three things in those two panels deserve separating out

The cliff is below 200 words, and it is steep. From 16.9% at 100–199 words to 84.6% at 300–399 is 68 percentage points across a gap of a hundred words. These bands are cut by achieved word count rather than by the length a generator was asked for, and that matters: binning by target length puts passages that came back at 200 words and more into a “100-word” band, which lifts the sub-200 figure above what the text at that length actually scores.

Above 300 words the trend is a gradient with noise on it, not a staircase. Short-form reads 84.6%, then 80.9%, then 94.2%. Long-form reads 88.5%, 90.7%, 97.7%, 97.1%, 99.2%. Drawing a smooth curve through those points would claim a precision the denominators do not support, which is why Figure 1 draws columns with intervals and leaves them unjoined.

Human false positives do not rise with length in any straight line. The one band that looks bad is 850–1,199 words at 2.9%, 30 of 1,050, three times the long-form corpus average. Twenty-two of those 30 documents are human fiction. Human fiction has a median length of 1,190 words, so 241 of the corpus’s 260 human stories land in that single band. Take fiction out and the band reads 8 of 809, 0.99%, in line with the corpus. Fiction inside the band reads 22 of 241, 9.1%, which is the known fiction weakness appearing as a length band because of where fiction happens to sit on the axis.

That cell is not evidence that detection is unreliable around a thousand words.

It is evidence that this detector is weak on fiction, measured somewhere the length axis makes it visible. Anyone reading Figure 2 as a length effect will draw the wrong conclusion and mistrust the wrong band.

04 / The headline

The headline describes long documents

The 95.8% figure comes from a corpus whose median machine document is 1,612 words. At 600–849 words the same detector on the same corpus reads 88.5%, 46 of 52. At 850–1,199 words it reads 90.7%, 175 of 193. Those are the lengths most people actually paste in.

What the headline is measured on 1,612 words

The median machine document in the 922-document long-form corpus. Half of what produced the 95.8% figure is longer than a typical long-read article.

What an 800-word post gets 88.5%

46 of 52 long-form machine documents in the 600–849 band, 95% CI 77.0–94.6. The interval is wide because the band holds 52 documents, and that width is part of the answer.

05 / The confound

Length confounds the model comparison, and the confound runs the wrong way

Per-model detection, at the same operating point, on the same 922 machine documents:

Per-model detection rate and median document length on 922 machine documents
ModelFlaggedDetection rate [95% CI]Median document length
openai/gpt-5.6-luna 121/121 100.0% [96.9–100.0] 2,263 words
qwen/qwen3.8-max 109/110 99.1% [95.0–99.8] not published
deepseek/deepseek-v4-pro-0813 128/131 97.7% [93.5–99.2] 2,144 words
google/gemini-3.7-flash 118/121 97.5% [93.0–99.2] not published
x-ai/grok-4.6 113/121 93.4% [87.5–96.6] 1,176 words
meta-llama/llama-4-maverick 80/101 79.2% [70.3–86.0] 905 words
all models 883/922 95.8% [94.3–96.9] 1,612 words

Source: DETECTION-BY-LENGTH-AND-MODEL.md Table 2. Further rows in that table hold fewer than 30 documents and print counts only; they are omitted here for the same reason the source file suppresses their rates.

The two lowest-scoring models wrote the shortest documents. Restrict the corpus to documents of 1,200 words or more and it reads 659 of 674, 97.8%; below that, 224 of 248, 90.3%. The model ordering in that table is therefore partly a length ordering, and reading it as a ranking of how detectable each model is would be reading the length axis by another name.

Separating the two would need a length-balanced generation run. A per-model-by-length cross-tabulation was deliberately not produced: 922 documents across 13 models and 8 bands would average nine documents per cell, well under the floor this work applies everywhere else.

06 / Certainty

The detector does not become less certain, it loses a tail

All thirteen models sit at a median machine probability between 98.7% and 99.0%, a spread of three tenths of a percentage point across the whole field, against a primary flag point of 98.55%. meta-llama/llama-4-maverick has the lowest detection rate at 79.2% and the lowest average at 98.4%, and its median is 98.7%, still above the flag point. Its misses are a minority tail rather than a shifted distribution.

Figure 3 Median against average machine probability, per model, either side of the flag point Each row runs from that model's average machine probability to its median, on a truncated 97.5–99.5% axis. The vertical rule is the primary flag point, 98.55%. Every median sits above it, including the two models whose average does not, which is what makes the detection differences tail losses rather than a detector that has become less certain about some models than others. Right-hand column is the model's detection rate at the shipped minimum-evidence pair, on the fp32 reference server, measured 30 August 2026. Seven of the thirteen models are drawn; the rows not drawn hold fewer than 30 documents each and the source prints counts for them and no rate.
Median and average machine probability per model against the primary flag point 97.5% 98.0% 98.5% 99.0% 99.5% Detection meta-llama/llama-4-maverick 101 documents 98.4% 98.7% 79.2% anthropic/claude-opus-5 23 documents · below the floor 98.5% 99.0% n below 30 x-ai/grok-4.6 121 documents 98.8% 99.0% 93.4% google/gemini-3.7-flash 121 documents 98.9% 99.0% 97.5% deepseek/deepseek-v4-pro-0813 131 documents 98.9% 99.0% 97.7% openai/gpt-5.6-luna 121 documents 99.0% 100.0% qwen/qwen3.8-max 110 documents 99.0% 99.1% primary flag point, 98.55% Value axis starts at 97.5%, not zero
Median machine probability Average machine probability Axis truncated at 97.5%; the whole field spans three tenths of a point
Average and median machine probability per model, with document counts
ModelDocumentsAverageMedianDetection rate
meta-llama/llama-4-maverick 101 98.4% 98.7% 79.2%
anthropic/claude-opus-5 23 98.5% 99.0% n below 30
x-ai/grok-4.6 121 98.8% 99.0% 93.4%
google/gemini-3.7-flash 121 98.9% 99.0% 97.5%
deepseek/deepseek-v4-pro-0813 131 98.9% 99.0% 97.7%
openai/gpt-5.6-luna 121 99.0% 99.0% 100.0%
qwen/qwen3.8-max 110 99.0% 99.0% 99.1%

DETECTION-BY-LENGTH-AND-MODEL.md Table 2. anthropic/claude-opus-5 holds 23 documents, below the 30-document floor, so it carries no published detection rate and is drawn for its distribution shape only. Where a model's median and average are equal only one marker is visible and only one figure is printed.

anthropic/claude-opus-5 shows the split most clearly: average 98.5%, median 99.0%, on 23 documents. Two documents drag the average below the flag point while the median sits above it. That is why both are printed, and why the machine and human populations are counted and averaged separately everywhere in this work, never as one number across both.

A related piece of arithmetic explains a column that looks alarming and is not. The average machine probability of human documents climbs steadily with length, from 17.8% under 100 words to 90.5% above 5,000 words, while the false-positive rate does not. The reported probability is the maximum across sections, and a longer document offers more sections for that maximum to be drawn from.

07 / Limits

What has not been measured

Nothing above 3,061 words. That is the longest machine document in any corpus on this project. The 3,500–4,999 and 5,000-and-above bands contain human documents only, 21 and 13 of them, both under the 30-document floor. No detection rate should be inferred there. The server refuses documents over 4,000 words and offers the browser route instead, so the unmeasured range begins inside what the server will accept and continues past it.

Nothing in the browser. Every figure on this page is fp32 through Python onnxruntime, and no browser figure appears in any cell or any sentence of it. The browser’s own full-corpus segmented length curve has never been measured. The two runtimes agree closely in the decision region, which is reassuring and is not the same as having measured the thing.

Nothing about edited text. Every machine document here is fully model-generated. Mixed documents, rewrites and human drafts a model tidied are a different problem with different numbers, and this corpus contains none of them.

A band with no rate in it is not a band with a low rate in it.

Wherever a cell on this page holds fewer than 30 documents, the count is printed and the rate is withheld. Figures 1 and 2 draw those cells as open outlines for the same reason: a rate we declined to compute and a rate that happens to be low must not look alike.

08 / Independence

The corpus is not fully held out, and by how much

The 5,558-document long-form corpus is described in some internal records as hash-quarantined against every training split. That wording is not accurate and should not be repeated.

Measured against cycle2-train/dataset.jsonl on normalised SHA-256, 268 of the 922 machine documents — 29.1% — appear in the cycle-2 dataset, 168 of them in the train split. The human half is effectively clean: 11 of 4,636, being 5 train, 3 calibration and 3 test.

The effect was measured rather than assumed. The subset cut was taken at a threshold of 0.984, which is the retired single-threshold rule and not the pair that ships, and the subset figures below are quoted at that retired point because that is where the cut exists.

Measured at a rule that no longer ships

Corpus
The 922 machine documents of the long-form corpus, split by whether each appears in the cycle-2 dataset on a normalised SHA-256 hash.
Retired flag point
0.984
Detector
tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d
Runtime
Python onnxruntime 1.29.0, CPU, fp32
Measured
30 August 2026
Also
Every subset figure below is at the retired single-threshold rule. None has been re-cut at the shipped pair.

The rule that ships today is 0.9855 / 0.9763. The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.

Detection by training-set overlap subset, at the retired 0.984 single-threshold rule
SubsetnDetected at the retired 0.984 rule
Never in the cycle-2 dataset654620 = 94.80% (retired 0.984 rule)
In the cycle-2 dataset, any split268257 = 95.90% (retired 0.984 rule)
Of which the train split168163 = 97.02% (retired 0.984 rule)

Source: corpus-reconciliation-2026-08-29/analysis.txt §2. These subset figures have not been re-cut at the shipped pair, and relabelling them as if they had would be the same error this page exists to correct.

Roughly 1.1 points separate the seen and unseen subsets at that threshold, worth about 0.3 points on the corpus headline. It changes no conclusion on this page. Where an argument rests entirely on unseen data, the 654-document independent subset is the honest denominator to use.

09 / Provenance

Where every figure on this page came from

Source file and section for every figure on this page
FigureFileSection
Every cell of the length curve; 883/922; 45/4,636; median 1,612 words; the 850–1,199 fiction decompositiondocs/measurements/DETECTION-BY-LENGTH-AND-MODEL.mdTable 1 and “What Table 1 says”
Per-model rows, median document lengths, 659/674 and 224/248, the median-against-average splitdocs/measurements/DETECTION-BY-LENGTH-AND-MODEL.mdTable 2 and “What Table 2 says”
No machine text above 3,061 words; no browser measurement; nothing about edited textdocs/measurements/DETECTION-BY-LENGTH-AND-MODEL.md“What is not here”
Detector SHA, operating point, T = 0.8324, segments-v3, onnxruntime 1.29.0docs/measurements/DETECTION-BY-LENGTH-AND-MODEL.mdTable 1 header
268/922, 168 in the train split, 11/4,636 human, and the two subset figures at the retired 0.984 rulecorpus-reconciliation-2026-08-29/analysis.txt§2
The shipped pair definition and the two-runtime fitdocs/programme/HANDOVER.md§4.2, §4.4

The measurement script behind Tables 1 and 2 reproduces six previously published figures exactly before it emits a single cell, and by_register_and_length.py exits without printing anything if any of them fails to reproduce. Two of those gates are the ones this page publishes, 883/922 and 45/4,636 at the shipped pair. The others are reproductions at retired operating points, and they are not printed here, because a figure at a rule that no longer ships has no business appearing under a current heading.

Every file named above ships with the open measurement repository, so any figure on this page can be re-run rather than taken on trust.

Read on

Length sets the ceiling. What the detector reads at that length is the next question.