Get in Touch
AI content toolsfrom Opace

Measured evidence · fp32 reference route · 30 August 2026

Detection rates In full

Three tables, every cell with its denominator and its confidence interval: how often the trained classifier flags AI writing and human writing, cut by document length, by the model that wrote the text, and by content type.

Evidence sheetBrowser-only
SourceExact text and protected factsHeld in memory for this run
SignalsNamed checks with limitsNo universal authorship score
Watermark scanKnown keysReal SynthID mathematics; Anthropic production keys stay private
Open the checker
On this page
  1. What these tables are
  2. What produced every cell
  3. Table 1 — by document length
  4. Table 2 — by the model that wrote it
  5. Table 3 — by content type
  6. The rest of the record
Published
30 August 2026
Measured
30 August 2026, re-derived 30 August 2026
Method contract
Contract 1.0.0 · segments-v3
Operating point
0.9855 primary / 0.9763 second-highest

01 / Read this first

What these tables are, and what they are not

They are measurements on a corpus that is mostly, but not entirely, independent of the model. Every document was scored through the same path the EU inference server runs, at the operating point that ships today, and every cell states how many documents it counted and how many were flagged, with a 95% confidence interval. Until 30 August 2026 this page called the corpus held out and said the model had never seen it. That was wrong for the AI half, and the correction is in the provenance block below rather than folded away.

They are not a guarantee about any individual document. A rate describes a population. It does not tell you the chance that the specific draft in front of you was written by a machine, and no number on this page should be used to make an accusation about a person. The checker is evidence for review: no score it produces identifies an author, and a quiet result does not establish who wrote the text.

They describe one check. The trained classifier is the only check in this tool that gives an AI reading. The hidden-character, lookalike, protected-fact, writing-signal and watermark checks are separate, are reported separately, and are not measured here.

They describe the server route. Every figure below is fp32 through Python onnxruntime, which is the reference-server scoring path. The in-browser route runs the same model in lower numerical precision; the two agree to a median of 0.0002 in the decision region, but a proxy is not the thing, and the browser's own full-corpus curve at this operating point is measured separately rather than borrowed from here.

The headline The trained classifier detects 883/922 (95.8%) of AI-written long-form documents on the EU server route and 889/922 (96.4%) in the browser, and wrongly flags 45/4,636 (0.97%) of human-written long-form documents on the server and 90/4,636 (1.94%) in the browser. Both figures are measured over the whole corpus at the operating point that ships.
Its weakest case, beside it 23/260 (8.8%) of human short stories are wrongly flagged on the server route, interval 6.0 to 12.9%.

The worst register measured, at the shipped operating point (0.9855/0.9763), over all 260 human fiction documents. Fiction sits in a dense cliff just below the flag point, so this rate moves faster with the threshold than any other register and has to be looked at separately from the corpus rate. See Table 3.

02 / Provenance

What produced every cell below

Detector
cycle-2, the fp32 parent of the int8 artefact that runs in your browser, and the exact file the EU inference server runs
Model file
tier3-cycle2-e5small-fp32.onnx
Model SHA-256
e313ab00de1fffd28d6157f014065b50bca8b59a8842746e54fe8b1504d2788d
Operating point
the shipped pair: a document is flagged when its strongest section reaches 98.55%, or its second-strongest reaches 97.63%
Pipeline
segments-v3 token-bounded segmentation. the document's AI probability is the maximum across its sections, exactly as /v1/check reports it; nothing is averaged
Calibration
probabilities are temperature-calibrated at T = 0.8324 before the flag rule is applied
Runtime
Python onnxruntime 1.29.0, CPU, fp32 — the reference-server scoring path. Not a browser measurement.
Long-form corpus
the 5,558-document long-form corpus of 28 August 2026: 922 AI documents across 13 models and 4,636 human documents from 8 sources. NOT fully held out — see the correction below.
Short-form corpora
816 AI passages generated on 29 August 2026 from two OpenAI models across four target lengths and three prompt styles, 432 further AI passages generated to be deliberately keyword-repetitive, and 4,368 human passages cut from the long-form corpus across 9 sources
Sparse cells
30 documents. Below it a cell prints its counts and says so, rather than printing a rate.
Measured
30 August 2026, re-derived 30 August 2026
Harness reproduction
The re-score reproduces 883/922 AI detected and 45/4,636 human false positives at the shipped pair, 877/922 and 56/4,636 at the prior 0.984 maximum-only rule, and the short-form pilot's 44/195, 175/206, 169/208 and 194/207 at the 0.9845 maximum-only threshold, before any cell here was emitted.

Correction, 30 August 2026: this corpus is not fully held out.

654 of the 922 AI documents are independent of every cycle-2 split. 268 are not, 168 of them in the training split. Every rate on this page pools both subsets, and no seen-against-unseen split has been measured at the operating point that ships.

The full record, with what was measured and at which flag point

CORRECTED 30 August 2026: this corpus was published as held out and hash-quarantined against every training split, and it is not. 654 of the 922 AI documents are independent of every cycle-2 split; 268 are not (168 training, 72 test, 28 calibration), as are 11 of the 4,636 human documents. The difference has been measured only at the superseded 0.984 single-threshold rule, where the independent subset reads 620/654 (94.80%) against 257/268 (95.90%) for the seen subset, a gap of 1.1 points. It has NOT been measured at the shipped 0.9855/0.9763 pair, so every rate on this page pools both subsets and no seen-against-unseen split is published for the operating point that ships.

What the difference measures, and at which flag point. The split has been scored at one operating point only, and it is not the one that ships. At the superseded 0.984 single-threshold rule the independent subset reads 620/654, 94.80%, against 257/268, 95.90% for the subset the model had seen, of which the training split alone reads 163/168, 97.02% — a gap of 1.1 percentage points. No seen-against-unseen split has been measured at the shipped 0.9855/0.9763 pair, so none is published for it, and the 1.1-point figure above must not be read as if it described the shipped rule. A re-measurement at the shipped pair is outstanding.

Source: services/local-engine/research/corpus-reconciliation-2026-08-29/analysis.txt, section 2, in the open measurement repository.

Percentages in the probability columns are the model's calibrated AI probability expressed as a percentage: a document that scored 0.9886 reads 98.9%. AI documents and human documents are counted, and their probabilities averaged, separately. A single average across both populations would describe neither.

03 / Table 1

By document length

Read as: of documents this long, how many were flagged, and what did the two populations score. Two corpora cover this range and they overlap between 400 and 849 words, so each is given its own figure rather than pooled: they were generated differently, by different models, in different registers.

Length is the dominant axis, and the cliff is below 200 words.

Binned by the words a passage actually has rather than the length a generator was asked for, 100–199 words detects 29 of 172 on the short-form corpus. Above 300 words the trend is upward but not monotone, so the axis is a gradient with noise on it rather than a clean staircase.
Figure 1 Long-form corpus, 5,558 documents Both populations on one 0–100% axis. The human line is labelled at every band because it sits near the floor throughout; the one band that rises is a register effect, explained below. Bands with no measured rate break the line rather than being joined across.
AI detection rate and human false-positive rate by document length, long-form corpus 0% 20% 40% 60% 80% 100% 400–599 600–849 850–1,199 1,200–1,699 1,700–2,399 2,400–3,499 3,500–4,999 5,000 and above n<30 nonenone 0.4% 0.2% 2.9% 0.2% 1.3% 0%n<30n<30
AI documents flagged Human documents wrongly flagged No rate quoted below 30 documents
AI detection rate and human false-positive rate by document length, long-form corpus
Words AI documents Human documents
Flagged, with 95% interval AI probability, avg / median Wrongly flagged, with 95% interval AI probability, avg / median
400–599
3/3n below 30 — no rate quoted
avg 98.8%
med 98.8%
0.4% 3/826 documents · 95% CI 0.1 to 1.1%
avg 40.5%
med 30.8%
600–849
88.5% 46/52 documents · 95% CI 77.0 to 94.6%
avg 98.7%
med 98.8%
0.2% 2/894 documents · 95% CI 0.1 to 0.8%
avg 41.4%
med 35.5%
850–1,199
90.7% 175/193 documents · 95% CI 85.7 to 94.0%
avg 98.8%
med 99.0%
2.9% 30/1,050 documents · 95% CI 2.0 to 4.0%
avg 47.4%
med 45.5%
1,200–1,699
97.7% 259/265 documents · 95% CI 95.1 to 99.0%
avg 98.9%
med 99.0%
0.2% 3/1,237 documents · 95% CI 0.1 to 0.7%
avg 36.9%
med 26.5%
1,700–2,399
97.1% 271/279 documents · 95% CI 94.4 to 98.5%
avg 98.9%
med 99.0%
1.3% 6/461 documents · 95% CI 0.6 to 2.8%
avg 57.2%
med 63.3%
2,400–3,499
99.2% 129/130 documents · 95% CI 95.8 to 99.9%
avg 98.9%
med 99.0%
0.0% 0/134 documents · 95% CI 0.0 to 2.8%
avg 65.7%
med 75.3%
3,500–4,999
no documents
avg
med
0/21n below 30 — no rate quoted
avg 65.8%
med 69.9%
5,000 and above
no documents
avg
med
1/13n below 30 — no rate quoted
avg 90.5%
med 96.6%
all lengths
95.8% 883/922 documents · 95% CI 94.3 to 96.9%
avg 98.9%
med 99.0%
1.0% 45/4,636 documents · 95% CI 0.7 to 1.3%
avg 43.9%
med 37.7%

Every point is drawn from the cell printed in the table below it. A band with no rate draws no point.

Figure 2 Short-form corpus, 816 AI and 4,368 human passages The same axes on the corpus built to probe short text. The shaded region is where the signal collapses: below 200 words most of what the model reads is gone.
AI detection rate and human false-positive rate by document length, short-form corpus Signal collapses below 200 words 0% 20% 40% 60% 80% 100% under 100 100–199 200–299 300–399 400–599 600–849 n<30 n<30 0% 0% 0.4% 0.8% 0.9% 0%
AI documents flagged Human documents wrongly flagged
AI detection rate and human false-positive rate by document length, short-form corpus
Words AI documents Human documents
Flagged, with 95% interval AI probability, avg / median Wrongly flagged, with 95% interval AI probability, avg / median
under 100
6/27n below 30 — no rate quoted
avg 82.5%
med 94.2%
0.0% 0/388 documents · 95% CI 0.0 to 1.0%
avg 17.8%
med 10.7%
100–199
16.9% 29/172 documents · 95% CI 12.0 to 23.2%
avg 84.6%
med 97.3%
0.0% 0/732 documents · 95% CI 0.0 to 0.5%
avg 16.8%
med 9.6%
200–299
16/19n below 30 — no rate quoted
avg 97.3%
med 98.9%
0.4% 3/822 documents · 95% CI 0.1 to 1.1%
avg 21.0%
med 8.8%
300–399
84.6% 193/228 documents · 95% CI 79.4 to 88.8%
avg 96.9%
med 98.9%
0.8% 9/1,170 documents · 95% CI 0.4 to 1.5%
avg 20.6%
med 7.7%
400–599
80.9% 161/199 documents · 95% CI 74.9 to 85.8%
avg 96.0%
med 98.9%
0.9% 10/1,167 documents · 95% CI 0.5 to 1.6%
avg 30.7%
med 18.2%
600–849
94.2% 161/171 documents · 95% CI 89.6 to 96.8%
avg 98.2%
med 99.0%
0.0% 0/86 documents · 95% CI 0.0 to 4.3%
avg 40.5%
med 33.7%
all lengths
69.4% 566/816 documents · 95% CI 66.1 to 72.4%
avg 93.9%
med 98.9%
0.5% 22/4,368 documents · 95% CI 0.3 to 0.8%
avg 22.9%
med 11.0%

A separate corpus from Figure 1, generated differently and by different models. The two overlap between 400 and 849 words and are never pooled.

The keyword-repetition arm, fenced separately

A further 432 AI passages were generated to be deliberately keyword-repetitive, as a test of the evasion axis that matters most commercially. They are an adversarial condition, not a sample of ordinary short text. Pooling them into the figures above would understate the detector against normal copy in the same way that omitting them would overstate it against SEO copy, so they sit here. There is no matched human corpus for this condition; only the AI side exists.

Figure 3 Keyword-repetition arm, 432 AI passages An adversarial condition with no human line, because no matched human corpus for it exists. Read against Figure 2: at every band the same detector reads roughly half as much of this text as it does of ordinary short copy.
AI detection rate by document length, keyword-repetition arm 0% 20% 40% 60% 80% 100% under 100 100–199 200–299 300–399 400–599 600–849 n<30 n<30
AI documents flagged
Show the numbers this was drawn from
AI detection rate by document length, keyword-repetition arm
Words Flagged, with 95% interval AI probability, avg / median Human documents
under 100
2/16n below 30 — no rate quoted
avg 63.6%
med 94.2%
no matched human corpus exists for this condition
100–199
14.0% 8/57 documents · 95% CI 7.3 to 25.3%
avg 63.0%
med 84.2%
no matched human corpus exists for this condition
200–299
3/9n below 30 — no rate quoted
avg 77.8%
med 93.6%
no matched human corpus exists for this condition
300–399
44.9% 61/136 documents · 95% CI 36.7 to 53.2%
avg 76.1%
med 98.1%
no matched human corpus exists for this condition
400–599
46.8% 58/124 documents · 95% CI 38.2 to 55.5%
avg 81.9%
med 98.3%
no matched human corpus exists for this condition
600–849
62.2% 56/90 documents · 95% CI 51.9 to 71.5%
avg 92.0%
med 98.8%
no matched human corpus exists for this condition
all lengths
43.5% 188/432 documents · 95% CI 38.9 to 48.2%
avg 78.9%
med 98.1%
no matched human corpus exists for this condition

Adversarial arm. These rates do not describe ordinary short copy and must not be quoted as if they did.

What Table 1 says

The headline is a long-document figure. It describes a corpus whose median AI document is 1,612 words. At 600–849 words the same detector reads 46 of 52 on that corpus. Someone pasting an 800-word blog post is not getting the headline rate, and the tool says so on the result rather than only here.

The one human band that looks bad is a register effect, not a length effect. The 850–1,199 band reads 30 of 1,050, about three times the corpus average. Twenty-two of those thirty are human fiction, whose median length is 1,190 words, so most of the corpus's human stories land in that single band. Excluding fiction the band is in line with the corpus as a whole. Read it as the fiction weakness reappearing where fiction happens to sit on the axis.

Above 3,061 words there is no AI data at all. That is the longest AI document any corpus in this project contains. The 3,500–4,999 and 5,000-and-above bands hold human documents only, both below the 30-document floor. No detection rate should be inferred above 3,061 words. The server refuses documents over 4,000 words, so the unmeasured range starts inside what it will accept.

The average human probability climbs with length while the false-positive rate does not. That is arithmetic, not drift: the reported probability is the maximum across sections, and a longer document offers more sections for the maximum to be drawn from. It is why the two populations are never averaged together, and why the median is printed beside the average.

04 / Table 2

By the model that wrote the text

Read as: of what this model writes, how much is flagged, and how confident is the detector when it reads it. There is no false-positive column here and there cannot be: a false positive belongs to a human writer, not to a model. The human false-positive rate at this same operating point is in Tables 1 and 3.

Figure 4 Detection by the model that generated the text, long-form corpus Ordered as the table is. Six models fall below the 30-document floor and are drawn with no bar at all, because thirteen documents estimate nothing; their counts are printed in their place and in the table.
AI detection rate by the model that generated the text 0% 25% 50% 75% 100% openai/gpt-5.6-luna 100.0% mistralai/mistral-medium-3-5 100.0% anthropic/claude-sonnet-5 26/26 — n below 30, no rate quotedmoonshotai/kimi-k3 26/26 — n below 30, no rate quotedgoogle/gemini-3.1-pro-preview 21/21 — n below 30, no rate quotedopenai/gpt-5.6-sol-pro 13/13 — n below 30, no rate quotedqwen/qwen3.8-max 99.1% z-ai/glm-5.3 98.5% deepseek/deepseek-v4-pro-0813 97.7% google/gemini-3.7-flash 97.5% x-ai/grok-4.6 93.4% anthropic/claude-opus-5 21/23 — n below 30, no rate quotedmeta-llama/llama-4-maverick 79.2% all models 95.8%
AI detection rate by the model that generated the text
Model Flagged, with 95% interval AI probability, avg / median
openai/gpt-5.6-luna
100.0% 121/121 documents · 95% CI 96.9 to 100.0%
avg 99.0%
med 99.0%
mistralai/mistral-medium-3-5
100.0% 41/41 documents · 95% CI 91.4 to 100.0%
avg 99.0%
med 99.0%
anthropic/claude-sonnet-5
26/26n below 30 — no rate quoted
avg 99.0%
med 99.0%
moonshotai/kimi-k3
26/26n below 30 — no rate quoted
avg 99.0%
med 99.0%
google/gemini-3.1-pro-preview
21/21n below 30 — no rate quoted
avg 99.0%
med 99.0%
openai/gpt-5.6-sol-pro
13/13n below 30 — no rate quoted
avg 99.0%
med 99.0%
qwen/qwen3.8-max
99.1% 109/110 documents · 95% CI 95.0 to 99.8%
avg 99.0%
med 99.0%
z-ai/glm-5.3
98.5% 66/67 documents · 95% CI 92.0 to 99.7%
avg 99.0%
med 99.0%
deepseek/deepseek-v4-pro-0813
97.7% 128/131 documents · 95% CI 93.5 to 99.2%
avg 98.9%
med 99.0%
google/gemini-3.7-flash
97.5% 118/121 documents · 95% CI 93.0 to 99.2%
avg 98.9%
med 99.0%
x-ai/grok-4.6
93.4% 113/121 documents · 95% CI 87.5 to 96.6%
avg 98.8%
med 99.0%
anthropic/claude-opus-5
21/23n below 30 — no rate quoted
avg 98.5%
med 99.0%
meta-llama/llama-4-maverick
79.2% 80/101 documents · 95% CI 70.3 to 86.0%
avg 98.4%
med 98.7%
all models
95.8% 883/922 documents · 95% CI 94.3 to 96.9%
avg 98.9%
med 99.0%

Model rank and document length are confounded here — see the reading note below the table.

Figure 5 Detection by provider, long-form corpus The same corpus rolled up to the provider that served each model. One provider row is below the floor and prints no rate.
AI detection rate by provider 0% 25% 50% 75% 100% openai 100.0% mistral 100.0% moonshot 26/26 — n below 30, no rate quotedqwen 99.1% zai 98.5% google 97.9% deepseek 97.7% anthropic 95.9% xai 93.4% meta 79.2% all providers 95.8%
Show the numbers this was drawn from
AI detection rate by provider
Provider Flagged, with 95% interval AI probability, avg / median
openai
100.0% 134/134 documents · 95% CI 97.2 to 100.0%
avg 99.0%
med 99.0%
mistral
100.0% 41/41 documents · 95% CI 91.4 to 100.0%
avg 99.0%
med 99.0%
moonshot
26/26n below 30 — no rate quoted
avg 99.0%
med 99.0%
qwen
99.1% 109/110 documents · 95% CI 95.0 to 99.8%
avg 99.0%
med 99.0%
zai
98.5% 66/67 documents · 95% CI 92.0 to 99.7%
avg 99.0%
med 99.0%
google
97.9% 139/142 documents · 95% CI 94.0 to 99.3%
avg 98.9%
med 99.0%
deepseek
97.7% 128/131 documents · 95% CI 93.5 to 99.2%
avg 98.9%
med 99.0%
anthropic
95.9% 47/49 documents · 95% CI 86.3 to 98.9%
avg 98.8%
med 99.0%
xai
93.4% 113/121 documents · 95% CI 87.5 to 96.6%
avg 98.8%
med 99.0%
meta
79.2% 80/101 documents · 95% CI 70.3 to 86.0%
avg 98.4%
med 98.7%
all providers
95.8% 883/922 documents · 95% CI 94.3 to 96.9%
avg 98.9%
med 99.0%

Short-form, per model

The short-form corpus was generated from two OpenAI models, so it supports a two-row table and no provider comparison at all. The ordinary arm first, then the keyword-repetition arm.

Short-form AI detection rate by model, ordinary arm
Model, short-form arm Flagged, with 95% interval AI probability, avg / median
openai/gpt-5.6-sol
70.6% 283/401 documents · 95% CI 65.9 to 74.8%
avg 93.6%
med 98.9%
openai/gpt-5.6-luna
68.2% 283/415 documents · 95% CI 63.6 to 72.5%
avg 94.2%
med 98.8%
all models
69.4% 566/816 documents · 95% CI 66.1 to 72.4%
avg 93.9%
med 98.9%
Short-form AI detection rate by model, keyword-repetition arm
Model, keyword-repetition arm Flagged, with 95% interval AI probability, avg / median
openai/gpt-5.6-sol
44.2% 88/199 documents · 95% CI 37.5 to 51.2%
avg 76.2%
med 98.0%
openai/gpt-5.6-luna
42.9% 100/233 documents · 95% CI 36.7 to 49.3%
avg 81.2%
med 98.1%
all models
43.5% 188/432 documents · 95% CI 38.9 to 48.2%
avg 78.9%
med 98.1%

What Table 2 says

Adding the probability distribution changes how the detection column should be read. Every model in the long-form corpus sits at a median probability inside a range narrower than half a percentage point. The detector is not less certain about the models it catches less often; it is equally certain about most of their documents and loses a tail.

The average and the median can disagree, which is why both are printed. Where a handful of documents drag an average below the flag point while the median sits above it, the average alone would have been misleading about that model.

Six rows are below the 30-document floor and print counts only. They are shown because omitting them would hide the corpus, not because thirteen documents estimate anything. A count of 21 out of 23 is not a percentage in any sense that should travel.

Model rank and document length are confounded here, and the confound runs the wrong way for a clean reading. The two lowest-scoring models also wrote the shortest documents. A per-model-by-length cross-tabulation was not produced: at 922 documents across 13 models and 8 bands the cells would average nine documents, below the floor this page applies everywhere else. The honest statement is that Table 2's ordering is partly a length ordering, and separating the two would need a length-balanced generation run this project has not done.

05 / Table 3

By content type

Read as: of writing of this kind, how much AI writing is flagged and how much human writing is wrongly flagged. Content types are the corpus's own labels. The rows are ordered by human false-positive rate, worst first, because that is the column with a person on the end of it: a missed AI document is a tool being unhelpful, and a wrongly flagged human document is a writer being accused.

Fiction is roughly four and a half times worse than the next worst content type, and it is the only row anywhere near one document in ten.

Everything below it sits under 2%. Alphabetical order would have buried that in the middle of the table; ranked order puts it where a novelist will see it before they trust a result.
Figure 6 Human writing wrongly flagged, by content type The axis runs 0 to 10%, not 0 to 100%, because every row but one sits under 2% and a full axis would show eleven flat lines. Read the numbers, not the bar lengths, when comparing this figure with Figures 1 to 5. The AI detection rate for each content type is in the table below.
Human false-positive rate by content type, ordered worst first 0% 2.5% 5% 7.5% 10% stories / fiction 8.8% academic conclusions 1.9% academic discussion 1.9% long-form journalism 0.4% academic introductions 0.2% white papers 0.2% company updates 0.2% academic literature reviews 0.0% research summaries 0.0% student essays 0.0% academic essays no documents
Human documents wrongly flagged, 0–10% axis
AI detection rate and human false-positive rate by content type, ordered by false-positive rate, worst first
Content type AI documents Human documents
Flagged, with 95% interval AI probability, avg / median Wrongly flagged, with 95% interval AI probability, avg / median
stories / fiction
93.9% 107/114 documents · 95% CI 87.9 to 97.0%
avg 98.7%
med 98.9%
8.8% 23/260 documents · 95% CI 6.0 to 12.9%
avg 44.6%
med 35.4%
academic conclusions
no documents
avg
med
1.9% 7/360 documents · 95% CI 0.9 to 4.0%
avg 56.7%
med 61.8%
academic discussion
95.6% 108/113 documents · 95% CI 90.1 to 98.1%
avg 98.9%
med 99.0%
1.9% 8/420 documents · 95% CI 1.0 to 3.7%
avg 67.5%
med 77.5%
long-form journalism
96.4% 132/137 documents · 95% CI 91.7 to 98.4%
avg 98.9%
med 99.0%
0.4% 3/840 documents · 95% CI 0.1 to 1.0%
avg 41.3%
med 35.8%
academic introductions
no documents
avg
med
0.2% 1/420 documents · 95% CI 0.0 to 1.3%
avg 57.8%
med 65.1%
white papers
99.0% 102/103 documents · 95% CI 94.7 to 99.8%
avg 99.0%
med 99.0%
0.2% 2/840 documents · 95% CI 0.1 to 0.9%
avg 41.6%
med 32.3%
company updates
100.0% 99/99 documents · 95% CI 96.3 to 100.0%
avg 99.0%
med 99.0%
0.2% 1/662 documents · 95% CI 0.0 to 0.9%
avg 27.5%
med 14.2%
academic literature reviews
94.4% 101/107 documents · 95% CI 88.3 to 97.4%
avg 98.9%
med 99.0%
0.0% 0/225 documents · 95% CI 0.0 to 1.7%
avg 51.0%
med 53.5%
research summaries
98.3% 115/117 documents · 95% CI 94.0 to 99.5%
avg 99.0%
med 99.0%
0.0% 0/189 documents · 95% CI 0.0 to 2.0%
avg 53.7%
med 53.2%
student essays
no documents
avg
med
0.0% 0/420 documents · 95% CI 0.0 to 0.9%
avg 22.8%
med 14.4%
academic essays
90.2% 119/132 documents · 95% CI 83.9 to 94.2%
avg 98.7%
med 99.0%
no documents
avg
med
all registers
95.8% 883/922 documents · 95% CI 94.3 to 96.9%
avg 98.9%
med 99.0%
1.0% 45/4,636 documents · 95% CI 0.7 to 1.3%
avg 43.9%
med 37.7%

Ordered by the human false-positive rate, worst first — the same order the table uses, and the ordering is itself the finding.

What the ordering reveals

And the contrast at the other end is worth saying out loud. Student essays are 0 of 420 — the content type with the most at stake for a real person, since a wrongly flagged essay is an academic misconduct allegation, is the safest row in the table. Academic literature reviews and research summaries are also 0. The tool is at its most reliable exactly where being wrong would cost the most, and at its least reliable on creative writing, where the model was never given a matched human corpus to learn from. That asymmetry is luck of the training data rather than design, and it is stated so that neither half is mistaken for a general property of the tool.

Why fiction fails. The model was never trained on human fiction. The training corpus holds AI fiction samples and no matched human set, and training on unmatched AI fiction would have taught it that fiction equals AI. That is an explanation, not an excuse. Fiction also sits in a dense cliff just below the flag point, so its rate moves faster with the threshold than any other content type and has to be read separately from the corpus rate. Novelists should not rely on this tool.

Three rows have no AI documents and one has no human documents. Academic conclusions, academic introductions and student essays exist in the human corpus only, so their AI cells say "no documents" rather than showing a rate; academic essays exist in the AI corpus only. That is the shape of the corpora, not a gap in the measurement, and it is printed rather than hidden by dropping the rows.

Detection and false positives do not trade off cleanly by content type. Company updates and white papers are near the top on detection and near the bottom on false positives at the same time. Fiction is the only content type that is poor on both counts, which is consistent with the model simply not having seen it.

06 / Everything else

The rest of the measurement record

These three tables are one cut of the evidence. The full reports, including the ones that record results less flattering than these, are in the open repository:

5,558 long-form documents: 922 written by 13 current AI models, 4,636 written by people, from Europe PMC, GOV.UK, CRS, Global Voices, Mongabay, SEC EDGAR and PERSUADE 2.0. 654 of the AI documents are independent of every training, test and calibration split; 268 are not, as are 11 of the human documents.

Use it with the limits attached

A rate describes a corpus. Your draft is one document.