On this page
01 / Read this first
What these tables are, and what they are not
They are measurements on a corpus that is mostly, but not entirely, independent of the model. Every document was scored through the same path the EU inference server runs, at the operating point that ships today, and every cell states how many documents it counted and how many were flagged, with a 95% confidence interval. Until 30 August 2026 this page called the corpus held out and said the model had never seen it. That was wrong for the AI half, and the correction is in the provenance block below rather than folded away.
They are not a guarantee about any individual document. A rate describes a population. It does not tell you the chance that the specific draft in front of you was written by a machine, and no number on this page should be used to make an accusation about a person. The checker is evidence for review: no score it produces identifies an author, and a quiet result does not establish who wrote the text.
They describe one check. The trained classifier is the only check in this tool that gives an AI reading. The hidden-character, lookalike, protected-fact, writing-signal and watermark checks are separate, are reported separately, and are not measured here.
They describe the server route. Every figure below is fp32 through Python onnxruntime, which is the reference-server scoring path. The in-browser route runs the same model in lower numerical precision; the two agree to a median of 0.0002 in the decision region, but a proxy is not the thing, and the browser's own full-corpus curve at this operating point is measured separately rather than borrowed from here.
The worst register measured, at the shipped operating point (0.9855/0.9763), over all 260 human fiction documents. Fiction sits in a dense cliff just below the flag point, so this rate moves faster with the threshold than any other register and has to be looked at separately from the corpus rate. See Table 3.
02 / Provenance
What produced every cell below
Correction, 30 August 2026: this corpus is not fully held out.
654 of the 922 AI documents are independent of every cycle-2 split. 268 are not, 168 of them in the training split. Every rate on this page pools both subsets, and no seen-against-unseen split has been measured at the operating point that ships.
The full record, with what was measured and at which flag point
CORRECTED 30 August 2026: this corpus was published as held out and hash-quarantined against every training split, and it is not. 654 of the 922 AI documents are independent of every cycle-2 split; 268 are not (168 training, 72 test, 28 calibration), as are 11 of the 4,636 human documents. The difference has been measured only at the superseded 0.984 single-threshold rule, where the independent subset reads 620/654 (94.80%) against 257/268 (95.90%) for the seen subset, a gap of 1.1 points. It has NOT been measured at the shipped 0.9855/0.9763 pair, so every rate on this page pools both subsets and no seen-against-unseen split is published for the operating point that ships.
What the difference measures, and at which flag point. The split has been scored at one operating point only, and it is not the one that ships. At the superseded 0.984 single-threshold rule the independent subset reads 620/654, 94.80%, against 257/268, 95.90% for the subset the model had seen, of which the training split alone reads 163/168, 97.02% — a gap of 1.1 percentage points. No seen-against-unseen split has been measured at the shipped 0.9855/0.9763 pair, so none is published for it, and the 1.1-point figure above must not be read as if it described the shipped rule. A re-measurement at the shipped pair is outstanding.
Source: services/local-engine/research/corpus-reconciliation-2026-08-29/analysis.txt, section 2, in the open measurement repository.
Percentages in the probability columns are the model's calibrated AI probability expressed as a percentage: a document that scored 0.9886 reads 98.9%. AI documents and human documents are counted, and their probabilities averaged, separately. A single average across both populations would describe neither.
03 / Table 1
By document length
Read as: of documents this long, how many were flagged, and what did the two populations score. Two corpora cover this range and they overlap between 400 and 849 words, so each is given its own figure rather than pooled: they were generated differently, by different models, in different registers.
Length is the dominant axis, and the cliff is below 200 words.
| Words | AI documents | Human documents | ||
|---|---|---|---|---|
| Flagged, with 95% interval | AI probability, avg / median | Wrongly flagged, with 95% interval | AI probability, avg / median | |
| 400–599 | 3/3n below 30 — no rate quoted | avg 98.8% med 98.8% | 0.4% 3/826 documents · 95% CI 0.1 to 1.1% | avg 40.5% med 30.8% |
| 600–849 | 88.5% 46/52 documents · 95% CI 77.0 to 94.6% | avg 98.7% med 98.8% | 0.2% 2/894 documents · 95% CI 0.1 to 0.8% | avg 41.4% med 35.5% |
| 850–1,199 | 90.7% 175/193 documents · 95% CI 85.7 to 94.0% | avg 98.8% med 99.0% | 2.9% 30/1,050 documents · 95% CI 2.0 to 4.0% | avg 47.4% med 45.5% |
| 1,200–1,699 | 97.7% 259/265 documents · 95% CI 95.1 to 99.0% | avg 98.9% med 99.0% | 0.2% 3/1,237 documents · 95% CI 0.1 to 0.7% | avg 36.9% med 26.5% |
| 1,700–2,399 | 97.1% 271/279 documents · 95% CI 94.4 to 98.5% | avg 98.9% med 99.0% | 1.3% 6/461 documents · 95% CI 0.6 to 2.8% | avg 57.2% med 63.3% |
| 2,400–3,499 | 99.2% 129/130 documents · 95% CI 95.8 to 99.9% | avg 98.9% med 99.0% | 0.0% 0/134 documents · 95% CI 0.0 to 2.8% | avg 65.7% med 75.3% |
| 3,500–4,999 | no documents | avg — med — | 0/21n below 30 — no rate quoted | avg 65.8% med 69.9% |
| 5,000 and above | no documents | avg — med — | 1/13n below 30 — no rate quoted | avg 90.5% med 96.6% |
| all lengths | 95.8% 883/922 documents · 95% CI 94.3 to 96.9% | avg 98.9% med 99.0% | 1.0% 45/4,636 documents · 95% CI 0.7 to 1.3% | avg 43.9% med 37.7% |
Every point is drawn from the cell printed in the table below it. A band with no rate draws no point.
| Words | AI documents | Human documents | ||
|---|---|---|---|---|
| Flagged, with 95% interval | AI probability, avg / median | Wrongly flagged, with 95% interval | AI probability, avg / median | |
| under 100 | 6/27n below 30 — no rate quoted | avg 82.5% med 94.2% | 0.0% 0/388 documents · 95% CI 0.0 to 1.0% | avg 17.8% med 10.7% |
| 100–199 | 16.9% 29/172 documents · 95% CI 12.0 to 23.2% | avg 84.6% med 97.3% | 0.0% 0/732 documents · 95% CI 0.0 to 0.5% | avg 16.8% med 9.6% |
| 200–299 | 16/19n below 30 — no rate quoted | avg 97.3% med 98.9% | 0.4% 3/822 documents · 95% CI 0.1 to 1.1% | avg 21.0% med 8.8% |
| 300–399 | 84.6% 193/228 documents · 95% CI 79.4 to 88.8% | avg 96.9% med 98.9% | 0.8% 9/1,170 documents · 95% CI 0.4 to 1.5% | avg 20.6% med 7.7% |
| 400–599 | 80.9% 161/199 documents · 95% CI 74.9 to 85.8% | avg 96.0% med 98.9% | 0.9% 10/1,167 documents · 95% CI 0.5 to 1.6% | avg 30.7% med 18.2% |
| 600–849 | 94.2% 161/171 documents · 95% CI 89.6 to 96.8% | avg 98.2% med 99.0% | 0.0% 0/86 documents · 95% CI 0.0 to 4.3% | avg 40.5% med 33.7% |
| all lengths | 69.4% 566/816 documents · 95% CI 66.1 to 72.4% | avg 93.9% med 98.9% | 0.5% 22/4,368 documents · 95% CI 0.3 to 0.8% | avg 22.9% med 11.0% |
A separate corpus from Figure 1, generated differently and by different models. The two overlap between 400 and 849 words and are never pooled.
The keyword-repetition arm, fenced separately
A further 432 AI passages were generated to be deliberately keyword-repetitive, as a test of the evasion axis that matters most commercially. They are an adversarial condition, not a sample of ordinary short text. Pooling them into the figures above would understate the detector against normal copy in the same way that omitting them would overstate it against SEO copy, so they sit here. There is no matched human corpus for this condition; only the AI side exists.
Show the numbers this was drawn from
| Words | Flagged, with 95% interval | AI probability, avg / median | Human documents |
|---|---|---|---|
| under 100 | 2/16n below 30 — no rate quoted | avg 63.6% med 94.2% | no matched human corpus exists for this condition |
| 100–199 | 14.0% 8/57 documents · 95% CI 7.3 to 25.3% | avg 63.0% med 84.2% | no matched human corpus exists for this condition |
| 200–299 | 3/9n below 30 — no rate quoted | avg 77.8% med 93.6% | no matched human corpus exists for this condition |
| 300–399 | 44.9% 61/136 documents · 95% CI 36.7 to 53.2% | avg 76.1% med 98.1% | no matched human corpus exists for this condition |
| 400–599 | 46.8% 58/124 documents · 95% CI 38.2 to 55.5% | avg 81.9% med 98.3% | no matched human corpus exists for this condition |
| 600–849 | 62.2% 56/90 documents · 95% CI 51.9 to 71.5% | avg 92.0% med 98.8% | no matched human corpus exists for this condition |
| all lengths | 43.5% 188/432 documents · 95% CI 38.9 to 48.2% | avg 78.9% med 98.1% | no matched human corpus exists for this condition |
Adversarial arm. These rates do not describe ordinary short copy and must not be quoted as if they did.
What Table 1 says
The headline is a long-document figure. It describes a corpus whose median AI document is 1,612 words. At 600–849 words the same detector reads 46 of 52 on that corpus. Someone pasting an 800-word blog post is not getting the headline rate, and the tool says so on the result rather than only here.
The one human band that looks bad is a register effect, not a length effect. The 850–1,199 band reads 30 of 1,050, about three times the corpus average. Twenty-two of those thirty are human fiction, whose median length is 1,190 words, so most of the corpus's human stories land in that single band. Excluding fiction the band is in line with the corpus as a whole. Read it as the fiction weakness reappearing where fiction happens to sit on the axis.
Above 3,061 words there is no AI data at all. That is the longest AI document any corpus in this project contains. The 3,500–4,999 and 5,000-and-above bands hold human documents only, both below the 30-document floor. No detection rate should be inferred above 3,061 words. The server refuses documents over 4,000 words, so the unmeasured range starts inside what it will accept.
The average human probability climbs with length while the false-positive rate does not. That is arithmetic, not drift: the reported probability is the maximum across sections, and a longer document offers more sections for the maximum to be drawn from. It is why the two populations are never averaged together, and why the median is printed beside the average.
04 / Table 2
By the model that wrote the text
Read as: of what this model writes, how much is flagged, and how confident is the detector when it reads it. There is no false-positive column here and there cannot be: a false positive belongs to a human writer, not to a model. The human false-positive rate at this same operating point is in Tables 1 and 3.
| Model | Flagged, with 95% interval | AI probability, avg / median |
|---|---|---|
openai/gpt-5.6-luna | 100.0% 121/121 documents · 95% CI 96.9 to 100.0% | avg 99.0% med 99.0% |
mistralai/mistral-medium-3-5 | 100.0% 41/41 documents · 95% CI 91.4 to 100.0% | avg 99.0% med 99.0% |
anthropic/claude-sonnet-5 | 26/26n below 30 — no rate quoted | avg 99.0% med 99.0% |
moonshotai/kimi-k3 | 26/26n below 30 — no rate quoted | avg 99.0% med 99.0% |
google/gemini-3.1-pro-preview | 21/21n below 30 — no rate quoted | avg 99.0% med 99.0% |
openai/gpt-5.6-sol-pro | 13/13n below 30 — no rate quoted | avg 99.0% med 99.0% |
qwen/qwen3.8-max | 99.1% 109/110 documents · 95% CI 95.0 to 99.8% | avg 99.0% med 99.0% |
z-ai/glm-5.3 | 98.5% 66/67 documents · 95% CI 92.0 to 99.7% | avg 99.0% med 99.0% |
deepseek/deepseek-v4-pro-0813 | 97.7% 128/131 documents · 95% CI 93.5 to 99.2% | avg 98.9% med 99.0% |
google/gemini-3.7-flash | 97.5% 118/121 documents · 95% CI 93.0 to 99.2% | avg 98.9% med 99.0% |
x-ai/grok-4.6 | 93.4% 113/121 documents · 95% CI 87.5 to 96.6% | avg 98.8% med 99.0% |
anthropic/claude-opus-5 | 21/23n below 30 — no rate quoted | avg 98.5% med 99.0% |
meta-llama/llama-4-maverick | 79.2% 80/101 documents · 95% CI 70.3 to 86.0% | avg 98.4% med 98.7% |
all models | 95.8% 883/922 documents · 95% CI 94.3 to 96.9% | avg 98.9% med 99.0% |
Model rank and document length are confounded here — see the reading note below the table.
Show the numbers this was drawn from
| Provider | Flagged, with 95% interval | AI probability, avg / median |
|---|---|---|
openai | 100.0% 134/134 documents · 95% CI 97.2 to 100.0% | avg 99.0% med 99.0% |
mistral | 100.0% 41/41 documents · 95% CI 91.4 to 100.0% | avg 99.0% med 99.0% |
moonshot | 26/26n below 30 — no rate quoted | avg 99.0% med 99.0% |
qwen | 99.1% 109/110 documents · 95% CI 95.0 to 99.8% | avg 99.0% med 99.0% |
zai | 98.5% 66/67 documents · 95% CI 92.0 to 99.7% | avg 99.0% med 99.0% |
google | 97.9% 139/142 documents · 95% CI 94.0 to 99.3% | avg 98.9% med 99.0% |
deepseek | 97.7% 128/131 documents · 95% CI 93.5 to 99.2% | avg 98.9% med 99.0% |
anthropic | 95.9% 47/49 documents · 95% CI 86.3 to 98.9% | avg 98.8% med 99.0% |
xai | 93.4% 113/121 documents · 95% CI 87.5 to 96.6% | avg 98.8% med 99.0% |
meta | 79.2% 80/101 documents · 95% CI 70.3 to 86.0% | avg 98.4% med 98.7% |
all providers | 95.8% 883/922 documents · 95% CI 94.3 to 96.9% | avg 98.9% med 99.0% |
Short-form, per model
The short-form corpus was generated from two OpenAI models, so it supports a two-row table and no provider comparison at all. The ordinary arm first, then the keyword-repetition arm.
| Model, short-form arm | Flagged, with 95% interval | AI probability, avg / median |
|---|---|---|
openai/gpt-5.6-sol | 70.6% 283/401 documents · 95% CI 65.9 to 74.8% | avg 93.6% med 98.9% |
openai/gpt-5.6-luna | 68.2% 283/415 documents · 95% CI 63.6 to 72.5% | avg 94.2% med 98.8% |
all models | 69.4% 566/816 documents · 95% CI 66.1 to 72.4% | avg 93.9% med 98.9% |
| Model, keyword-repetition arm | Flagged, with 95% interval | AI probability, avg / median |
|---|---|---|
openai/gpt-5.6-sol | 44.2% 88/199 documents · 95% CI 37.5 to 51.2% | avg 76.2% med 98.0% |
openai/gpt-5.6-luna | 42.9% 100/233 documents · 95% CI 36.7 to 49.3% | avg 81.2% med 98.1% |
all models | 43.5% 188/432 documents · 95% CI 38.9 to 48.2% | avg 78.9% med 98.1% |
What Table 2 says
Adding the probability distribution changes how the detection column should be read. Every model in the long-form corpus sits at a median probability inside a range narrower than half a percentage point. The detector is not less certain about the models it catches less often; it is equally certain about most of their documents and loses a tail.
The average and the median can disagree, which is why both are printed. Where a handful of documents drag an average below the flag point while the median sits above it, the average alone would have been misleading about that model.
Six rows are below the 30-document floor and print counts only. They are shown because omitting them would hide the corpus, not because thirteen documents estimate anything. A count of 21 out of 23 is not a percentage in any sense that should travel.
Model rank and document length are confounded here, and the confound runs the wrong way for a clean reading. The two lowest-scoring models also wrote the shortest documents. A per-model-by-length cross-tabulation was not produced: at 922 documents across 13 models and 8 bands the cells would average nine documents, below the floor this page applies everywhere else. The honest statement is that Table 2's ordering is partly a length ordering, and separating the two would need a length-balanced generation run this project has not done.
05 / Table 3
By content type
Read as: of writing of this kind, how much AI writing is flagged and how much human writing is wrongly flagged. Content types are the corpus's own labels. The rows are ordered by human false-positive rate, worst first, because that is the column with a person on the end of it: a missed AI document is a tool being unhelpful, and a wrongly flagged human document is a writer being accused.
Fiction is roughly four and a half times worse than the next worst content type, and it is the only row anywhere near one document in ten.
| Content type | AI documents | Human documents | ||
|---|---|---|---|---|
| Flagged, with 95% interval | AI probability, avg / median | Wrongly flagged, with 95% interval | AI probability, avg / median | |
| stories / fiction | 93.9% 107/114 documents · 95% CI 87.9 to 97.0% | avg 98.7% med 98.9% | 8.8% 23/260 documents · 95% CI 6.0 to 12.9% | avg 44.6% med 35.4% |
| academic conclusions | no documents | avg — med — | 1.9% 7/360 documents · 95% CI 0.9 to 4.0% | avg 56.7% med 61.8% |
| academic discussion | 95.6% 108/113 documents · 95% CI 90.1 to 98.1% | avg 98.9% med 99.0% | 1.9% 8/420 documents · 95% CI 1.0 to 3.7% | avg 67.5% med 77.5% |
| long-form journalism | 96.4% 132/137 documents · 95% CI 91.7 to 98.4% | avg 98.9% med 99.0% | 0.4% 3/840 documents · 95% CI 0.1 to 1.0% | avg 41.3% med 35.8% |
| academic introductions | no documents | avg — med — | 0.2% 1/420 documents · 95% CI 0.0 to 1.3% | avg 57.8% med 65.1% |
| white papers | 99.0% 102/103 documents · 95% CI 94.7 to 99.8% | avg 99.0% med 99.0% | 0.2% 2/840 documents · 95% CI 0.1 to 0.9% | avg 41.6% med 32.3% |
| company updates | 100.0% 99/99 documents · 95% CI 96.3 to 100.0% | avg 99.0% med 99.0% | 0.2% 1/662 documents · 95% CI 0.0 to 0.9% | avg 27.5% med 14.2% |
| academic literature reviews | 94.4% 101/107 documents · 95% CI 88.3 to 97.4% | avg 98.9% med 99.0% | 0.0% 0/225 documents · 95% CI 0.0 to 1.7% | avg 51.0% med 53.5% |
| research summaries | 98.3% 115/117 documents · 95% CI 94.0 to 99.5% | avg 99.0% med 99.0% | 0.0% 0/189 documents · 95% CI 0.0 to 2.0% | avg 53.7% med 53.2% |
| student essays | no documents | avg — med — | 0.0% 0/420 documents · 95% CI 0.0 to 0.9% | avg 22.8% med 14.4% |
| academic essays | 90.2% 119/132 documents · 95% CI 83.9 to 94.2% | avg 98.7% med 99.0% | no documents | avg — med — |
| all registers | 95.8% 883/922 documents · 95% CI 94.3 to 96.9% | avg 98.9% med 99.0% | 1.0% 45/4,636 documents · 95% CI 0.7 to 1.3% | avg 43.9% med 37.7% |
Ordered by the human false-positive rate, worst first — the same order the table uses, and the ordering is itself the finding.
What the ordering reveals
And the contrast at the other end is worth saying out loud. Student essays are 0 of 420 — the content type with the most at stake for a real person, since a wrongly flagged essay is an academic misconduct allegation, is the safest row in the table. Academic literature reviews and research summaries are also 0. The tool is at its most reliable exactly where being wrong would cost the most, and at its least reliable on creative writing, where the model was never given a matched human corpus to learn from. That asymmetry is luck of the training data rather than design, and it is stated so that neither half is mistaken for a general property of the tool.
Why fiction fails. The model was never trained on human fiction. The training corpus holds AI fiction samples and no matched human set, and training on unmatched AI fiction would have taught it that fiction equals AI. That is an explanation, not an excuse. Fiction also sits in a dense cliff just below the flag point, so its rate moves faster with the threshold than any other content type and has to be read separately from the corpus rate. Novelists should not rely on this tool.
Three rows have no AI documents and one has no human documents. Academic conclusions, academic introductions and student essays exist in the human corpus only, so their AI cells say "no documents" rather than showing a rate; academic essays exist in the AI corpus only. That is the shape of the corpora, not a gap in the measurement, and it is printed rather than hidden by dropping the rows.
Detection and false positives do not trade off cleanly by content type. Company updates and white papers are near the top on detection and near the bottom on false positives at the same time. Fiction is the only content type that is poor on both counts, which is consistent with the model simply not having seen it.
06 / Everything else
The rest of the measurement record
These three tables are one cut of the evidence. The full reports, including the ones that record results less flattering than these, are in the open repository:
- Detection by length and by model — The source report for Tables 1 and 2, with the harness reproduction that licenses them.
- Aggregation, rhythm and the flag rule — Why the strongest section decides and nothing is averaged, and how the two-section rule was fitted.
- The segmentation token fix — 1,348 of 23,318 sections were silently overflowing the model's window. What that cost, and what it reads now.
- Short-form corpus and retrain — Where the short-form and keyword-repetition arms come from, including the type-token-ratio cliff.
- Route parity, server against browser — How far the fp32 and int8 runtimes disagree, and why both share one flag point.
- Per-model detection at the earlier flag point — The same corpus through the prior 0.984 rule. Two operating points: do not place a row from one beside a row from the other.
- Measured findings — Prompt-style evasion, register beating model choice, sentence occlusion, and one figure withdrawn.
- Test evidence index — Every published claim traced to the run, report or research document that produced it.
- The watermark lab — SynthID-Text detection mathematics, the wrong-key collapse, and what paraphrase does to a mark.
5,558 long-form documents: 922 written by 13 current AI models, 4,636 written by people, from Europe PMC, GOV.UK, CRS, Global Voices, Mongabay, SEC EDGAR and PERSUADE 2.0. 654 of the AI documents are independent of every training, test and calibration split; 268 are not, as are 11 of the human documents.