Get in Touch
AI content toolsfrom Opace

Measurement paper · 5,558 documents

The corpus, the quarantine, and the file it missed

Every accuracy figure this tool publishes is measured on 5,558 long-form documents: 4,636 from licensed human sources and 922 generated through OpenRouter. Each candidate was hashed against an index of material it had to be independent of. The index held four files. The training file was not one of them.

On this page
  1. The finding
  2. What is in it, and under what licence
  3. What was refused
  4. The machine-written half
  5. The quarantine, as designed
  6. What it was pointed at
  7. What the leak was worth
  8. What this corpus does not cover
  9. Where every figure comes from
Published
30 August 2026
Measured
Corpus built 28 August 2026; contamination measured 29 August 2026; re-verified against source 30 August 2026
Corpus
5,558 long-form documents: 4,636 human, 922 AI from 13 models
Independence
654 of 922 AI documents are independent of every cycle-2 split; 268 are not. Human: 11 of 4,636.

01 / The finding

The quarantine was well designed and pointed at the wrong file list

Every accuracy figure this tool publishes is measured on 5,558 documents assembled for that purpose: 4,636 human documents drawn under an explicit licence, and 922 machine-written documents generated through OpenRouter after the shipped model had finished training. Each candidate was hashed against the material it had to be independent of, with an exact collision set to abort the build.

The index held 6,114 hashes from four held-out files, and the build returned zero exact collisions against all four. cycle2-train/dataset.jsonl was not one of the four.

Measured afterwards on normalised SHA-256: 268 of the 922 machine-written documents — 29.1% — appear in the cycle-2 dataset, 168 of them in the training split. The human half is effectively clean at 11 of 4,636. The corpus manifest publishes the four per-file counts and no total; 6,114 is their sum, computed here.
Human half, licensed sources4,636

Eight source pools, each row carrying its licence string. Where two authorities disagree on a licence, both are recorded rather than resolved.

Machine half, after the model trained922

13 current models through OpenRouter, 800–2,000 word targets, eight long-form registers, three prompt styles.

Not independent of cycle 2268 of 922

168 in the train split, 72 in test, 28 in calibration. The independent subset is 654 documents, and that is the population any unseen-data argument has to use.

That belongs at the top of this page rather than in a footnote at the bottom of it. The design was sound: normalise, hash, abort on a collision. The list it checked against was incomplete, which is a more useful lesson than a clean result would have been, and it is why the phrase “hash-quarantined against every training split” has been withdrawn from this project’s records and corrected at the head of the file it appeared in.

The size of the effect is the second half of the finding, and it cuts the other way. Measured directly, seeing the training data was worth roughly one percentage point. That is small. It is also not the correction: a figure published as a measurement on unseen documents was, for 29.1% of the machine-written half, a measurement on documents the model had fitted on.

02 / Composition

What is in it, and under what licence

Every row carries its licence string. The human sources named in the shipped measurement file are Europe PMC, GOV.UK, CRS, Global Voices, Mongabay, SEC EDGAR and PERSUADE 2.0, and the eighth pool is the Internet Archive’s Creative Commons text collection, taken per item under whatever the uploader declared. PERSUADE 2.0 is the one row with a conflict: the dataset is CC BY 4.0 upstream and the mirror used declares MIT. Both are recorded, because they differ.

Figure 1 The human half by source, with the licence each row was taken under Document counts, not rates: the value axis runs to 1,500 and no percentage is printed, because a share of the human half is not what any of these rows was collected to measure. Eight source pools totalling 4,636 documents. Per-source and per-register capping limits how far one publisher can dominate a register: at most three passages per SEC filing, at most four items per Internet Archive uploader.
Human-side document counts by source, with the licence each source was taken under 0 300 600 900 1,200 1,500 Europe PMC open access CC BY family, or CC0 1,425 GOV.UK research and policy Open Government Licence v3.0 851 Congressional Research Service US government work, 17 U.S.C. 105 420 Global Voices CC BY 3.0 420 Mongabay CC BY-ND 4.0, stored verbatim 420 SEC EDGAR 10-K Item 7 mandatory public disclosure 420 PERSUADE 2.0 student essays CC BY 4.0 upstream, MIT on the mirror used 420 Internet Archive CC texts per item, the uploader’s licenceurl 260
Human-side composition by source, licence and register
SourceDocumentsLicenceRegisters
Europe PMC open access1,425CC BY family, or CC0academic introduction, literature review, discussion, conclusion
GOV.UK research and policy851Open Government Licence v3.0white paper, research summary, company update
Congressional Research Service420US government work, 17 U.S.C. 105white paper
Global Voices420CC BY 3.0long-form journalism
Mongabay420CC BY-ND 4.0, stored verbatimlong-form journalism
SEC EDGAR 10-K Item 7420mandatory public disclosurecompany update
PERSUADE 2.0 student essays420CC BY 4.0 upstream, MIT on the mirror usedstudent essay
Internet Archive CC texts260per item, the uploader’s licenceurlstory
Total, human half4,636eight licences, one of them disputedeight long-form registers

longform-corpus/MANIFEST.md. Corpus built 28 August 2026. Counts are documents after de-duplication, capping and the prose test; no detector, no runtime and no operating point applies to any value here.

The exclusions show the judgement as clearly as the inclusions. Abstracts were dropped from Europe PMC because they are formulaic and are not what anyone pastes into a detector. PDF-only GOV.UK publications were skipped rather than OCR-guessed. Financial tables were stripped from 10-K filings before chunking, and any passage still mostly figures failed the prose test.

03 / What was refused

Four sources were examined and turned down

The Conversation would have been the single best fit for long-form journalism written by academics. Its republishing terms state that its Creative Commons licence prohibits using its content as training data for AI systems, so it is not here. Strange Horizons reserves all rights on behalf of its authors. BAWE is registration-gated with no redistribution right, and ICLE is commercially licensed.

Excluding the best-fitting source on the publisher’s own terms is the largest single loss in the build, and it is visible downstream: the journalism register rests on two outlets rather than three, and the student-essay register rests on one dataset.

04 / The machine half

922 documents from 13 models, generated after the model was trained

The machine-written half was generated through OpenRouter with 800–2,000 word targets, across eight long-form registers and three prompt styles, from a fixed bank of 68 topics. It cost $12.33 of the $13 authorised, and the per-call cost is stored on every machine-written row so the total is recomputable from the delivered file rather than taken on trust.

Figure 2 The machine half by model Document counts only. No detection rate is plotted here: per-model detection is a separate measurement with its own denominators and its own operating point, and putting it beside a generation count would invite the two to be read as one figure. 922 documents from 13 models. Model mix is a function of what OpenRouter served inside the budget, not a designed balance.
Machine-written document counts by model 0 35 70 105 140 deepseek-v4-pro-0813 131 openai/gpt-5.6-luna 121 google/gemini-3.7-flash 121 x-ai/grok-4.6 121 qwen/qwen3.8-max 110 meta-llama/llama-4-maverick 101 z-ai/glm-5.3 67 mistralai/mistral-medium-3-5 41 anthropic/claude-sonnet-5 26 moonshotai/kimi-k3 26 anthropic/claude-opus-5 23 google/gemini-3.1-pro-preview 21 openai/gpt-5.6-sol-pro 13
Show the per-model counts
Machine-written document counts by model identifier
ModelDocuments
deepseek-v4-pro-0813131
openai/gpt-5.6-luna121
google/gemini-3.7-flash121
x-ai/grok-4.6121
qwen/qwen3.8-max110
meta-llama/llama-4-maverick101
z-ai/glm-5.367
mistralai/mistral-medium-3-541
anthropic/claude-sonnet-526
moonshotai/kimi-k326
anthropic/claude-opus-523
google/gemini-3.1-pro-preview21
openai/gpt-5.6-sol-pro13
Total, machine half922

docs/measurements/DETECTION-BY-LENGTH-AND-MODEL.md Table 2, denominators column. Generated 28 August 2026, after the cycle-2 model finished training. Counts only: no detector, no runtime and no operating point applies to any value here.

Three models contribute under 30 documents each. Any per-model figure drawn from those rows carries an interval wide enough to make a ranking meaningless, which is why per-model detection is published with its denominators on the detection-rates page rather than summarised here.

05 / The quarantine

The quarantine, as designed

Every candidate document is hashed on normalised text: NFKC, lower-cased, punctuation stripped, whitespace collapsed. That hash is checked against an index of material the corpus must be independent of, and an exact collision aborts the build, because an exact collision is proof rather than evidence.

Near-duplicates are screened separately, at 12-word shingle overlap with a 25% threshold. Those rows are dropped and listed in the manifest rather than aborting the build. A shingle hit is a heuristic, and a build that halts on a heuristic will be disabled by whoever runs it next.

Splits are group-aware: 60/15/25 by SHA-256 of a group key, with every variant of a source sitting in the same split as its source. In the cycle-3 build the group key is the post slug, chosen so that one key covers both “same source URL” and “same topic”.

The guard is not decorative, and there is a measurement showing it fires. The cycle-4 corpus build checked every candidate against the normalised hashes of all 11,004 documents in five measurement sets and their 2,684 source references, and caught and excluded three rows: duplicate uploads of the same public-domain work. That build asserts if more than 25 rows are caught, on the reasoning that a large number would mean the exclusion-by-source step had failed rather than that coincidence had happened.

Figure 3 The pipeline every candidate document passed through, and the file it was never checked against Six stages, left column, in order. The box on the right is the index the exact-collision check was run against: four files, 6,114 normalised hashes, zero collisions. The dashed box below it is cycle2-train/dataset.jsonl, drawn outside the index because that is where it sat. No candidate was ever hashed against it. The full pipeline is written out in order in the table beneath.
The corpus build pipeline, its hash index, and the training file left out of that index Six pipeline stages run down the left: fetch, normalise, SHA-256, exact-collision check against the index, a twelve-word shingle screen at twenty-five per cent, and a group-aware sixty fifteen twenty-five split. The exact-collision check points right to a box listing the four indexed files and their hash counts, totalling 6,114, against which the build returned zero collisions. Below that box, outside it and drawn with a dashed outline, is cycle2-train/dataset.jsonl, marked as absent from the index; 268 of the 922 machine-written documents were later found in it. 1 Fetch One candidate document from a licensed source, with its licence string attached 2 Normalise NFKC · lower-case · punctuation stripped · whitespace collapsed 3 SHA-256 One norm_sha256 per document. Hashes are compared, never the raw text 4 Exact-collision check Against the index →. A collision aborts the build: it is proof, not evidence 5 Shingle screen 12-word shingles at 25% overlap. Drops the row and logs it. Never fatal 6 Group-aware split 60/15/25 by SHA-256 of the group key; every variant lands with its source RESULT AGAINST THOSE FOUR FILES 0 exact collisions 2 near-duplicates dropped at the shingle screen 505 internal duplicates dropped The mechanism worked. An exact collision is proof, so it aborts the whole build. A shingle hit is a heuristic, so it drops one row and is logged. Every candidate was checked against everything in the index. What follows is not a failure of the check. It is a failure of the list. The index this build was pointed at Four held-out files · 6,114 normalised hashes eval-samples.json 34 provider-eval/eval-set.jsonl 1,896 tests/battery/human-corpus-v1.json 40 tests/battery/human-corpus-v2.json 4,144 Total hashes in the index 6,114 NOT IN THE INDEX · NEVER HASHED AGAINST cycle2-train/dataset.jsonl The training file for the model this corpus was built to evaluate, built the same day. The index held the evaluation material and the human battery, and not the training set itself. 268 of 922 machine-written documents were later found in it 168 in the train split, 72 in test, 28 in calibration. Human half: 11 of 4,636.
Inside the index: checked against every candidate Outside the index: never checked
The corpus build pipeline stage by stage, with what each stage produced
StepStageWhat happensWhat it produced
1FetchOne candidate document is pulled from a licensed source, under the licence recorded on its row.A candidate with its licence string attached
2NormaliseNFKC normalisation, lower-casing, punctuation stripped, whitespace collapsed.One canonical text per document
3SHA-256The normalised text is hashed.One norm_sha256 per document
4Exact-collision checkThe hash is compared against the index of material the corpus must be independent of. A collision aborts the whole build, because an exact collision is proof rather than evidence.0 exact collisions against the four indexed files
5Shingle screen12-word shingle overlap at a 25% threshold. A hit drops the row and lists it in the manifest rather than halting the build, because a build that halts on a heuristic gets disabled by whoever runs it next.2 near-duplicates dropped, 505 internal duplicates dropped
6Group-aware split60/15/25 by SHA-256 of a group key, with every variant of a source landing in the same split as its source. In the cycle-3 build the group key is the post slug.Train, calibration and test splits with no source spanning two of them
Not in the indexcycle2-train/dataset.jsonl, the training file for the model this corpus was built to evaluate, built the same day. No candidate was ever hashed against it.268 of 922 AI documents were later found in it, 168 of them in the train split

longform-corpus/MANIFEST.md for the pipeline and the index contents; corpus-reconciliation-2026-08-29/analysis.txt §2 for the counts in the dashed box. Schematic, not a measurement: the only figures in it are the four index counts, their total, and the overlap counts named in §2.

06 / The index

What it was pointed at, and what it missed

The four files in the quarantine index, and the file that was not in it
Held-out sourceTextsIn the index?
eval-samples.json34yes
provider-eval/eval-set.jsonl1,896yes
tests/battery/human-corpus-v1.json40yes
tests/battery/human-corpus-v2.json4,144yes
Total hashes checked against6,114
cycle2-train/dataset.jsonlnot countedno

Against what it was pointed at, the design worked exactly as specified: 0 exact collisions, 2 near-duplicates dropped, 505 internal duplicates dropped. Source: longform-corpus/MANIFEST.md.

cycle2-train/dataset.jsonl is the training file for the model this corpus was built to evaluate, and it was built the same day. The index covered the evaluation material and the human battery, and not the training set itself.

Overlap between each half of the corpus and the cycle-2 dataset, measured after the build
Corpus fileRowsAlso present in the cycle-2 dataset
longform-corpus/ai-longform.jsonl922268 — 29.1%: 168 train, 72 test, 28 calibration
longform-corpus/human-longform.jsonl4,63611 — 0.24%: 5 train, 3 calibration, 3 test

Measured afterwards on normalised SHA-256. Source: corpus-reconciliation-2026-08-29/analysis.txt §2, recorded in CORPUS-RECONCILIATION-2026-08-29.md §2.1.

One caveat the manifest carried from the start.

PERSUADE 2.0 is not held-out material. It also appears in the cycle-2 training corpus, and the manifest says plainly that anyone combining the two corpora must deduplicate on norm_sha256. That was recorded on the day the corpus was built and it was correct. It is the most important single caveat on this page and it is not softened anywhere on this site.

07 / What it cost

What the leak was worth, in points

The effect was measured directly, by splitting the machine-written half into the documents that appear nowhere in the cycle-2 dataset and those that do. That comparison exists at one operating point and one only: the superseded 0.984 single-threshold rule, which is not the rule this tool ships.

Measured at a rule that no longer ships

Corpus
The 922-document machine-written half, split into 654 documents independent of every cycle-2 split and 268 that are not, of which 168 are in the train split.
Retired flag point
0.984
Detector
tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d
Runtime
Python onnxruntime, CPU, fp32, maximum-over-sections aggregation
Measured
29 August 2026
Also
No seen-against-unseen split has been measured at the shipped pair. These three rows are the only scoring of that split this project has.

The rule that ships today is 0.9855 / 0.9763. The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.

Figure 4 Detection on the documents the model had fitted on, against the documents it had not Detector tier3-cycle2-e5small-fp32.onnx, fp32 Python onnxruntime, maximum-over-sections aggregation, at the SUPERSEDED 0.984 single-threshold rule. This is not the shipped operating point and these rates are not the tool's current accuracy. Denominators travel with every bar. The value axis starts at 90% so a one-point gap is visible at all; the gap being one point is the finding.
Detection rate on the independent subset, the seen subset and the train split alone, at the retired 0.984 rule 90% 92% 94% 96% 98% 100% Never in the cycle-2 dataset 620 of 654 documents 94.80% In the cycle-2 dataset, any split 257 of 268 documents 95.90% In the cycle-2 train split alone 163 of 168 documents 97.02% Value axis starts at 90%, not zero
Detection by contamination subset at the retired 0.984 single-threshold rule
SubsetWhat it isnDetectedRate
Never in the cycle-2 datasetthe independent subset65462094.80%
In the cycle-2 dataset, any splittrain, test or calibration26825795.90%
In the cycle-2 train split alonethe documents the model fitted on16816397.02%

corpus-reconciliation-2026-08-29/analysis.txt §2. 654 independent, 268 seen in any cycle-2 split, 168 in the train split alone; the three subsets are drawn from the same 922 machine-written documents, so the second and third rows overlap by construction.

Both gaps should be quoted with the pair they describe. Against the seen subset the independent subset reads 1.1 points lower. Against the training split alone it reads 2.2 points lower. Weighted across the corpus, contamination is worth about 0.3 points on the headline detection figure.

Seeing the training data was worth roughly one percentage point. The correction is not that the effect is small.

It is that a figure published as a measurement on unseen documents was, for 29.1% of the machine-written half, a measurement on documents the model had fitted on. Where a page’s whole argument rests on independent data, the 654-document independent subset is the population to use, and it should be named as such rather than described as “the corpus”.

The three rows above are at 0.984 and must not be relabelled or reprinted under a shipped-pair heading. A re-measurement at the shipped pair is outstanding.

08 / Limits

What this corpus does not cover

The seen-against-unseen split has never been measured at the shipped pair.

The 1.1-point gap in Figure 4 was measured at the retired 0.984 single-threshold rule, which is the only operating point that split has ever been scored at. No seen-against-unseen figure exists for the pair that ships, and the 1.1 points must not be quoted under a shipped-pair heading. Until that re-measurement is run, the honest statement is that the corpus is not fully held out and the size of the effect at the current rule is unknown.

  • Humanities full text at scale. Europe PMC is a biomedical index first. Twenty-two subject queries widened it deliberately, into education, sociology, linguistics, history, ethics, policy and law, anthropology, business, economics and media, and the discipline label is carried per row. It remains health-adjacent humanities rather than literary criticism, theology or philosophy proper. A detector tuned on this material should not be claimed to be validated on humanities essays.
  • Modern open-licensed short fiction barely exists, and the fiction register is the weakest source in the corpus for that reason. What remains is the Internet Archive’s Creative Commons text pool, which is uneven self-publishing and scanned material; some passages are creative non-fiction or essays about literature rather than fiction, and some are OCR of scanned pages. This register’s numbers deserve more suspicion than the others, and it is also the register with this project’s worst human false-positive rate.
  • The white-paper register is two governments, the UK and the US. Think-tank, NGO and standards-body publications were sought and are largely PDF-first or non-commercially licensed.
  • Global Voices is substantially translated into English, which is a different kind of English from an originating newsroom.
  • The delivered GOV.UK era span is 2014–2022, not the 2018–2022 the filter asked for, because the search filter and the recorded year are different fields. This is stated in the manifest.
  • Pre-2022 corporate blogs under an open licence were searched for and not found at useful scale. SEC management discussion and analysis stands in for corporate communications, and it reads differently: it is written under legal review.
  • The machine-written side is the easy case. These are single-shot generations from a fixed bank of 68 topics across three prompt styles. Nothing here is edited, human-revised or adversarially rewritten, and nothing measures hybrid human-and-machine text.
  • The corpus holds no mixed documents and no machine-written document above 3,061 words. The server accepts up to 4,000 words, so the unmeasured range starts inside what it will score.
  • Register labels are machine-assigned. Every per-register figure anywhere in this project inherits that.
  • The 25,723-document figure from the signal study is a different corpus — six pooled sources, de-duplicated — and must never be conflated with these 5,558.

09 / Provenance

Where every figure on this page comes from

Source file and section for every figure on this page
FigureFileSection
Corpus 5,558 = 4,636 human + 922 AI; built 28 August 2026services/local-engine/research/longform-corpus/
Per-source human counts, licences, registers, capping rules and exclusionslongform-corpus/MANIFEST.mdwhole file
Per-model machine counts, 13 models totalling 922docs/measurements/DETECTION-BY-LENGTH-AND-MODEL.mdTable 2, denominators
Spend $12.33 of $13 authorised, per-call cost stored per rowlongform-corpus/MANIFEST.mdgeneration run
Pipeline, index contents (34 / 1,896 / 40 / 4,144 = 6,114), 0 collisions, 2 near-duplicates, 505 internal duplicateslongform-corpus/MANIFEST.mdquarantine
Cycle-4 build guard: 11,004 documents, 2,684 references, 3 rows excluded, assert above 25longform-corpus/MANIFEST.mdbuild guard
268/922 and 168 train-split overlap; human 11/4,636corpus-reconciliation-2026-08-29/analysis.txt§2
620/654, 257/268, 163/168 at the retired 0.984 rulecorpus-reconciliation-2026-08-29/analysis.txt§2
The same overlap figures as published on this sitedocs/measurements/CORPUS-RECONCILIATION-2026-08-29.md§2.1

Every file named above ships with the open measurement repository, so any figure on this page can be re-run rather than taken on trust. The counts in the masthead and in the prose are read at build time from the same thresholds.json the browser fetches, so a correction to the measurement lands here without anyone editing this page.

Read on

This is the corpus every rate on this site is measured on. Here is what those rates are.