The research
Published so the measurement can be repeated rather than believed
Each paper below states its corpus, its runtime, the operating point it was measured at and a confidence interval on every rate. Where a figure was measured at a rule this tool no longer uses, it is labelled as such on the page and in the chart caption. Where a figure has been withdrawn, it is published beside what replaced it.
The rule that decides verdicts today is the minimum-evidence pair 0.9855 / 0.9763 under segmentation contract segments-v3, running tier3-cycle2-e5small-fp32.onnx. Every reproducible report ships with the open measurement repository.
02 / What the signals are actually worth
What the signals are actually worth
Four papers on the things people believe separate machine writing from human writing, each measured against a corpus rather than asserted.
- Rule pack en-signals:2026.08.6
116 writing rules that are editorial advice, and the six that point the wrong way
The rule tier reached mixed signals or above on 45.1% of AI writing while flagging 24.8% of human writing, so it stopped producing a verdict. Six of its rules fire more often on human prose than on machine prose, and one withdrawn figure is published beside its re-measurement.
Read the paper - Three prompt styles, 654 documents
The most popular evasion instruction, measured twice
Telling a model to write like a human moves full-length detection from 96.6% to 94.6% on 654 independent documents, an overlap rather than a difference. On the first 512 words it does more. The version of this result that circulated in 2025 described a model that no longer ships.
Read the paper - Four capabilities, none shipped
Four things we built, measured, and did not ship
A GPT-2 surprisal tier, nine reimplemented zero-shot detectors, per-sentence highlighting and two rejected retrains. Log perplexity reaches AUROC 0.715 and detects nothing at all at a 1% false-positive budget. Declining to ship a measured capability is part of the method.
Read the paper - Fourteen manipulations
What the model keys on, and what it does not notice
Lowercasing the text changes detection by nothing. Deleting every “AI vocabulary” word costs 0.8 points. Making the text repeat itself more costs 33. The classifier reads a property of the prose rather than a list of tells, and the repetition cliff is its published weakness.
Read the paper
03 / How one verdict gets made
How one verdict gets made
Four papers on the machinery between a pasted document and a result: how it is cut up, how the pieces are combined, what length is worth, and whether the two runtimes agree.
- Twelve length bands
Length dominates every other variable measured
Detection runs from 16.9% on 100-to-199-word passages to 99.2% above 2,400 words, on the same model at the same operating point. No register, no provider and no prompt style moves the result as far. Bands with too few documents to publish a rate print their raw count instead.
Read the paper - Seven candidate rules
A long document is many readings: why the verdict is the strongest section
On 700 synthetic half-AI documents the strongest section catches 638; the document average catches 29. Averaging a document's sections buries the AI half under the human half, which is why the shipped rule is a minimum-evidence pair rather than a mean.
Read the paper - Two runtimes, one answer
The same text, scored in your browser and on our server
In the score band where the flag points sit, the browser's two execution providers differ by a median of 0.000017. Lower in the range they differ by far more, and int8 against fp32 differs more again. Each route publishes the accuracy measured on the runtime that produced it.
Read the paper - 684 documents, silently truncated
The text the model never saw: a bug that hid inside a passing test
The segmentation rule counted words and the model counts tokens, so 12.31% of documents had their ends dropped without a warning and 2.98% of all tokens were never scored. The test suite passed throughout. Coverage is now 0 of 21,093 sections over the window.
Read the paper
04 / The corpus, the method and the rates
The corpus, the method and the rates
Where the numbers come from, what the method promises, and every published rate with its denominator and confidence interval.
- 5,558 documents, eight sources
5,558 documents, a quarantine, and the file it was never pointed at
A hash quarantine, a shingle screen and a group-aware split, pointed at four index files and not at a fifth. 268 of the 922 AI documents sit inside a cycle-2 split. What that overlap was worth is measured rather than estimated: about one point.
Read the paper - Contract 1.0.0
AI Content Integrity methodology
How exact text checks, editorial patterns, protected facts, detector states, privacy routes and reproducible receipts are kept separate, with every published rate at the operating point that ships and two corrections stated rather than quietly edited.
Read the paper - Every measured cell
Detection rates in full
Every measured cell by document length, by the model that wrote the text and by content type, each with its denominator, its corpus, its runtime, its operating point and a 95% confidence interval. The shipped route detects 883/922 (95.8%) of AI long-form writing.
Read the paper
05 / What travels with every figure
Three conditions that qualify the whole set
It was published as held out and hash-quarantined against every training split, and it is not. 268 of the 922 AI documents appear in a cycle-2 split, 168 of them in the train split, and 11 of the 4,636 human documents. What that overlap was worth is measured on the corpus page rather than estimated.
Long-form only. Every detector figure across these papers describes prose of roughly 600 words and up. short marketing, SEO and social copy has never been measured on independent data, because every sample this programme owns for those registers sits inside the training set.
A retired flag point is labelled, not relabelled. Several papers publish figures measured at rules that no longer ship, because those measurements answered questions worth publishing. Each carries its own flag point in the prose, in the table and in the chart caption. None of them describes the tool as it runs today.