AI detector testing, in the open
AI detector accuracy: measured, not claimed
Every piece of AI detector testing below states its corpus, its runtime, the operating point it was measured at and a confidence interval on every rate. Where a figure was measured at a rule this tool no longer uses, it is labelled as such on the page and in the chart caption. Where a figure has been withdrawn, it is published beside what replaced it.
The rule that decides verdicts today is the minimum-evidence rule margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562) under segmentation contract segments-v3, running tier3-cycle5-full-e5small-fp32.onnx. Every reproducible report ships with the open measurement repository.
02 / What the signals are actually worth
What the signals are actually worth
Four papers on the things people believe separate machine writing from human writing, each measured against a corpus rather than asserted.
- Rule pack en-signals:2026.08.6
116 writing rules that are editorial advice, and the six that point the wrong way
The rule tier reached mixed signals or above on 45.1% of AI writing while flagging 24.8% of human writing, so it stopped producing a verdict. Six of its rules fire more often on human prose than on machine prose, and one withdrawn figure is published beside its re-measurement.
Read the paper - Three prompt styles, 654 documents
The most popular evasion instruction, measured twice
Telling a model to write like a human moves full-length detection from 96.6% to 94.6% on 654 independent documents, an overlap rather than a difference. On the first 512 words it does more. The version of this result that circulated in 2025 described a model that no longer ships.
Read the paper - 98 folk phrases, two corpora
The tells we tested, and the folk lore that failed
Twenty-one of ninety-eight famous 'AI phrases' survive a dual-corpus 2× gate; thirteen point at humans, and 'delve into' reads 0.2× on 2026 models. The strongest shape tell is one nobody quotes: paragraphs of near-identical length, 13.2% of AI documents (506 of 3,839) against 0.8% of structured human ones.
Read the paper - 18 rows, every one an interval
Eighteen phrases we can honestly put a number on
The first phrase table ranked 'per cent of' at 30.1× — a British-spelling detector in disguise, caught before shipping and published. What survives spelling normalisation, a register control and held-out validation: eighteen three-word phrases with document counts on every row, and one artefact cut by judgement with its numbers shown.
Read the paper - One ear, 5,558 documents
The rhythm you can hear, measured — and refused a vote
Opace's founder reads a paragraph rhythm in humanised AI by ear. Measured, his ear is real: the signals fire on the model's misses at 3.56× the rate they fire on cleared humans (20 of 45 against 571 of 4,580). Letting it change the score would buy five catches for five false accusations, so it explains results instead.
Read the paper - 418 never-trained documents
The 27.3% problem: the weakness our own evaluation found
Handed modern structured human writing with its markdown intact, the superseded tier3-cycle2 model wrongly flagged 27.3% of it (114 of 418). The mechanism was the syntax itself — the same documents read 0.0% stripped — and the trained replacement, tier3-cycle5-v1, deployed 1 September 2026 and reads the same 418 documents at 0.2%.
Read the paper - Twelve humanisers, one build
The humaniser weakness, published in full
Against tier3-cycle2 (retired 1 September 2026), Undetectable.ai escaped on 27 of the 28 eligible texts this build first caught, and StealthGPT on 24 of 25. Ordinary LLM rewriting failed as an attack — 95.6% of detected AI documents stayed detected (526 of 550) — while 21.0% of heavily rewritten human originals (57 of 272) were flagged; tier3-cycle5-v1 reads that same axis at 28.5% on its own held-out rows, the disclosed trade of adopting it.
Read the paper - Four capabilities, none shipped
Four things we built, measured, and did not ship
A GPT-2 surprisal tier, nine reimplemented zero-shot detectors, per-sentence highlighting and two rejected retrains. Log perplexity reaches AUROC 0.715 and detects nothing at all at a 1% false-positive budget. Declining to ship a measured capability is part of the method.
Read the paper - Fourteen manipulations
What the model keys on, and what it does not notice
Lowercasing the text changes detection by nothing. Deleting every “AI vocabulary” word costs 0.8 points. Making the text repeat itself more costs 33. The classifier reads a property of the prose rather than a list of tells, and the repetition cliff is its published weakness.
Read the paper
03 / How one verdict gets made
How one verdict gets made
Four papers on the machinery between a pasted document and a result: how it is cut up, how the pieces are combined, what length is worth, and whether the two runtimes agree.
- Twelve length bands
Length dominates every other variable measured
Detection runs from 16.9% on 100-to-199-word passages to 99.2% above 2,400 words, on the same model at the same operating point. No register, no provider and no prompt style moves the result as far. Bands with too few documents to publish a rate print their raw count instead.
Read the paper - Seven candidate rules
A long document is many readings: why the verdict is the strongest section
On 700 synthetic half-AI documents the strongest section catches 638; the document average catches 29. Averaging a document's sections buries the AI half under the human half, which is why the shipped rule is a minimum-evidence pair rather than a mean.
Read the paper - 2,302 paired documents
The verdict we refuse to give
'Likely AI but human edited' needs rewritten AI to separate from pure AI. It measures AUROC 0.448 — below chance, and worse the heavier the rewrite. The mirror label is real (0.866 on heavy edits) and still catches only 21.3% at an honest 1% false-label budget. The fifth measured decline.
Read the paper - 269,732 sentences
Why no sentence gets a number
A sentence score reaches AUROC 0.764 where the document score reaches 0.9695, 57.4% of the sentences inside AI documents read human, and one spelling change moved a sentence by 0.475. What ships is a ranking behind a counted floor — 25 false marks in 200,890 human sentences — and never a percentage.
Read the paper - 18 operating points
The price of strictness, at every notch
Competitors ship sensitivity dials with no prices. The whole corpus re-scored at eighteen operating points against tier3-cycle2 (retired 1 September 2026): one notch looser than the default bought 0.9 points of detection for double the wrongly-flagged humans (45 → 91 of 4,636), and below 0.96 the dial was close to pure cost.
Read the paper - Two runtimes, one answer
The same text, scored in your browser and on our server
In the score band where the flag points sit, the browser's two execution providers differ by a median of 0.000017. Lower in the range they differ by far more, and int8 against fp32 differs more again. Each route publishes the accuracy measured on the runtime that produced it.
Read the paper - 684 documents, silently truncated
The text the model never saw: a bug that hid inside a passing test
The segmentation rule counted words and the model counts tokens, so 12.31% of documents had their ends dropped without a warning and 2.98% of all tokens were never scored. The test suite passed throughout. Coverage is now 0 of 21,093 sections over the window.
Read the paper
04 / AI detector test results: the corpus, the method and the rates
AI detector test results: the corpus, the method and the rates
Where the numbers come from, what the method promises, and every published rate with its denominator and confidence interval.
- 5,558 documents, eight sources
5,558 documents, a quarantine, and the file it was never pointed at
A hash quarantine, a shingle screen and a group-aware split, pointed at four index files and not at a fifth. 268 of the 922 AI documents sit inside a cycle-2 split. What that overlap was worth is measured rather than estimated: about one point.
Read the paper - Contract 1.0.0
AI detector methodology: how every rate is measured
How exact text checks, editorial patterns, protected facts, detector states, privacy routes and reproducible receipts are kept separate, with every published rate at the operating point that ships and two corrections stated rather than quietly edited.
Read the paper - Every measured cell
Detection rates in full
Every measured cell by document length, by the model that wrote the text and by content type, each with its denominator, its corpus, its runtime, its operating point and a 95% confidence interval. The shipped route detects 902/922 (97.8%) of AI long-form writing.
Read the paper - Watermarking, not detection
Claude, SynthID-Text and what a watermark can actually prove
The one boundary this tool cannot cross by measuring harder. Anthropic's commitment covers models launched from 2 August 2026, no shipping model falls under it yet, the production keys are private and no public verifier exists. What the published SynthID-Text mathematics does prove, on the three public demo keys, is run in your browser on every assessment.
Read the paper
05 / What travels with every figure
How accurate are AI detectors? Three limits that apply to every result
The honest answer to “how accurate are AI detectors” starts with these three limits, before any headline rate.
It was published as held out and hash-quarantined against every training split, and it is not. 268 of the 922 AI documents appear in a cycle-2 split, 168 of them in the train split, and 11 of the 4,636 human documents. What that overlap was worth is measured on the corpus page rather than estimated.
Long-form only. Every detector figure across these papers describes prose of roughly 600 words and up. short marketing, SEO and social copy has never been measured on independent data, because every sample this programme owns for those registers sits inside the training set.
A retired flag point is labelled, not relabelled. Several papers publish figures measured at rules that no longer ship, because those measurements answered questions worth publishing. Each carries its own flag point in the prose, in the table and in the chart caption. None of them describes the tool as it runs today.