Get in Touch

21 papers · Updated 2 Sep 2026

AI Detector Research & Testing

AI detector research done in the open: AI detector test results with their denominators, the failed methods, the withdrawn figures and the corpus limits behind the free AI content checker. AI detector accuracy tested, not claimed.

AI detector testing, in the open

AI detector accuracy: measured, not claimed

Every piece of AI detector testing below states its corpus, its runtime, the operating point it was measured at and a confidence interval on every rate. Where a figure was measured at a rule this tool no longer uses, it is labelled as such on the page and in the chart caption. Where a figure has been withdrawn, it is published beside what replaced it.

The rule that decides verdicts today is the minimum-evidence rule margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562) under segmentation contract segments-v3, running tier3-cycle5-full-e5small-fp32.onnx. Every reproducible report ships with the open measurement repository.

02 / What the signals are actually worth

What the signals are actually worth

Four papers on the things people believe separate machine writing from human writing, each measured against a corpus rather than asserted.

03 / How one verdict gets made

How one verdict gets made

Four papers on the machinery between a pasted document and a result: how it is cut up, how the pieces are combined, what length is worth, and whether the two runtimes agree.

04 / AI detector test results: the corpus, the method and the rates

AI detector test results: the corpus, the method and the rates

Where the numbers come from, what the method promises, and every published rate with its denominator and confidence interval.

05 / What travels with every figure

How accurate are AI detectors? Three limits that apply to every result

The honest answer to “how accurate are AI detectors” starts with these three limits, before any headline rate.

The long-form corpus is not fully held out.

It was published as held out and hash-quarantined against every training split, and it is not. 268 of the 922 AI documents appear in a cycle-2 split, 168 of them in the train split, and 11 of the 4,636 human documents. What that overlap was worth is measured on the corpus page rather than estimated.

Long-form only. Every detector figure across these papers describes prose of roughly 600 words and up. short marketing, SEO and social copy has never been measured on independent data, because every sample this programme owns for those registers sits inside the training set.

A retired flag point is labelled, not relabelled. Several papers publish figures measured at rules that no longer ship, because those measurements answered questions worth publishing. Each carries its own flag point in the prose, in the table and in the chart caption. None of them describes the tool as it runs today.

Apply it

Check a document and view the measured accuracy beside the result.