Get in Touch
AI content toolsfrom Opace

12 measurement papers · re-verified 30 August 2026

AI content detection, measured

Every claim this tool makes is a measurement with a denominator behind it. These papers publish those measurements in full, including the ones that did not work, the figures that were withdrawn and the corpus problem we found in our own quarantine.

The research

Published so the measurement can be repeated rather than believed

Each paper below states its corpus, its runtime, the operating point it was measured at and a confidence interval on every rate. Where a figure was measured at a rule this tool no longer uses, it is labelled as such on the page and in the chart caption. Where a figure has been withdrawn, it is published beside what replaced it.

The rule that decides verdicts today is the minimum-evidence pair 0.9855 / 0.9763 under segmentation contract segments-v3, running tier3-cycle2-e5small-fp32.onnx. Every reproducible report ships with the open measurement repository.

02 / What the signals are actually worth

What the signals are actually worth

Four papers on the things people believe separate machine writing from human writing, each measured against a corpus rather than asserted.

03 / How one verdict gets made

How one verdict gets made

Four papers on the machinery between a pasted document and a result: how it is cut up, how the pieces are combined, what length is worth, and whether the two runtimes agree.

04 / The corpus, the method and the rates

The corpus, the method and the rates

Where the numbers come from, what the method promises, and every published rate with its denominator and confidence interval.

05 / What travels with every figure

Three conditions that qualify the whole set

The long-form corpus is not fully held out.

It was published as held out and hash-quarantined against every training split, and it is not. 268 of the 922 AI documents appear in a cycle-2 split, 168 of them in the train split, and 11 of the 4,636 human documents. What that overlap was worth is measured on the corpus page rather than estimated.

Long-form only. Every detector figure across these papers describes prose of roughly 600 words and up. short marketing, SEO and social copy has never been measured on independent data, because every sample this programme owns for those registers sits inside the training set.

A retired flag point is labelled, not relabelled. Several papers publish figures measured at rules that no longer ship, because those measurements answered questions worth publishing. Each carries its own flag point in the prose, in the table and in the chart caption. None of them describes the tool as it runs today.

Apply it

Run a document, and keep the limits beside the result.