Get in Touch

Verdict study · 2,302 documents

The verdict we refuse to give

Four-way human and edited-AI labels failed on 2,302 documents, so the checker does not claim to provide them.

On this page
  1. The finding
  2. How it was measured
  3. The impossible half
  4. The real but unshippable half
  5. What the label would cost
  6. What this does not prove
Published
31 August 2026
Measured
Scoring 31 August 2026; full analysis completed the same day and confirmed against the interim figures
Corpus
2,302 rows from 600 source lineages: 300 human and 300 AI originals, each LLM-rewritten at three strengths by five models
Operating point
tier3-cycle2-v1's live configuration throughout: fp32, segments-v3, pair 0.9855/0.9763 — the pair that shipped when this corpus was scored on 31 August 2026; superseded 1 September 2026 by tier3-cycle5-v1's margin-space rule, and this 2,302-row separability study has not been re-run against it

01 / The finding

Half the four-way verdict is impossible, and the other half is unaffordable

AI-then-rewritten against pure AI0.448

Below the 0.5 of a coin (CI 0.431–0.466), and it worsens with rewrite strength: 0.487 light, 0.438 medium, 0.424 heavy. Rewritten AI reads MORE machine-like, not less. "Likely AI but human edited" has nothing to stand on.

Human against human-then-AI-edited, heavy edits0.866

A real signal — the detector noticing that a machine produced the words. Overall 0.751; the paired shift on heavy edits is +3.039 in margin, upward in 96.7% of 272 pairs. Real, and still unshippable below.

What the second label catches at an honest budget21.3%

At a 1% false-label budget — the discipline every verdict here keeps — the best available detector catches 21.3% of AI-edited human documents overall, and 2.6% of light copy-edits, which are what people actually do.

If "has this been through a machine" were a signal distinct from "did a machine write this", rewritten AI would separate from pure AI. It measures 0.448.

The evidence cited for the four-way idea is the evidence that disproves it. This is the programme's fifth measured decline, published like the other four.

02 / Method

Paired documents, the shipped score, and a correction stated rather than buried

Measured at a rule that no longer ships

Corpus
600 sources (300 human, 300 AI), 1,702 kept rewrites at three strengths across five rewriting models, 2,302 rows; every rewrite carries its source lineage, and intervals are cluster-bootstrapped over the 600 lineages. Median document 343 words.
Retired flag point
0.9855 / 0.9763
Detector
tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d (cycle 2, superseded 1 September 2026)
Runtime
tier3-cycle2-e5small-fp32.onnx, segments-v3, temperature 0.8324, shipped pair 0.9855/0.9763. Harness re-proved against the published 883/922 and 45/4,636 before scoring.
Measured
30 August 2026
Also
Deterministic run, seeded. Every row is generic LLM paraphrase — commercial_humaniser: false on all of them.

The rule that ships today is margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562). The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.

A correction, in the open. The first analysis pass crashed partway and one of its probe figures carried a feature leak that produced a giveaway AUROC of 1.000. Both defects were fixed before the analysis was committed, and the complete run of 31 August 2026 confirms every interim figure unchanged — except the leaked probe, which falls to its honest value of about 0.60. A number that flatters the conclusion you want is the one to check hardest; this one was.

Figure 1 Every pairwise contrast on the shipped document score AUROC with 95% cluster-bootstrap intervals. Five of the six contrasts separate respectably — they are all versions of the human/machine axis the model learned. The sixth is the one the four-way verdict needs, and it sits below chance.
Pairwise AUROC for every contrast the four-way verdict would need, drawn either side of chance 0.00 0.25 0.50 0.75 1.00 Human v AI+rewrite n = 300/841 · CI 0.969–0.986 0.978 Human v pure AI n = 300/300 · CI 0.946–0.976 0.962 Human+AI-edit v AI+rewrite n = 861/841 · CI 0.875–0.918 0.897 Human+AI-edit v pure AI n = 861/300 · CI 0.831–0.885 0.858 Human v human+AI-edit n = 300/861 · CI 0.731–0.771 0.751 AI+rewrite v pure AI n = 841/300 · CI 0.431–0.466 0.448 chance, 0.500
Pairwise AUROC on the shipped score
ContrastnAUROC95% CI
Human v AI+rewrite300/8410.9780.969–0.986
Human v pure AI300/3000.9620.946–0.976
Human+AI-edit v AI+rewrite861/8410.8970.875–0.918
Human+AI-edit v pure AI861/3000.8580.831–0.885
Human v human+AI-edit300/8610.7510.731–0.771
AI+rewrite v pure AI841/3000.4480.431–0.466

fourway-separability-2026-08-31/RESULTS.md §1. AI+rewrite v pure AI by strength: 0.487 light (258/300), 0.438 medium (294/300), 0.424 heavy (289/300).

03 / The impossible half

"Likely AI but human edited" — the classes collapse into each other

The label needs AI-then-rewritten text to look different from pure AI text. It does not, and every attempt to force the distinction failed on its own numbers. An oracle three-class boundary — fitted on the very data it is scored on, the most generous test that exists — reaches 62.4% balanced accuracy but recovers only 20 of 300 pure-AI documents: the class simply collapses into AI+rewrite. A group-aware probe over the shipped outputs plus honest surface features manages 57.5–57.8% balanced accuracy on held-out data (n = 598), an AUROC of about 0.60. Nothing there supports telling a user "AI, then edited" apart from "AI".

There is also no operating point for the mirror-image accusation: pure AI wrongly called "AI, rewritten" at a 1% budget catches 2.3% of what it looks for. The ordering is inverted — a rewrite makes machine text read more machine-like — and an inverted ordering cannot be thresholded into honesty.

04 / The real half

"Likely human but AI edited" — a real signal that fails every honest budget

The other half is genuinely there. A human original put through a heavy LLM rewrite separates from untouched human writing at 0.866, and the paired movement is large and consistent: the median document moves +3.039 in margin, upward in 96.7% of 272 heavy pairs, against a flag point at 3.512. The detector is noticing exactly what it was built to notice — that a machine produced the words on the page.

Three things stop that becoming a verdict. The strong version of the measurement compares a document against its own original, which inference does not have. The signal lives almost entirely in heavy rewrites, while the edit most people actually make is a light polish — flagged at 1.4% (4 of 290), deliberately, since a person tidying their own prose is not a case to accuse. And the label is an accusation of the same kind as "AI-generated", so it must be held to the same 1% false-label discipline — where the arithmetic in the next section takes over.

05 / The price

Hold the false-label rate at 1%, and read off what the label catches

Detection of AI-edited human writing at each false-label budget
False-label budgetAll intensitiesLight copy-editStructural paraphraseFull rewrite
1% — the tool's discipline21.3% [17.8–25.4]2.6%20.6%43.1% [35.1–51.4]
2%39.3%7.2%47.1%66.4%
5%46.7%14.4%54.8%73.7%
20%60.4%35.9%63.9%83.9%

In plain terms: ship the label at the 1% discipline and it is wrong about a genuinely human writer one time in a hundred while missing 79% of the documents it exists to catch — and 97% of light edits. Loosen to 5% to catch about half, and one human writer in twenty is told their own work was machine-edited. A light copy-edit is invisible at every budget, and a light copy-edit is what most people actually do. The probe figures above are the best available detector; the shipped score alone reads 11.7% at the 1% budget.

What ships instead is the honest subset of the finding: the measured fact that 21.0% of heavily rewritten human originals get flagged is published as a weakness of the plain verdict, with its denominators, rather than dressed up as a fourth verdict the measurement cannot support.

06 / Limits

What this page does not prove

The corpus is LLM paraphrase, not commercial-humaniser output. The largest limitation, and it applies to every number here. Purpose-built humanisers escaped this build 96.4% and 96.0% of the time; nothing on this page describes them.

"AI then genuinely human-edited" does not exist in this corpus. It was measured against AI-then-LLM-rewritten as a proxy. The proxy is the generous case and it still fails, which strengthens the negative — but a corpus of real human edits could in principle behave differently, and nothing here rules that out.

Bounded, not universal. The probe is a linear model over a dozen features; a larger model trained end-to-end might beat 0.60 on the AI side. 9.0% of the AI sources touch the shipped model's training set, which inflates pure-AI scores and thereby widens the gap the failed contrast needed — so contamination cannot be the reason it failed. Median document 343 words; absolute rates are not comparable to the long-form headline figures. And the corpus contains no mixed half-and-half documents, which remain the aggregation rule's separately recorded weakness.

Apply the method

Check a document and review each result separately.