On this page
01 / The finding
The strongest attack works, the popular one does not, and both numbers are published
A detector that publishes its detection rates owes its readers the escape rates too. Three findings sit on this page, each from a measurement against tier3-cycle2-v1 at the operating point that shipped on 30–31 August 2026 (retired 1 September 2026), and they point in different directions.
Of the 28 source texts this build first caught, 27 transformed outputs of at least 200 words fell below the verdict rule. StealthGPT: 24 of 25. These are direct results against the build live at the time, not third-party figures, and this challenge has not been re-run against tier3-cycle5-v1.
526 of 550. Asking an ordinary LLM to reword machine text did not hide it — the flag rate on AI documents rose with rewrite strength, from 65.3% untouched to 79.2% after a full rewrite.
57 of 272 full rewrites. A person's document, rewritten by a machine, was called machine-like about one time in five. tier3-cycle5-v1 reads the same axis at 28.5% (39/137), up from cycle-2's 21.2% on its own held-out rows — a direct, disclosed consequence of training heavy rewrites as machine-written, which they are. See §5.
A purpose-built humaniser is a different attack from an LLM rewrite, and the two must never share a number.
02 / The commercial challenge
Twelve commercial humanisers against tier3-cycle2, the build live at the time
Thirty-three source texts — eleven each from Claude Sonnet 5, Gemini 3.5 Flash and GPT-5.5 — were first scored untouched. The build caught 28 of 33: all eleven Claude texts, all eleven Gemini texts, and six of the eleven GPT-5.5 texts. Each source then went through twelve humaniser services, producing 396 transformed outputs, and each output was scored by the same build under the same rule.
Measured at a rule that no longer ships
- Corpus
- 33 source texts, 11 each from three current models; 396 transformed outputs from 12 commercial humanisers. Eligible outputs: at least 200 words, from a source the build first caught.
- Retired flag point
- 0.9855 / 0.9763
- Detector
- tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d (cycle 2, superseded 1 September 2026)
- Runtime
- The server build live at the time of the run — tier3-cycle2, e313ab00de1fffd2, segments-v3, superseded 1 September 2026 by tier3-cycle5-v1 — live health rechecked before the run, 30 August 2026.
- Measured
- 30 August 2026
- Also
- Escape is conditional: the share of eligible outputs that fell below the verdict rule. It is not an accuracy claim, a quality claim, or a statement that the outputs read as human. Not re-challenged against tier3-cycle5-v1.
The rule that ships today is margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562). The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.
| Humaniser | Escaped / eligible | Rate |
|---|---|---|
| Undetectable.ai | 27 of 28 | 96.4% |
| StealthGPT | 24 of 25 | 96.0% |
| WriteHuman | 24 of 27 | 88.9% |
| Humanize AI Pro | 23 of 26 | 88.5% |
| AI Humanize io | 22 of 25 | 88.0% |
| Phrasly | 23 of 28 | 82.1% |
| HIX Bypass | 20 of 27 | 74.1% |
| Stealth Writer | 20 of 27 | 74.1% |
| Humbot | 18 of 27 | 66.7% |
| Walter Writes | 17 of 28 | 60.7% |
| Super Humanizer | 15 of 27 | 55.6% |
| Grammarly | 2 of 27 | 7.4% |
research/HUMANIZER-AND-DETECTOR-COMPETITOR-STUDY.md, 'Exact current Opace challenge result', run 30 August 2026 against tier3-cycle2. Full feature deltas and reproduction instructions are in the study's evidence file.
Read the top row correctly. 96.4% means: of the 28 source articles this specific build initially caught, 27 transformed outputs of at least 200 words fell below the verdict rule. It does not mean 96.4% accuracy, 96.4% quality, or a 96.4% chance of fooling every detector. The corpus is small — 33 sources from three generating models — and the run is a snapshot of products that change without notice.
The bottom row is the control that makes the table readable: Grammarly, which rewrites for clarity rather than for evasion, freed 2 of 27. The spread from 7.4% to 96.4% is the finding — these services are not interchangeable, and the strongest of them beat this detector almost every time.
03 / Two direct checks
QuillBot failed its one test; JustDone remains unmeasured, and the reason is published
QuillBot's free humaniser failed the one direct test it was given. A 450-word Claude article scoring 0.9897 came back as three recombined free-tier chunks totalling 331 words — still scoring 0.9897, still flagged, with 39.6% of the original four-word sequences intact. One sample is not a rate, and the free quota ended before a controlled run could follow. It is published as one direct observation and nothing more.
JustDone was not measured, on purpose. Its Terms of Use prohibit the automated access the evaluation required, in terms that working slowly or discarding the output would not cure. The evaluation stopped before a single document was submitted; the harness proof, the selected sample and the reference profile are preserved so that one lawful script run completes it. Until then, JustDone's paid mode is recorded as untested against this build — a prediction either way would be a guess wearing a percentage.
04 / The popular attack
Asking an LLM to reword AI text makes it easier to catch, not harder
The attack most people can actually mount is not a specialist service. It is pasting the text into a chatbot and asking for a rewrite. That attack was measured on a purpose-built paired corpus: 600 source documents — 300 AI-written, 300 human-written — each rewritten at three strengths (copy-edit, structural paraphrase, full rewrite) by five different models, giving 2,302 rows in total, every rewrite carrying the lineage of its source.
Measured at a rule that no longer ships
- Corpus
- 2,302 rows: 600 sources (300 AI, 300 human) and 1,702 kept rewrites at three strengths across five rewriting models, one family held out. Median source length 372 words — far shorter than the corpus behind the headline rates, which is why no figure here may be read against them.
- Retired flag point
- 0.9855 / 0.9763
- Detector
- tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d (cycle 2, superseded 1 September 2026)
- Runtime
- Server route, tier3-cycle2-e5small-fp32.onnx, segments-v3, cycle-2's live pair 0.9855/0.9763 — retired 1 September 2026. Harness re-proved against 883/922 and 45/4,636 before any new figure was taken. The matched browser subset agrees within about a point on every cell.
- Measured
- 30 August 2026
- Also
- Every row is generic LLM paraphrase — commercial_humaniser: false. Nothing in this section describes what the services in Figure 1 do. Not re-measured against tier3-cycle5-v1, which reads the heavy-rewrite-of-human-original axis higher on its own held-out corpus — see §5.
The rule that ships today is margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562). The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.
| Strength | AI originals flagged | Human originals flagged |
|---|---|---|
| Untouched | 65.3% (196/300) | 1.0% (3/300) |
| Copy-edit | 67.4% (174/258) | 1.4% (4/290) |
| Structural paraphrase | 74.8% (220/294) | 11.0% (33/299) |
| Full rewrite | 79.2% (229/289) | 21.0% (57/272) |
LLM-REWRITE-ROBUSTNESS.md §3.1, §3.3; per-strength AI figures from the fp32 results record. Untouched human false positives on this corpus: 1.0% (3/300).
The paired reading is the strong one. Of the AI sources the build detects, 95.6% of their rewrites are still detected — 526 of 550 — and even at full-rewrite strength, where 9% of the original four-word sequences survive, 92.0% (172 of 187) stay caught. No rewriting model defeated it: the five ranged from 91.9% to 99.1% survival, with the held-out family mid-pack. Humaniser products imply the opposite, and for this class of attack the implication is measured and wrong.
The 65.3% baseline needs its own sentence: these sources are short, at a median of 372 words and one scored section, and length dominates every other variable. The baseline is a property of the corpus, not a revision of any published rate.
05 / The other side of the ledger
One in five heavily rewritten human documents gets flagged — and the rewriter decides
The same corpus answers the question that matters to people who never tried to cheat: what happens when a machine rewrites a person's work? The ladder in Figure 2 behaves correctly at its base — a copy-edit of your own prose is left alone at 1.4%, near the untouched rate, which is deliberate. At the top, where almost nothing of the author's wording survives, 21.0% of documents are flagged. Polish your writing through a chatbot aggressively enough and there is about a one-in-five chance of being called machine-written — which is fair comment on the words, since a machine wrote them, and still a real cost to the person who started with an honest draft.
| Rewriting model | Flagged / rewrites | Rate |
|---|---|---|
google/gemini-3.7-flash | 39 of 196 | 19.9% |
mistralai/mistral-medium-3-5 (unseen family) | 24 of 134 | 17.9% |
deepseek/deepseek-v4-pro-0813 | 13 of 173 | 7.5% |
openai/gpt-5.6-luna | 12 of 174 | 6.9% |
meta-llama/llama-4-maverick | 6 of 184 | 3.3% |
LLM-REWRITE-ROBUSTNESS.md §3.4. The mistral family was held out of every split as an unseen family.
The registers where this cost lands hardest are the two already published as the tool's weakest: academic writing at 18.2% (32 of 176) and fiction at 16.9% (29 of 172) on the rewritten-human arm. That is the known weakness reappearing on a new axis rather than a new one.
On cycle 5's own 1,199-row held-out humaniser-pairs corpus (a different corpus from the 272-row figure above, so the two numbers are not directly subtracted), a heavy LLM rewrite of a human original now flags 28.5% (39/137), up from the superseded cycle-2 model's 21.2% on the same rows — a direct consequence of training heavy rewrites as machine-written, which their words are. Light copy-edits stay clear: 0.7% (1/153). Source: CYCLE5-REPORT.md §Recommendation point 3 and §5.
06 / Limits
What this page does not prove
33 sources from three generating models, one snapshot run, products that change weekly, and two absent names: JustDone (terms boundary, above) and QuillBot beyond a single direct test. The 96.4% row is conditional on sources the build first caught and on outputs of at least 200 words.
The paired corpus is LLM paraphrase throughout. It measures the attack most users can mount, not what a purpose-built humaniser does — the services in Figure 1 beat this build 96% of the time while ordinary rewriting fails. A training response to the humaniser gap needs humaniser output, which is a licensing question as much as a technical one.
Short documents, short sections. The paired corpus has a median of 372 words and one scored section per document, so its absolute rates sit far below the long-form headline figures and may not be compared with them. The paired, same-lineage comparison — did the rewrite of a caught document escape? — is the only reading that survives the length confound.
Escape is not authorship. An output that falls below the verdict rule has evaded one detector at one operating point. Nothing on this page says such text reads well, passes any other tool, or is safe to publish as human work.