Get in Touch

Humaniser study · 28 texts

The humaniser weakness

Undetectable.ai escaped the retired detector on 27 of 28 eligible texts. The current model has not yet repeated the test.

On this page
  1. The finding
  2. Twelve commercial humanisers
  3. QuillBot and JustDone
  4. Ordinary LLM rewriting fails
  5. The cost to real writers
  6. What this does not prove
Published
31 August 2026
Measured
Commercial challenge 30 August 2026 · LLM-rewrite corpus 31 August 2026, both against tier3-cycle2-v1, the build live at the time — superseded 1 September 2026 by tier3-cycle5-v1, against which neither study has been re-run
Operating point
tier3-cycle2-v1's live minimum-evidence pair, 0.9855 primary / 0.9763 secondary, segments-v3 — retired 1 September 2026
Corpora
33 source texts × 12 humanisers (396 outputs) · 2,302-row paired rewrite corpus from 600 sources

01 / The finding

The strongest attack works, the popular one does not, and both numbers are published

A detector that publishes its detection rates owes its readers the escape rates too. Three findings sit on this page, each from a measurement against tier3-cycle2-v1 at the operating point that shipped on 30–31 August 2026 (retired 1 September 2026), and they point in different directions.

Undetectable.ai, conditional long-form escape (tier3-cycle2, 30 Aug 2026)27/28

Of the 28 source texts this build first caught, 27 transformed outputs of at least 200 words fell below the verdict rule. StealthGPT: 24 of 25. These are direct results against the build live at the time, not third-party figures, and this challenge has not been re-run against tier3-cycle5-v1.

Detected AI documents still detected after an LLM rewrite (tier3-cycle2)95.6%

526 of 550. Asking an ordinary LLM to reword machine text did not hide it — the flag rate on AI documents rose with rewrite strength, from 65.3% untouched to 79.2% after a full rewrite.

Heavily rewritten human originals flagged (tier3-cycle2)21.0%

57 of 272 full rewrites. A person's document, rewritten by a machine, was called machine-like about one time in five. tier3-cycle5-v1 reads the same axis at 28.5% (39/137), up from cycle-2's 21.2% on its own held-out rows — a direct, disclosed consequence of training heavy rewrites as machine-written, which they are. See §5.

A purpose-built humaniser is a different attack from an LLM rewrite, and the two must never share a number.

Every figure on this page names which of the two it describes. The paired-corpus rows carry commercial_humaniser: false on every row; the commercial challenge used the live products.

02 / The commercial challenge

Twelve commercial humanisers against tier3-cycle2, the build live at the time

Thirty-three source texts — eleven each from Claude Sonnet 5, Gemini 3.5 Flash and GPT-5.5 — were first scored untouched. The build caught 28 of 33: all eleven Claude texts, all eleven Gemini texts, and six of the eleven GPT-5.5 texts. Each source then went through twelve humaniser services, producing 396 transformed outputs, and each output was scored by the same build under the same rule.

Measured at a rule that no longer ships

Corpus
33 source texts, 11 each from three current models; 396 transformed outputs from 12 commercial humanisers. Eligible outputs: at least 200 words, from a source the build first caught.
Retired flag point
0.9855 / 0.9763
Detector
tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d (cycle 2, superseded 1 September 2026)
Runtime
The server build live at the time of the run — tier3-cycle2, e313ab00de1fffd2, segments-v3, superseded 1 September 2026 by tier3-cycle5-v1 — live health rechecked before the run, 30 August 2026.
Measured
30 August 2026
Also
Escape is conditional: the share of eligible outputs that fell below the verdict rule. It is not an accuracy claim, a quality claim, or a statement that the outputs read as human. Not re-challenged against tier3-cycle5-v1.

The rule that ships today is margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562). The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.

Figure 1 Conditional long-form escape, by humaniser Share of eligible transformed outputs — at least 200 words, source first caught — that fell below the verdict rule tier3-cycle2 shipped at the time (retired 1 September 2026). Each row prints its own numerator and denominator; denominators differ because short outputs are excluded per row. Rows at 90% or above are drawn emphasised.
Conditional long-form escape rate for twelve commercial humanisers against the tier3-cycle2 Opace build, 30 August 2026 0% 25% 50% 75% 100% Undetectable.ai 27/28 eligible outputs 96.4% StealthGPT 24/25 eligible outputs 96.0% WriteHuman 24/27 eligible outputs 88.9% Humanize AI Pro 23/26 eligible outputs 88.5% AI Humanize io 22/25 eligible outputs 88.0% Phrasly 23/28 eligible outputs 82.1% HIX Bypass 20/27 eligible outputs 74.1% Stealth Writer 20/27 eligible outputs 74.1% Humbot 18/27 eligible outputs 66.7% Walter Writes 17/28 eligible outputs 60.7% Super Humanizer 15/27 eligible outputs 55.6% Grammarly 2/27 eligible outputs 7.4%
Escape counts and rates by commercial humaniser
HumaniserEscaped / eligibleRate
Undetectable.ai27 of 2896.4%
StealthGPT24 of 2596.0%
WriteHuman24 of 2788.9%
Humanize AI Pro23 of 2688.5%
AI Humanize io22 of 2588.0%
Phrasly23 of 2882.1%
HIX Bypass20 of 2774.1%
Stealth Writer20 of 2774.1%
Humbot18 of 2766.7%
Walter Writes17 of 2860.7%
Super Humanizer15 of 2755.6%
Grammarly2 of 277.4%

research/HUMANIZER-AND-DETECTOR-COMPETITOR-STUDY.md, 'Exact current Opace challenge result', run 30 August 2026 against tier3-cycle2. Full feature deltas and reproduction instructions are in the study's evidence file.

Read the top row correctly. 96.4% means: of the 28 source articles this specific build initially caught, 27 transformed outputs of at least 200 words fell below the verdict rule. It does not mean 96.4% accuracy, 96.4% quality, or a 96.4% chance of fooling every detector. The corpus is small — 33 sources from three generating models — and the run is a snapshot of products that change without notice.

The bottom row is the control that makes the table readable: Grammarly, which rewrites for clarity rather than for evasion, freed 2 of 27. The spread from 7.4% to 96.4% is the finding — these services are not interchangeable, and the strongest of them beat this detector almost every time.

03 / Two direct checks

QuillBot failed its one test; JustDone remains unmeasured, and the reason is published

QuillBot's free humaniser failed the one direct test it was given. A 450-word Claude article scoring 0.9897 came back as three recombined free-tier chunks totalling 331 words — still scoring 0.9897, still flagged, with 39.6% of the original four-word sequences intact. One sample is not a rate, and the free quota ended before a controlled run could follow. It is published as one direct observation and nothing more.

JustDone was not measured, on purpose. Its Terms of Use prohibit the automated access the evaluation required, in terms that working slowly or discarding the output would not cure. The evaluation stopped before a single document was submitted; the harness proof, the selected sample and the reference profile are preserved so that one lawful script run completes it. Until then, JustDone's paid mode is recorded as untested against this build — a prediction either way would be a guess wearing a percentage.

04 / The popular attack

Asking an LLM to reword AI text makes it easier to catch, not harder

The attack most people can actually mount is not a specialist service. It is pasting the text into a chatbot and asking for a rewrite. That attack was measured on a purpose-built paired corpus: 600 source documents — 300 AI-written, 300 human-written — each rewritten at three strengths (copy-edit, structural paraphrase, full rewrite) by five different models, giving 2,302 rows in total, every rewrite carrying the lineage of its source.

Measured at a rule that no longer ships

Corpus
2,302 rows: 600 sources (300 AI, 300 human) and 1,702 kept rewrites at three strengths across five rewriting models, one family held out. Median source length 372 words — far shorter than the corpus behind the headline rates, which is why no figure here may be read against them.
Retired flag point
0.9855 / 0.9763
Detector
tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d (cycle 2, superseded 1 September 2026)
Runtime
Server route, tier3-cycle2-e5small-fp32.onnx, segments-v3, cycle-2's live pair 0.9855/0.9763 — retired 1 September 2026. Harness re-proved against 883/922 and 45/4,636 before any new figure was taken. The matched browser subset agrees within about a point on every cell.
Measured
30 August 2026
Also
Every row is generic LLM paraphrase — commercial_humaniser: false. Nothing in this section describes what the services in Figure 1 do. Not re-measured against tier3-cycle5-v1, which reads the heavy-rewrite-of-human-original axis higher on its own held-out corpus — see §5.

The rule that ships today is margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562). The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.

Figure 2 Flag rate by rewrite strength, both directions The same sources, untouched and then rewritten at each strength. The AI line rises — an LLM asked to reword machine prose produces more machine-like prose — while the human line stays near its floor until almost none of the author's wording survives. Denominators differ by strength because 98 failed rewrites were quarantined; the exclusions understate the light-strength AI cells rather than flattering them.
Flag rate on AI originals and human originals, untouched and at three rewrite strengths 0% 25% 50% 75% 100% Untouched Copy-edit Paraphrase Full rewrite 65.3% (196/300) 79.2% (229/289) AI originals flagged server route, tier3-cycle2's pair (retired 1 Sept 2026) 1.0% (3/300) 21.0% (57/272) Human originals flagged server route, tier3-cycle2's pair (retired 1 Sept 2026)
Flag rates by rewrite strength for AI and human originals
StrengthAI originals flaggedHuman originals flagged
Untouched65.3% (196/300)1.0% (3/300)
Copy-edit67.4% (174/258)1.4% (4/290)
Structural paraphrase74.8% (220/294)11.0% (33/299)
Full rewrite79.2% (229/289)21.0% (57/272)

LLM-REWRITE-ROBUSTNESS.md §3.1, §3.3; per-strength AI figures from the fp32 results record. Untouched human false positives on this corpus: 1.0% (3/300).

The paired reading is the strong one. Of the AI sources the build detects, 95.6% of their rewrites are still detected — 526 of 550 — and even at full-rewrite strength, where 9% of the original four-word sequences survive, 92.0% (172 of 187) stay caught. No rewriting model defeated it: the five ranged from 91.9% to 99.1% survival, with the held-out family mid-pack. Humaniser products imply the opposite, and for this class of attack the implication is measured and wrong.

The 65.3% baseline needs its own sentence: these sources are short, at a median of 372 words and one scored section, and length dominates every other variable. The baseline is a property of the corpus, not a revision of any published rate.

05 / The other side of the ledger

One in five heavily rewritten human documents gets flagged — and the rewriter decides

The same corpus answers the question that matters to people who never tried to cheat: what happens when a machine rewrites a person's work? The ladder in Figure 2 behaves correctly at its base — a copy-edit of your own prose is left alone at 1.4%, near the untouched rate, which is deliberate. At the top, where almost nothing of the author's wording survives, 21.0% of documents are flagged. Polish your writing through a chatbot aggressively enough and there is about a one-in-five chance of being called machine-written — which is fair comment on the words, since a machine wrote them, and still a real cost to the person who started with an honest draft.

Figure 3 Flagged rewritten human originals, by rewriting model All strengths pooled, per rewriting model. The same human documents and the same instruction produce a six-fold spread depending only on which model did the rewriting, so any single published number for this case is an average over this range.
Share of rewritten human originals flagged, by rewriting model 0% 5% 10% 15% 20% 25% google/gemini-3.7-flash 39/196 rewritten human originals 19.9% mistralai/mistral-medium-3-5 24/134 rewritten human originals · unseen family 17.9% deepseek/deepseek-v4-pro-0813 13/173 rewritten human originals 7.5% openai/gpt-5.6-luna 12/174 rewritten human originals 6.9% meta-llama/llama-4-maverick 6/184 rewritten human originals 3.3%
Flagged rewritten human originals by rewriting model
Rewriting modelFlagged / rewritesRate
google/gemini-3.7-flash39 of 19619.9%
mistralai/mistral-medium-3-5 (unseen family)24 of 13417.9%
deepseek/deepseek-v4-pro-081313 of 1737.5%
openai/gpt-5.6-luna12 of 1746.9%
meta-llama/llama-4-maverick6 of 1843.3%

LLM-REWRITE-ROBUSTNESS.md §3.4. The mistral family was held out of every split as an unseen family.

The registers where this cost lands hardest are the two already published as the tool's weakest: academic writing at 18.2% (32 of 176) and fiction at 16.9% (29 of 172) on the rewritten-human arm. That is the known weakness reappearing on a new axis rather than a new one.

tier3-cycle5-v1, deployed 1 September 2026, raises this cost — disclosed, not hidden.

On cycle 5's own 1,199-row held-out humaniser-pairs corpus (a different corpus from the 272-row figure above, so the two numbers are not directly subtracted), a heavy LLM rewrite of a human original now flags 28.5% (39/137), up from the superseded cycle-2 model's 21.2% on the same rows — a direct consequence of training heavy rewrites as machine-written, which their words are. Light copy-edits stay clear: 0.7% (1/153). Source: CYCLE5-REPORT.md §Recommendation point 3 and §5.

06 / Limits

What this page does not prove

The commercial corpus is small, dated and incomplete.

33 sources from three generating models, one snapshot run, products that change weekly, and two absent names: JustDone (terms boundary, above) and QuillBot beyond a single direct test. The 96.4% row is conditional on sources the build first caught and on outputs of at least 200 words.

The paired corpus is LLM paraphrase throughout. It measures the attack most users can mount, not what a purpose-built humaniser does — the services in Figure 1 beat this build 96% of the time while ordinary rewriting fails. A training response to the humaniser gap needs humaniser output, which is a licensing question as much as a technical one.

Short documents, short sections. The paired corpus has a median of 372 words and one scored section per document, so its absolute rates sit far below the long-form headline figures and may not be compared with them. The paired, same-lineage comparison — did the rewrite of a caught document escape? — is the only reading that survives the length confound.

Escape is not authorship. An output that falls below the verdict rule has evaded one detector at one operating point. Nothing on this page says such text reads well, passes any other tool, or is safe to publish as human work.

Apply the method

Check a document and compare the result with the humaniser tests.