On this page
01 / Result vocabulary
Method status is explicit
Canonical results use pass, attention, fail, inconclusive, unsupported, not configured, not run and error. Unsupported, missing and failed checks do not share pass styling or counts. The Anthropic production watermark is reported as not assessed, because its keys are private and no public verifier exists; that boundary is stated as its own line rather than dressed up as a check.
02 / Evidence types
Exact checks and writing patterns are different
Unicode inspection reports named code points, exact UTF-16 and code-point offsets, source-bound hashes and a limited treatment. Writing-pattern rules are versioned editorial prompts: the current signal set, en-signals:2026.08.6, runs 116 named rules across 113 weighted categories, covering stock phrasing, structure, sentence rhythm and cadence, chat-export formatting furniture, and chatbot artefact traces attributed to the model families that leave them. They do not estimate authorship probability.
Uploaded image and PDF files receive a separate local Content Credentials (C2PA) read. A trained beta classifier adds the one check here that gives an AI reading. It scores on Opace's own EU server by default, one request for the whole document, or entirely in your browser after an explicit one-off 34.5 MB download, and its measured accuracy is disclosed with every result. Every assessment also runs the published SynthID-Text detection mathematics against three public demo keys in the browser and reports the score for each key, with fewer than 40 scoreable positions returning no verdict at all.
03 / Protected content
Protected facts block unsafe changes
Names, organisations, figures, dates, times, units, links, email addresses, quotations, citations and code can be recorded as protected spans. Candidate text fails a hard gate when an exact protected item disappears, changes unexpectedly or is duplicated.
04 / Deliberate changes
Safe treatments require selection
The browser can preview only allowlisted changes for a selected finding. It does not remove joiners or combining marks automatically, replace lookalikes, rewrite prose, or alter protected links, quotations, citations and code.
05 / Evidence receipts
Receipts record evidence without retaining the draft
Hash-only receipts are the default. They use RFC 8785 JSON Canonicalization Scheme and SHA-256, record method versions and preserve content only when an explicit retain-content policy agrees with the receipt state.
06 / Data routes
Privacy is a route property
Browser-only inspection holds text in memory and sends no content-bearing request. Text, file names, URLs, hashes, findings, protected spans and receipt IDs do not enter analytics, browser storage or the page address.
The trained model is the one check with a choice of route, and the page says which one is running where you paste rather than in the small print: by default the document is sent over HTTPS in a single request to Opace's own server in Google Cloud Run europe-west1, scored in memory and then discarded, and every result prints how many words were sent. The browser route sends nothing at all; its download size, cache behaviour and local-only execution are stated before consent is requested, and nothing fetches until the user agrees. Interacting with the form fires a standard site analytics event that carries no text, file or result content.
07 / Reproduction limits
Limits and reproduction
The method cannot prove authorship. A result applies only to the named method and version that ran. Public watermark fixtures also require an immutable source, licence, known configuration and matched detector version before they can appear in the Claude Watermark Readiness Lab.
08 / Published results and evidence
Measured results, stated plainly
Correction, 29 August 2026: two withdrawn measurements.
The rule tier's 66.7% figure and the trained model's zero false positives at 0.857 are both withdrawn and must not be quoted. What replaces them is below, each with its corpus, its flag point and its denominators.
What was published, and why each was withdrawn
The figures this section carried between 26 and 28 August came from two measurements that have since been superseded, and both are withdrawn. The rule tier's 66.7% detection at zero false positives on 169 human documents, and the twelve-slice per-provider table published with it, were measured on a corpus whose human half was 76% encyclopaedic and question-and-answer text: a register these rules barely react to. Re-tested on representative long-form writing the same rules flag roughly one human document in four. The trained model's zero false positives across 116 verified human texts at a 0.857 threshold described its first training cycle, which no longer ships. Please do not quote either figure.
Correction, 30 August 2026: the flag point on this page was wrong, and the corpus is not fully held out.
The rule that ships is the minimum-evidence pair 0.9855 / 0.9763, not the single 0.984 threshold this page described. And 268 of the 922 AI documents were already seen by the model in a cycle-2 split.
Both corrections in full, with the measured effect of each
Until 30 August 2026 this page described a single 0.984 threshold and quoted 877 of 922 AI documents and 56 of 4,636 human ones. That rule is superseded, it no longer ships, and those are not the shipped figures. Every measured number on this page is now read from the file the browser fetches, so neither this page nor the detection-rates page beside it can hold its own copy again.
Corrected 30 August 2026. This corpus was published as held out and hash-quarantined against every training split. It is not, for the AI half: of the 922 AI documents, 654 are independent of every cycle-2 split and 268 are not (168 in the training split, 72 in test, 28 in calibration). The human half is effectively independent: 11 of 4,636 appear in a cycle-2 split. The effect was measured at the superseded 0.984 single-threshold rule, which is the only operating point the split has ever been scored at: independent 620/654 (94.80%), seen-in-cycle-2 257/268 (95.90%), of which the training split alone 163/168 (97.02%) — a gap of 1.1 percentage points. It has NOT been measured at the shipped 0.9855/0.9763 pair, so no seen-against-unseen split is published for the operating point that ships, and the 1.1-point figure must not be quoted under a shipped-pair heading. Source: services/local-engine/research/corpus-reconciliation-2026-08-29/analysis.txt, section 2.
Both tiers were re-measured in August 2026 against 5,558 long-form documents, 922 written by current models and 4,636 by people. 654 of the AI documents are independent of every training, test and calibration split; 268 are not, as are 11 of the human documents. Every rate below pools both subsets, and the full breakdown is on the detection-rates page.
The trained model, which produces the AI reading
At the operating point that ships — a document is flagged when its strongest section reaches 0.9855, or its second-strongest reaches 0.9763 — the EU server route flags 883/922 AI documents, 95.8%, and 45/4,636 human documents, 0.97%.
Loosening the rule to a single 0.980 threshold buys 10 more AI documents and costs 52 more human ones: 893/922, 96.9%, at 97/4,636, 2.09%. We publish both ends of that trade because the second number is the one that lands on a real writer.
The browser route runs the same model at the same operating point in lower numerical precision, and it is not a copy of the server's result. At the shipped pair it flags 889/922 AI documents, 96.4%, and 90/4,636 human documents, 1.94% — it catches a little more AI writing and wrongly flags roughly twice as many people. Each route therefore reports the accuracy measured on the runtime that produced it, and never borrows the other's.
| Route and operating point | AI documents flagged | Human documents wrongly flagged |
|---|---|---|
| EU server, fp32, shipped pair 0.9855 / 0.9763 | 883/922 (95.8%) | 45/4,636 (0.97%) |
| Browser, int8, shipped pair 0.9855 / 0.9763 | 889/922 (96.4%) | 90/4,636 (1.94%) |
| EU server, fp32, superseded single 0.980 | 893/922 (96.9%) | 97/4,636 (2.09%) |
Server route: Python onnxruntime fp32, the reference-server scoring path. Browser route: the int8 runtime that runs in your browser. Both over the whole 5,558-document corpus.
Where it is weakest
Its weakest ground is published rather than buried. The model was never trained on human fiction, because the corpus has AI fiction and no matched human set, so novelists should not rely on it. Short passages lose most of the signal: measured on naturally short text at the same operating point, the classifier flags 29/172 (16.9%) of AI passages of 100 to 199 words against 193/228 (84.6%) at 300 to 399. Under 100 words, and 200 to 299 words, hold fewer than 30 AI passages each, so no detection rate is quoted for either band. Treat a draft under about 300 words as unreliable. AI text that has been rewritten from a human original is the hardest case of any register measured. Business and white-paper writing clears the floor on a set too small to call it settled. Every one of those is measured in the open repository rather than asserted here.
Withdrawn on 30 August 2026: this paragraph previously said detection fell to 67.0% at 200 words, 50.3% at 150 and 19.0% at 100. Those figures came from long documents truncated to those lengths, were scored at the superseded 0.980 threshold, recorded no per-length AI denominator and were never re-measured on either shipping runtime. They must not be quoted.
The writing-signal rules, which no longer produce one
The rules reached mixed signals or above on 45.1% of the AI writing while flagging 24.8% of the human writing — worse than the trained model on both counts at once.
Show the numbers this was drawn from
| Tier | AI writing flagged | Human writing flagged | Corpus |
|---|---|---|---|
| The trained classifier | 883/922 (95.8%) | 45/4,636 (0.97%) | Whole corpus, shipped pair, fp32 |
| The 113 writing-signal rules | 45.1% | 24.8% | a 5,558-document long-form corpus; these rules have no training set, so none of it was fitted to them, 922 AI and 1,200 human documents |
Trained classifier: whole 5,558-document corpus at the shipped pair, fp32. Writing-signal rules: 922 AI and 1,200 human documents from the same corpus. The two denominators differ and are printed rather than reconciled.
Two conditions travel with that figure, and we publish them wherever it appears. First, much of what these rules catch is chat-export formatting: bold runs, heading lines and dense bullet layouts. Text pasted through an editor that strips those markers loses most of the signal, and hand-cleaned plain prose falls back to the smaller subset that survives. Second, they read register rather than authorship. Short conversational question-and-answer text measurably carries almost none of the style tells they look for, whatever produced it, and genuine human marketing copy triggers the cliché-vocabulary rules routinely. That is the mechanism behind the 24.8%.
A measured model we decided not to ship
During the first training cycle we built a second model, a GPT-2 surprisal-rhythm engine, measured it, and declined it. Ensembled with the classifier of the day it lifted that cycle's small evaluation from 2 of 23 clean-prose passages to 6, and from 2 of 30 AI samples to 8, but it added 2 false positives across that cycle's 116 human texts, 1.73%, and raised the one-off browser download from 34.5 MB to 238.8 MB. Those are first-cycle numbers on a first-cycle corpus and they measure the decision, not today's accuracy. It remains gated off. We record it because declining to ship a measured capability is part of the method, not an omission from it.
The full tables are published rather than summarised
Detection rates in full carries every measured cell — by document length, by the model that wrote the text, and by content type — each with its denominator, its corpus, its runtime, its operating point and a 95% confidence interval. That is where the fiction weakness, the short-text collapse and the bands with too few documents to quote are all visible at once.
Full reproducible reports ship with the open-source repository: the segmentation and threshold measurements behind the figures above, the per-register breakdowns, the route-parity comparison, the calibration corpus documentation, the machine-readable model evaluation and the complete test-suite evidence index. Any figure on this page can be re-run rather than taken on trust.
09 / The signal pack, measured
What the 116 rules are, and which of them work
The pack is counted, not estimated: 116 named rules across 113 weighted categories, counted from the built packs rather than from prose, and a test fails the build if the constant and the packs ever disagree. But a count of rules is a capability claim, and a capability claim without a measurement is what this project keeps having to retract. So every named rule is scored against two corpora and the register is published whole, including the rules that never fire and the ones that point the wrong way.
These rules produce no AI verdict, and nothing in this section changes that.
Since 28 August 2026 the pack contributes nothing to the AI reading. It is editorial feedback on phrasing and structure. The figures below describe what the rules fire on, not what anything concludes.
The denominators here are not the detection corpus. Rule liveness is measured over 5,743 AI documents and 4,353 human documents drawn from five corpora, which is a different population from the 5,558-document corpus every classifier figure on this page comes from. The two are never pooled and never compared cell to cell.
The five corpora behind every count in this section
- generated-2026-08 — 4,016 AI documents. 4,016 usable current-model long-form articles, 21 models, 10 providers, generated by this project (ours to publish). Published register — the register users actually paste.
- provider-eval-ai — 1,727 AI documents. The 1,727 AI documents of the 1,896-sample provider-eval set. Chat-reply register, which is why it under-reports rules that measure published-prose cadence.
- provider-eval-human — 169 human documents. The 169 held-out human documents of the provider-eval set.
- human-corpus-v2 — 4,144 human documents. 4,144 modern human samples, 1,233 of them business and marketing copy. Test-only: its manifest forbids training use.
- human-corpus-v1 — 40 human documents. The 40-sample genre-matched verified-human calibration corpus.
Register generated by tests/battery/rule-liveness.mjs at en-signals:2026.08.6, measured 2026-08-30.
How many of the 116 can fire at all
| Rule | State | Recorded reason |
|---|---|---|
tier3-phrase-cluster | inactive | Cannot fire on realistic English prose. The gate needs 3 or more DISTINCT phrases from TIER3_PHRASES in one document; the measured maximum across 10,096 documents (5,743 AI, 4,353 human) is 1. The list is inherited crypto/web3 whitepaper vocabulary: 'decentralized compute', 'reward emissions', 'tokenized incentive structures' and 'emerging sector/space/category/industry' match no document in any corpus, and the only two that match anything ('the integration of', 'the intersection of') are register-neutral English that fires on humans at a comparable rate. Not counted as a live capability. See docs/CAPABILITIES.md 3.4a. |
transition-density | dormant-register | Needs more than 4 transition words per 100 words over at least 40 words. signals.transition catches the same texts at a lower density first, so the stricter en-gb gate is never the binding one on real documents. Probe-verified reachable. |
ai-citation-markup | dormant-forensic | Exposed chatbot citation markup (citeturn / oaicite forms). Provenance marker, not a style rule: it fires only on text pasted straight out of a chat interface without cleaning. Near-zero false-positive risk, kept as insurance. Probe-verified reachable. |
ai-citation-token | dormant-forensic | Leaked citation tokens ([web:N], glyph-marker forms). Same provenance rationale as ai_citation_markup. Probe-verified reachable. |
ai-utm-source | dormant-forensic | utm_source=chatgpt.com-style URL fingerprints. Provenance marker; corpora are prose, not link-carrying web copy. Probe-verified reachable. |
placeholder-token | dormant-forensic | Unfilled placeholders (INSERT_NAME_HERE). Fires on unedited generated drafts; every corpus document is a finished sample. Probe-verified reachable. |
math-alphanumeric | dormant-forensic | Mathematical-alphanumeric character-set leakage. Provenance marker. Probe-verified reachable. |
pua-character | dormant-forensic | Private Use Area character leakage. Provenance marker. Probe-verified reachable. |
reasoning-artifact | dormant-forensic | Reasoning-trace leaks ('let me think step by step'). Fires on chat exports, not on published prose. Probe-verified reachable. |
rhetorical-question | dormant-register | Fired on 1 of 4,353 human documents and 0 of 5,743 AI. The corpora are articles and reports; the direct-address rhetorical question belongs to blog and social register. Probe-verified reachable. |
rhetorical-qa | dormant-register | The 'The result? ... The catch? ...' question-answer cadence. Absent from the measured registers. Probe-verified reachable. |
future-narrative | dormant-register | Fired on 1 of 4,353 human documents and 0 of 5,743 AI. Speculative-decade framing; absent from the measured registers. Probe-verified reachable. |
despite-challenges-arc | dormant-register | The 'despite these challenges ... continues to thrive' narrative arc. Absent from the measured registers. Probe-verified reachable. |
legacy-framing | dormant-register | 'indelible mark', 'enduring legacy' framing. Absent from the measured registers. Probe-verified reachable. |
narrative-cliche | dormant-register | 'poignant reminder'-class narrative cliches. Absent from the measured registers. Probe-verified reachable. |
notability-canned | dormant-register | Canned encyclopaedic notability phrasing. The corpora contain no encyclopaedia articles. Probe-verified reachable. |
kobak-density | dormant-register | The Kobak et al. excess-vocabulary set at density. A vocabulary rule with a documented decay problem: the tells it names were 2023-24 markers. Probe-verified reachable, but treat a fire as weak evidence. |
fiction-claudeism | dormant-register | Fiction-specific tells ('ministrations', 'despite herself'). The measured corpora hold almost no fiction, which is also the register with the project's worst human false-positive rate. Probe-verified reachable; untested where it matters. |
transition-stacking | dormant-register | Fired on 1 of 4,353 human documents and 0 of 5,743 AI. Requires consecutive paragraphs each opening on a stacked transition. Probe-verified reachable. |
directive-colon-bullets | dormant-register | Imperative-verb bullets with a colon gloss. Needs list structure the corpora largely lack. Probe-verified reachable. |
invalid-isbn | dormant-forensic | ISBN-13 checksum failure — a fabricated-citation marker. Probe-verified reachable, with a negative control proving a valid ISBN does not fire. |
Every dormant and inactive rule carries a recorded category and reason, listed under the figure. A rule that starts or stops firing fails the engine's own liveness test, so this register cannot go stale in either direction.
Which rules actually separate the two populations
A rule's likelihood ratio is how much more often it fires on AI writing than on human writing, both as rates over their own denominators. It is multiplicative, so it is drawn on a log axis: a rule at 0.1 is exactly as wrong as a rule at 10 is right, and a linear axis would draw the first as a rounding error.
| Rule | Fired on AI documents | Fired on human documents | Likelihood ratio |
|---|---|---|---|
formatting | 2,399 of 5,743 | 1 of 4,353 | 1818.36× |
markdown-bold | 2,900 of 5,743 | 2 of 4,353 | 1099.05× |
markdown-heading | 2,726 of 5,743 | 11 of 4,353 | 187.84× |
markdown-furniture | 4,092 of 5,743 | 20 of 4,353 | 155.08× |
uniform-list-items | 715 of 5,743 | 11 of 4,353 | 49.27× |
chatbot | 151 of 5,743 | 4 of 4,353 | 28.61× |
hashtag-stuff | 147 of 5,743 | 4 of 4,353 | 27.86× |
ai-placeholder | 95 of 5,743 | 3 of 4,353 | 24× |
conditional-compression | 88 of 5,743 | 3 of 4,353 | 22.23× |
arrow-decoration | 43 of 5,743 | 2 of 4,353 | 16.3× |
The top of this list is formatting, not prose. markdown-bold, markdown-heading and markdown-furniture measure chat-export layout — text pasted through an editor that strips those markers loses most of it, which is the mechanism behind the tier's demotion.
Six rules run backwards
Six shipped rules fire more often on human writing than on AI writing. They are still in the pack, and they are published rather than quietly dropped.
| Rule | Fired on AI documents | Fired on human documents | Likelihood ratio |
|---|---|---|---|
parenthetical-hedge | 2 of 5,743 | 14 of 4,353 | 0.11× |
quote-inconsistency | 29 of 5,743 | 116 of 4,353 | 0.19× |
passive-ratio | 31 of 5,743 | 91 of 4,353 | 0.26× |
low-specificity | 24 of 5,743 | 62 of 4,353 | 0.29× |
adjacent-lemma-repeat | 473 of 5,743 | 932 of 4,353 | 0.38× |
tier1-clarity | 623 of 5,743 | 991 of 4,353 | 0.48× |
Selected from the per-rule validation as the backwards rules that are statistically supported. Several further rules sit below 1.0 on single-figure counts where the ordering is noise, and they are not drawn as findings.
And one of them was a small-corpus artefact
token-cutoff fires on text that names a model's training cut-off. It was published as one of the clearest backwards rules in the pack. Re-measured on a corpus twenty-five times larger on the human side, it points the right way and is one of the more discriminating rules there is. Both readings are drawn below, because the pairing is the finding: it is a direct measurement of what a 169-document human corpus did to a figure this project published.
| Reading | AI documents | Human documents | Likelihood ratio |
|---|---|---|---|
| As published, provider-eval set — withdrawn | 10 of 1,727 | 6 of 169 | 0.16× |
| Re-measured, current corpora | 232 of 5,743 | 22 of 4,353 | 7.99× |
The withdrawn row is a retracted figure from the 1,896-sample provider-eval set, whose human half was 169 documents. It is reproduced here so the retraction can be read, and it must not be quoted as current.
The technique families, end to end
What the checker actually runs, by family, with what each one is allowed to conclude. The right-hand column is the part that matters: only one family produces an AI reading, and it is not any of the rule families.
| Tier | Family | What it measures | What it may conclude |
|---|---|---|---|
| A — deterministic | Invisible-character carriers | Named code points with exact UTF-16 and code-point offsets: zero-width joiners, non-joiners, spaces and other invisible carriers, each reported as a located span rather than a score. | Exact evidence |
| A — deterministic | Homoglyph and lookalike analysis | Characters from one script disguised as another, located and named. A normalisation pre-pass swaps them before pattern matching, with an offset map so every span still addresses the original text. | Exact evidence |
| A — deterministic | Protected content | Twelve span kinds — names, organisations, figures, dates, times, units, links, email addresses, quotations, citations, code and identifiers — extracted and hard-gated against unsafe rewriting. | Exact evidence |
| A — deterministic | Provenance, C2PA | Content Credentials read locally from uploaded JPEG, PNG, WebP and PDF files through the official CAI SDK. Certificate trust lists are deliberately not consulted, and the interface says so. | Exact evidence |
| B — editorial | Phrase and lexical rules | Tier 1/2/3 vocabulary, stock phrasing, hollow intensifiers, significance and novelty inflation, false concessions, template sentences, the "isn't just X, it's Y" contrast. | Editorial suggestions only |
| B — editorial | Structural rules | Uniform section lengths, uniform list items, heading inflation, bold-label bullets, repeated openings, directive colon bullets. | Editorial suggestions only |
| B — editorial | Stylometric and rhythm rules | Sentence-length flatline, cross-paragraph burstiness, function-word trigram entropy, punctuation distribution, type-token ratio, em-dash density, adjacent lemma repetition. | Editorial suggestions only |
| B — forensic | Chat-export artefacts | Chatbot citation markup, leaked citation tokens, AI URL parameters, unfilled placeholders, reasoning-trace leaks, private-use-area and mathematical-alphanumeric leakage, ISBN checksum failure. | Provenance markers, dormant on prose |
| C — trained | The classifier | The one check that gives an AI reading. Segmented at segments-v3, temperature-calibrated, scored on the EU server by default or in the browser on consent. | The AI reading |
| C — watermark | SynthID-Text key scan | The published SynthID-Text detection mathematics run in the browser against three public demo keys, with a per-key score. Anthropic production keys are private, so that watermark is reported as not assessed. | Per-key score, never a verdict |