Get in Touch
AI content toolsfrom Opace

Measurement paper · rule pack en-signals:2026.08.6

116 writing rules, and six that point the wrong way

The tool ships 116 named writing rules: passive voice, hedging, quote inconsistency, adjacent lemma repetition and a hundred and twelve others. None of them counts towards the authorship verdict. Six of them fire more often on human writing than on machine writing.

On this page
  1. The finding
  2. How to read a likelihood ratio
  3. Why the tier stopped counting
  4. The six that run backwards
  5. The figure that is withdrawn
  6. The same result, independently
  7. What this does not prove
Published
30 August 2026
Measured
Rule liveness re-measured 30 August 2026
Rule pack
en-signals:2026.08.6 · 116 named rules across 113 weighted categories
Corpus
5,743 AI and 4,353 human documents for the firing rates; 922 AI and 4,636 human for the classifier

01 / The finding

The writing rules are editorial advice, and they are not detection

The rules were removed from the authorship verdict on 28 August 2026, on the day the whole tier was measured against the trained classifier and lost on both axes at once. They still run, they still produce suggestions, and a reader of the checker still sees them. What they no longer do is move a verdict.

Six named rules fire more often on human writing than on machine writing, at Benjamini–Hochberg q < 0.05. Following those six as advice makes prose look less like the human documents in these corpora, not more.

Measured on 5,743 AI and 4,353 human documents, rule pack en-signals:2026.08.6. Source: tests/battery/rule-liveness.json.
Named rules in the pack116

Across 113 weighted categories. Ninety-five of them fired on at least one of 5,743 AI documents. None contributes to the reading.

Rules that run backwards6

They fire on human prose more often than on machine prose, at q < 0.05. Two of them fire on roughly one human document in five.

Withdrawn rather than dropped0.16 → 8.0

token-cutoff was recorded as the clearest backwards rule in the pack on 169 human documents. On 4,353 it points the other way, and strongly.

Every figure below belongs to one rule pack version, en-signals:2026.08.6. That includes the count in the title. A pack bump changes which rules exist, which fire and how often, and this page would have to be re-measured rather than edited.

02 / Reading the figures

How to read a likelihood ratio on this page

A rule’s likelihood ratio is its firing rate on AI documents divided by its firing rate on human documents. Above 1.0 the rule fires more often on machine writing. Below 1.0 it fires more often on people. Counts are documents on which the rule fired at least once, not occurrences.

These are firing rates over a corpus, not detector outputs. No operating point applies to them and none is printed. The only detector measurement on this page is the classifier comparison in the next section, and it carries its own conditions.

Likelihood ratio
AI firing rate divided by human firing rate, on 5,743 AI and 4,353 human documents. A ratio of 0.11 means the rule fires roughly nine times more often on a human document than on a machine one.
q < 0.05
Benjamini–Hochberg false-discovery control across the whole pack, which is what stops a hundred and sixteen simultaneous tests manufacturing significant results.
Rule pack en-signals:2026.08.6
The versioned set of rule definitions and thresholds these rates were measured against: 51 v2 rules, 55 v3, 7 v4 rhythm rules and 3 en-gb-v1 rules.

Measurement conditions

Corpus
5,743 AI documents (4,016 generated long-form articles from 21 models across 10 providers, plus the 1,727 AI documents of the provider-eval set) and 4,353 human documents (4,144 modern samples, test-only by manifest, plus the 169 provider-eval humans and the 40-sample genre-matched calibration corpus).
Runtime
No model and no runtime. Every rule figure is a firing rate computed over text.
Measured
30 August 2026
Also
Rule pack en-signals:2026.08.6. Source: tests/battery/rule-liveness.json, measured_utc 2026-08-30.

03 / The demotion

Why the tier stopped counting towards the verdict

Measured on the fresh long-form corpus, the whole rule tier reaches 45.1% detection at a 24.8% human false-positive rate, on 922 AI and 1,200 human documents. A quarter of human writers wrongly accused, for less than half the machine writing caught. That is not a detector at any setting, and the tier became editorial suggestions only.

The 45.1% / 24.8% pair carries no stated operating point.

What rules-score threshold produces it is not recorded anywhere in the measurement repository. Until that gap is closed the pair should be read as a direction rather than as a measurement, and it is published here in that spirit: it is not given a heading of its own, and the caveat travels with it into the figure below.

The classifier side of that comparison has been rebuilt for this page. The figure the demotion decision was originally recorded against — 90.3% detection at 1.34% false positives — is browser measured, from before segmentation existed, at a threshold that no longer runs, and it carries no raw counts. It is not printed here as a current figure. What replaces it is the shipped fp32 measurement at the operating point the tool runs today: 883/922 (95.8%) AI documents detected at 45/4,636 (0.97%) human documents wrongly flagged.

Figure 1 The rule tier against the trained classifier, on both axes Rule pack en-signals:2026.08.6. The rule tier's two rates are measured on 922 AI and 1,200 human documents and have no recorded operating point, which is printed in the right-hand column rather than left to a footnote. The classifier's rates are 883/922 (95.8%) and 45/4,636 (0.97%) on the whole 5,558-document long-form corpus, fp32, at the shipped pair 0.9855 / 0.9763. The two systems are measured on different human denominators and the bars are not a like-for-like scoreboard.
Detection and human false-positive rates for the rule tier and the trained classifier 0% 25% 50% 75% 100% Operating point Rule tier · detection 45.1% of 922 AI documents 45.1% none recorded Rule tier · human false positives 24.8% of 1,200 human documents 24.8% none recorded Trained classifier · detection 883/922 (95.8%) AI documents 95.8% 0.9855 / 0.9763 Trained classifier · human false positives 45/4,636 (0.97%) human documents 0.97% 0.9855 / 0.9763
Rule tier, no recorded operating point Trained classifier at 0.9855 / 0.9763, fp32
Detection and false-positive rates for the rule tier and the trained classifier, with denominators and operating points
SystemDetectionHuman false positivesOperating point
Rule tier, pack en-signals:2026.08.645.1% of 92224.8% of 1,200none recorded
Trained classifier, fp32883/922 (95.8%)45/4,636 (0.97%)0.9855 / 0.9763

Rule tier: docs/MEASURED-FINDINGS.md §4. Classifier: thresholds.json, read at build time through published-figures.ts, so the values on this page cannot drift from the file the browser fetches. The engine repository's docs/assets/charts/rules-vs-model.svg covers the same comparison with a pre-segmentation browser model series at a superseded operating point, and was rebuilt rather than reused.

Measurement conditions

Corpus
The 5,558-document long-form corpus: 922 AI documents and 4,636 human documents. 268 of the 922 AI documents sit in a cycle-2 split, 168 of them in the train split; on the human side 11 of 4,636.
Operating point
0.9855 / 0.9763 · contract segments-v3 · T = 0.8324
Detector
tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d
Runtime
Python onnxruntime 1.29.0, CPU, fp32
Measured
30 August 2026
Also
Applies to the classifier row of Figure 1 only. The rule firing rates elsewhere on this page are feature statistics with no operating point.

04 / Wrong direction

The six that run backwards

Re-measured on 5,743 AI and 4,353 human documents, six named rules fire more often on human prose than on machine prose at q < 0.05.

The two that matter in practice are the last two. adjacent-lemma-repeat and tier1-clarity fire on roughly one human document in five. Repeating a word across adjacent sentences, and using the vocabulary tier1-clarity objects to, are both things human writers do more than models do. Anyone editing to sound less like a machine will be pushed the wrong way by both.

didactic-note also points backwards, at 12 AI documents against 15 human ones. At q = 0.26 it is not distinguishable from chance, so it is listed and not claimed.

Figure 2 Six rules that fire more often on people, and the one figure that reversed Rule pack en-signals:2026.08.6, on 5,743 AI and 4,353 human documents. Bars run left from the neutral line at 1.0, so a longer bar is a rule that fires further in favour of human writing. Counts and denominators are printed under every rule name. The two token-cutoff rows are the same rule measured twice: the greyed, struck row is the withdrawn original on 1,727 AI and 169 human documents, and the row beside it is the re-measurement. Likelihood ratio on a log axis; these are firing rates, so no operating point applies.
Likelihood ratio of eight rule firings, on a log axis around the neutral value of 1.0 0.10 0.25 0.50 1 2 4 8 12 parenthetical-hedge 2 / 5,743 AI · 14 / 4,353 human 0.11 quote-inconsistency 29 / 5,743 AI · 116 / 4,353 human 0.19 passive-ratio 31 / 5,743 AI · 91 / 4,353 human 0.26 low-specificity 24 / 5,743 AI · 62 / 4,353 human 0.29 adjacent-lemma-repeat 473 / 5,743 AI · 932 / 4,353 human 0.38 tier1-clarity 623 / 5,743 AI · 991 / 4,353 human 0.48 token-cutoff (withdrawn) 10 / 1,727 AI · 6 / 169 human · original corpus 0.16 withdrawn token-cutoff (re-measured) 232 / 5,743 AI · 22 / 4,353 human 8.0 1.0 — fires equally on both
Below 1.0: fires more often on human writing Above 1.0: fires more often on AI writing Withdrawn, superseded by a larger corpus
Likelihood ratios, firing counts and denominators for the six backwards rules and for token-cutoff measured twice
RuleAI documentsHuman documentsLikelihood ratio
parenthetical-hedge2 / 5,743 = 0.03%14 / 4,353 = 0.32%0.11
quote-inconsistency29 / 5,743 = 0.50%116 / 4,353 = 2.66%0.19
passive-ratio31 / 5,743 = 0.54%91 / 4,353 = 2.09%0.26
low-specificity24 / 5,743 = 0.42%62 / 4,353 = 1.42%0.29
adjacent-lemma-repeat473 / 5,743 = 8.24%932 / 4,353 = 21.41%0.38
tier1-clarity623 / 5,743 = 10.85%991 / 4,353 = 22.77%0.48
token-cutoff, withdrawn original10 / 1,727 = 0.58%6 / 169 = 3.55%0.16, withdrawn
token-cutoff, re-measured232 / 5,743 = 4.04%22 / 4,353 = 0.51%8.0

tests/battery/rule-liveness.json, measured_utc 2026-08-30, verified row by row against the raw file. The withdrawn original is from research/rule-validation/RULE-VALIDATION.md, 1,727 AI and 169 human documents.

Ten of the seventeen rules recorded as backwards in the original measurement reversed direction on the larger corpora. not-just-contrast went from 0.18% to 2.05% of AI documents, a likelihood ratio of 11.2; normalization-flag reads 6.6, and hollow-intensifier 2.7.

05 / Withdrawn

The figure that is withdrawn

token-cutoff fires on text that names a model’s training cut-off. It had been recorded as one of the clearest backwards rules in the pack: 3.55% of humans, 6 of 169, against 0.58% of AI, 10 of 1,727, for a likelihood ratio of 0.16 at q = 0.012. Re-measured on a corpus roughly three hundred times larger on the human side, it points the right way and it is one of the more discriminating rules in the pack, at 232 of 5,743 AI documents against 22 of 4,353 human ones.

The original rested on six human documents. Six documents can produce a likelihood ratio with a respectable-looking q value and no stability at all.

The 0.16 figure should not be quoted, and it is not published here as a finding. The same 169-document human corpus is on record in this project as having caused three separate claims to be retracted.

It is shown struck through in Figure 2 rather than quietly deleted. How a project handles a figure that will not reproduce says more than the figure did.

06 / Corroboration

The same conclusion from an independent direction

A transparent 24-feature scorecard was fitted twice, on exactly the data the deployed model was trained on, and evaluated on exactly the data it was validated on: 793 machine and 4,179 human fresh long-form documents. Once with every feature available, and once with every formatting feature withheld.

The formatting-free version is better. Detection at a 1% false-positive budget rises from 62.5% to 72.1%, and on creative writing specifically from 28.2% to 62.1%. On a separate human population of 3,767 modern web, business, marketing, academic and non-native documents, false positives at a 5% budget fall from 3.61% to 1.06%.

Figure 3 Detection at four false-positive budgets, with and without formatting features Two fits of the same 24-feature scorecard, evaluated on 793 machine and 4,179 human fresh long-form documents. The prose-only fit is better at the two strict budgets, where a published tool has to operate, and the two converge by the 5% budget. This is the same result the rule tier produced under pack en-signals:2026.08.6, reached from the opposite direction: withholding hand-written formatting and phrasing signals improved the model.
Detection rate at 1, 2, 3 and 5 per cent false-positive budgets for the unrestricted and prose-only scorecards 0% 25% 50% 75% 100% 1% budget 793 machine · 4,179 human 62.5% 72.1% 2% budget 793 machine · 4,179 human 77.6% 81.1% 3% budget 793 machine · 4,179 human 84.9% 85.4% 5% budget 793 machine · 4,179 human 89.9% 88.8%
Every feature available Formatting features withheld
Detection rate by false-positive budget for the unrestricted and prose-only scorecards, on 793 machine and 4,179 human documents
False-positive budgetEvery featureFormatting withheld
1% budget62.5%72.1%
2% budget77.6%81.1%
3% budget84.9%85.4%
5% budget89.9%88.8%

SIGNAL-SCIENCE.md §3. Detection is measured on 793 machine and 4,179 human documents; the false-positive figures in Figure 4 are on a different population of 3,767 human documents and the two are not pooled.

Figure 4 False positives on an independent human population of 3,767 documents The same two fits, checked on a human population they were not tuned against: 3,767 modern web, business, marketing, academic and non-native documents. The prose-only fit is worse at the 1% budget and better at every budget above it, ending at 1.06% against 3.61%. These are a different population and a different denominator from Figure 3, and the two sets of rates are not pooled. Rule pack en-signals:2026.08.6 governs the rule figures elsewhere on this page, not this scorecard.
False-positive rate on 3,767 independent human documents at four false-positive budgets 0% 1% 2% 3% 4% 1% budget 3,767 human documents 0.13% 0.27% 2% budget 3,767 human documents 0.96% 0.42% 3% budget 3,767 human documents 1.67% 0.66% 5% budget 3,767 human documents 3.61% 1.06%
Every feature available Formatting features withheld
False-positive rate on 3,767 independent human documents by budget, for the unrestricted and prose-only scorecards
False-positive budgetEvery featureFormatting withheld
1% budget0.13%0.27%
2% budget0.96%0.42%
3% budget1.67%0.66%
5% budget3.61%1.06%

SIGNAL-SCIENCE.md §3. Independent human population, 3,767 documents. The budget on the left is the budget the threshold was set to on the evaluation set, not the rate achieved on this population, which is what the bars report.

Hand-written formatting and phrasing signals detect register and provenance. Forbidding them made a detector better on both axes at once.

07 / Limits

What this does not prove

The AI and human sides are different corpora, not two halves of one. The AI side is dominated by 4,016 generated long-form articles; the human side by 4,144 modern samples weighted towards business and marketing copy. Register differs between them, and register confounding authorship is this project’s central finding. These likelihood ratios describe how a rule behaves across these corpora. They are strong enough to say a rule points the wrong way. They are not strong enough to fix a weight.

A rule that does not fire is not the same as a rule that is wrong. Of the 116, 95 fired on at least one of 5,743 AI documents. tier3-phrase-cluster is recorded as inactive because it cannot fire on realistic English prose, and twenty more are dormant on every corpus measured. Some of those twenty are deliberate forensic markers kept as insurance, such as leaked citation tokens and unfilled placeholders, and their absence from a prose corpus is the expected result rather than a failure.

One proposed loosening was measured and rejected. The proposed gate for punchline-fragment-density reaches zero documents on the long-form corpus: the AI corpus maximum punchline rate is 0.136 and its 99th percentile 0.084, against a shipped gate of 0.18 and a proposed gate of 0.10. It is recorded as rejected so that it cannot be mistaken for a pending improvement.

These are English rules measured on English prose. Nothing here is measured in any other language.

Two open items travel with this page.

The 45.1% / 24.8% rule-tier pair has no recorded operating point, and should not be moved into a headline position until one exists. The 90.3% / 1.34% classifier figure the original comparison was made against is superseded: browser measured, pre-segmentation, at a threshold that no longer runs, with no raw counts. It is named on this page only to say that it is not the current figure.

The long-form corpus is not fully held out.

268 of its 922 AI documents sit in a cycle-2 split, 168 of them in the train split; on the human side 11 of 4,636. That overlap is immaterial to the rule firing rates, which involve no model and no training, and it is not immaterial to the classifier row of Figure 1.

Provenance

Source file and section for every figure on this page
FigureFileSection
Rule pack en-signals:2026.08.6: 116 named rules across 113 weighted categoriesen-signals pack manifest
Likelihood ratios, firing counts, 5,743 AI and 4,353 human denominatorstests/battery/rule-liveness.jsonmeasured_utc 2026-08-30
Original 1,727 AI and 169 human measurement, including token-cutoff 0.16research/rule-validation/RULE-VALIDATION.mdwhole record
Rule tier 45.1% / 24.8%, and the 28 August 2026 demotiondocs/MEASURED-FINDINGS.md§4
Classifier 883/922 (95.8%) at 45/4,636 (0.97%), fp32, 0.9855 / 0.9763thresholds.json via published-figures.tsserver_fp32_segmented
24-feature scorecard, both fits, both populationsresearch/signal-science/SIGNAL-SCIENCE.md§3
268/922 and 168 train-split overlap; human 11/4,636corpus-reconciliation-2026-08-29/analysis.txt§2

One date is unreconciled at publication: docs/MEASURED-FINDINGS.md §4 dates the re-measurement 29 August 2026, while tests/battery/rule-liveness.json records measured_utc 2026-08-30. This page uses the measurement file’s own date, 30 August 2026.

Every file named above ships with the open measurement repository, so any figure on this page can be re-run rather than taken on trust.

Read on

The rules were the readable half. The measured half is next door.