On this page
01 / The finding
The writing rules are editorial advice, and they are not detection
The rules were removed from the authorship verdict on 28 August 2026, on the day the whole tier was measured against the trained classifier and lost on both axes at once. They still run, they still produce suggestions, and a reader of the checker still sees them. What they no longer do is move a verdict.
Six named rules fire more often on human writing than on machine writing, at Benjamini–Hochberg q < 0.05. Following those six as advice makes prose look less like the human documents in these corpora, not more.
Across 113 weighted categories. Ninety-five of them fired on at least one of 5,743 AI documents. None contributes to the reading.
They fire on human prose more often than on machine prose, at q < 0.05. Two of them fire on roughly one human document in five.
token-cutoff was recorded as the clearest backwards rule in the pack on 169 human documents. On 4,353 it points the other way, and strongly.
Every figure below belongs to one rule pack version, en-signals:2026.08.6. That includes the count in the title. A pack bump changes which rules exist, which fire and how often, and this page would have to be re-measured rather than edited.
02 / Reading the figures
How to read a likelihood ratio on this page
A rule’s likelihood ratio is its firing rate on AI documents divided by its firing rate on human documents. Above 1.0 the rule fires more often on machine writing. Below 1.0 it fires more often on people. Counts are documents on which the rule fired at least once, not occurrences.
These are firing rates over a corpus, not detector outputs. No operating point applies to them and none is printed. The only detector measurement on this page is the classifier comparison in the next section, and it carries its own conditions.
- Likelihood ratio
- AI firing rate divided by human firing rate, on 5,743 AI and 4,353 human documents. A ratio of 0.11 means the rule fires roughly nine times more often on a human document than on a machine one.
- q < 0.05
- Benjamini–Hochberg false-discovery control across the whole pack, which is what stops a hundred and sixteen simultaneous tests manufacturing significant results.
- Rule pack
en-signals:2026.08.6 - The versioned set of rule definitions and thresholds these rates were measured against: 51 v2 rules, 55 v3, 7 v4 rhythm rules and 3 en-gb-v1 rules.
Measurement conditions
- Corpus
- 5,743 AI documents (4,016 generated long-form articles from 21 models across 10 providers, plus the 1,727 AI documents of the provider-eval set) and 4,353 human documents (4,144 modern samples, test-only by manifest, plus the 169 provider-eval humans and the 40-sample genre-matched calibration corpus).
- Runtime
- No model and no runtime. Every rule figure is a firing rate computed over text.
- Measured
- 30 August 2026
- Also
- Rule pack en-signals:2026.08.6. Source: tests/battery/rule-liveness.json, measured_utc 2026-08-30.
03 / The demotion
Why the tier stopped counting towards the verdict
Measured on the fresh long-form corpus, the whole rule tier reaches 45.1% detection at a 24.8% human false-positive rate, on 922 AI and 1,200 human documents. A quarter of human writers wrongly accused, for less than half the machine writing caught. That is not a detector at any setting, and the tier became editorial suggestions only.
What rules-score threshold produces it is not recorded anywhere in the measurement repository. Until that gap is closed the pair should be read as a direction rather than as a measurement, and it is published here in that spirit: it is not given a heading of its own, and the caveat travels with it into the figure below.
The classifier side of that comparison has been rebuilt for this page. The figure the demotion decision was originally recorded against — 90.3% detection at 1.34% false positives — is browser measured, from before segmentation existed, at a threshold that no longer runs, and it carries no raw counts. It is not printed here as a current figure. What replaces it is the shipped fp32 measurement at the operating point the tool runs today: 883/922 (95.8%) AI documents detected at 45/4,636 (0.97%) human documents wrongly flagged.
| System | Detection | Human false positives | Operating point |
|---|---|---|---|
| Rule tier, pack en-signals:2026.08.6 | 45.1% of 922 | 24.8% of 1,200 | none recorded |
| Trained classifier, fp32 | 883/922 (95.8%) | 45/4,636 (0.97%) | 0.9855 / 0.9763 |
Rule tier: docs/MEASURED-FINDINGS.md §4. Classifier: thresholds.json, read at build time through published-figures.ts, so the values on this page cannot drift from the file the browser fetches. The engine repository's docs/assets/charts/rules-vs-model.svg covers the same comparison with a pre-segmentation browser model series at a superseded operating point, and was rebuilt rather than reused.
Measurement conditions
- Corpus
- The 5,558-document long-form corpus: 922 AI documents and 4,636 human documents. 268 of the 922 AI documents sit in a cycle-2 split, 168 of them in the train split; on the human side 11 of 4,636.
- Operating point
- 0.9855 / 0.9763 · contract segments-v3 · T = 0.8324
- Detector
tier3-cycle2-e5small-fp32.onnx, SHA-256e313ab00de1fffd2…4d2788d- Runtime
- Python onnxruntime 1.29.0, CPU, fp32
- Measured
- 30 August 2026
- Also
- Applies to the classifier row of Figure 1 only. The rule firing rates elsewhere on this page are feature statistics with no operating point.
04 / Wrong direction
The six that run backwards
Re-measured on 5,743 AI and 4,353 human documents, six named rules fire more often on human prose than on machine prose at q < 0.05.
The two that matter in practice are the last two. adjacent-lemma-repeat and tier1-clarity fire on roughly one human document in five. Repeating a word across adjacent sentences, and using the vocabulary tier1-clarity objects to, are both things human writers do more than models do. Anyone editing to sound less like a machine will be pushed the wrong way by both.
didactic-note also points backwards, at 12 AI documents against 15 human ones. At q = 0.26 it is not distinguishable from chance, so it is listed and not claimed.
| Rule | AI documents | Human documents | Likelihood ratio |
|---|---|---|---|
parenthetical-hedge | 2 / 5,743 = 0.03% | 14 / 4,353 = 0.32% | 0.11 |
quote-inconsistency | 29 / 5,743 = 0.50% | 116 / 4,353 = 2.66% | 0.19 |
passive-ratio | 31 / 5,743 = 0.54% | 91 / 4,353 = 2.09% | 0.26 |
low-specificity | 24 / 5,743 = 0.42% | 62 / 4,353 = 1.42% | 0.29 |
adjacent-lemma-repeat | 473 / 5,743 = 8.24% | 932 / 4,353 = 21.41% | 0.38 |
tier1-clarity | 623 / 5,743 = 10.85% | 991 / 4,353 = 22.77% | 0.48 |
token-cutoff, withdrawn original | 10 / 1,727 = 0.58% | 6 / 169 = 3.55% | 0.16, withdrawn |
token-cutoff, re-measured | 232 / 5,743 = 4.04% | 22 / 4,353 = 0.51% | 8.0 |
tests/battery/rule-liveness.json, measured_utc 2026-08-30, verified row by row against the raw file. The withdrawn original is from research/rule-validation/RULE-VALIDATION.md, 1,727 AI and 169 human documents.
Ten of the seventeen rules recorded as backwards in the original measurement reversed direction on the larger corpora. not-just-contrast went from 0.18% to 2.05% of AI documents, a likelihood ratio of 11.2; normalization-flag reads 6.6, and hollow-intensifier 2.7.
05 / Withdrawn
The figure that is withdrawn
token-cutoff fires on text that names a model’s training cut-off. It had been recorded as one of the clearest backwards rules in the pack: 3.55% of humans, 6 of 169, against 0.58% of AI, 10 of 1,727, for a likelihood ratio of 0.16 at q = 0.012. Re-measured on a corpus roughly three hundred times larger on the human side, it points the right way and it is one of the more discriminating rules in the pack, at 232 of 5,743 AI documents against 22 of 4,353 human ones.
The original rested on six human documents. Six documents can produce a likelihood ratio with a respectable-looking q value and no stability at all.
It is shown struck through in Figure 2 rather than quietly deleted. How a project handles a figure that will not reproduce says more than the figure did.
06 / Corroboration
The same conclusion from an independent direction
A transparent 24-feature scorecard was fitted twice, on exactly the data the deployed model was trained on, and evaluated on exactly the data it was validated on: 793 machine and 4,179 human fresh long-form documents. Once with every feature available, and once with every formatting feature withheld.
The formatting-free version is better. Detection at a 1% false-positive budget rises from 62.5% to 72.1%, and on creative writing specifically from 28.2% to 62.1%. On a separate human population of 3,767 modern web, business, marketing, academic and non-native documents, false positives at a 5% budget fall from 3.61% to 1.06%.
| False-positive budget | Every feature | Formatting withheld |
|---|---|---|
| 1% budget | 62.5% | 72.1% |
| 2% budget | 77.6% | 81.1% |
| 3% budget | 84.9% | 85.4% |
| 5% budget | 89.9% | 88.8% |
SIGNAL-SCIENCE.md §3. Detection is measured on 793 machine and 4,179 human documents; the false-positive figures in Figure 4 are on a different population of 3,767 human documents and the two are not pooled.
| False-positive budget | Every feature | Formatting withheld |
|---|---|---|
| 1% budget | 0.13% | 0.27% |
| 2% budget | 0.96% | 0.42% |
| 3% budget | 1.67% | 0.66% |
| 5% budget | 3.61% | 1.06% |
SIGNAL-SCIENCE.md §3. Independent human population, 3,767 documents. The budget on the left is the budget the threshold was set to on the evaluation set, not the rate achieved on this population, which is what the bars report.
Hand-written formatting and phrasing signals detect register and provenance. Forbidding them made a detector better on both axes at once.
07 / Limits
What this does not prove
The AI and human sides are different corpora, not two halves of one. The AI side is dominated by 4,016 generated long-form articles; the human side by 4,144 modern samples weighted towards business and marketing copy. Register differs between them, and register confounding authorship is this project’s central finding. These likelihood ratios describe how a rule behaves across these corpora. They are strong enough to say a rule points the wrong way. They are not strong enough to fix a weight.
A rule that does not fire is not the same as a rule that is wrong. Of the 116, 95 fired on at least one of 5,743 AI documents. tier3-phrase-cluster is recorded as inactive because it cannot fire on realistic English prose, and twenty more are dormant on every corpus measured. Some of those twenty are deliberate forensic markers kept as insurance, such as leaked citation tokens and unfilled placeholders, and their absence from a prose corpus is the expected result rather than a failure.
One proposed loosening was measured and rejected. The proposed gate for punchline-fragment-density reaches zero documents on the long-form corpus: the AI corpus maximum punchline rate is 0.136 and its 99th percentile 0.084, against a shipped gate of 0.18 and a proposed gate of 0.10. It is recorded as rejected so that it cannot be mistaken for a pending improvement.
These are English rules measured on English prose. Nothing here is measured in any other language.
The 45.1% / 24.8% rule-tier pair has no recorded operating point, and should not be moved into a headline position until one exists. The 90.3% / 1.34% classifier figure the original comparison was made against is superseded: browser measured, pre-segmentation, at a threshold that no longer runs, with no raw counts. It is named on this page only to say that it is not the current figure.
268 of its 922 AI documents sit in a cycle-2 split, 168 of them in the train split; on the human side 11 of 4,636. That overlap is immaterial to the rule firing rates, which involve no model and no training, and it is not immaterial to the classifier row of Figure 1.
Provenance
| Figure | File | Section |
|---|---|---|
| Rule pack en-signals:2026.08.6: 116 named rules across 113 weighted categories | en-signals pack manifest | — |
| Likelihood ratios, firing counts, 5,743 AI and 4,353 human denominators | tests/battery/rule-liveness.json | measured_utc 2026-08-30 |
Original 1,727 AI and 169 human measurement, including token-cutoff 0.16 | research/rule-validation/RULE-VALIDATION.md | whole record |
| Rule tier 45.1% / 24.8%, and the 28 August 2026 demotion | docs/MEASURED-FINDINGS.md | §4 |
| Classifier 883/922 (95.8%) at 45/4,636 (0.97%), fp32, 0.9855 / 0.9763 | thresholds.json via published-figures.ts | server_fp32_segmented |
| 24-feature scorecard, both fits, both populations | research/signal-science/SIGNAL-SCIENCE.md | §3 |
| 268/922 and 168 train-split overlap; human 11/4,636 | corpus-reconciliation-2026-08-29/analysis.txt | §2 |
One date is unreconciled at publication: docs/MEASURED-FINDINGS.md §4 dates the re-measurement 29 August 2026, while tests/battery/rule-liveness.json records measured_utc 2026-08-30. This page uses the measurement file’s own date, 30 August 2026.
Every file named above ships with the open measurement repository, so any figure on this page can be re-run rather than taken on trust.