Get in Touch

Rhythm study · 5,558 documents

The rhythm you can hear

The rhythm scorer fired 3.5 times more often on model misses, but did not justify changing a score.

On this page
  1. The finding
  2. The ear, made checkable
  3. It fires where the model fails
  4. Why it still cannot change the score
  5. The lead that refuted itself
  6. What ships instead
  7. What this does not prove
Published
31 August 2026
Measured
29 August 2026, re-verified at source 31 August 2026
Corpus
5,558 long-form documents: 922 AI, 4,636 human; plus the founder's own nine articles
Operating point
Conditional analysis at the retired 0.984 single threshold, segments-v2 — labelled throughout; nothing here describes the shipped pair

01 / The finding

The ear is measurable. The arithmetic still says no.

Most detection research starts from a corpus. This measurement started from a person: the founder kept reading passages the model had cleared and saying they were machine-written, and describing why in terms no rule in the pack measured — a claim, an instruction, then a consequence, forty words, no mess. The question was whether that rhythm exists outside his head.

The three rhythm signals, on the model's misses3.56×

They fire on 20 of the 45 AI documents the model missed, against 571 of the 4,580 human documents it correctly cleared — 44.4% against 12.5%, interval 2.40–4.75, which excludes 1. The previous rhythm rules scored 1.2 on the same test and their interval did not.

His own nine articles, strongest paragraph score8 v 3

The humanised AI article the model scores 80.8% reaches 8 at a gate of 4; the unedited AI reaches 7; his six genuine human articles reach 3, 1, 0, 0, 0 and 0. The module finds both passages he quoted, verbatim, with the role sequences he described.

What escalating on it would cost5 for 5

The only defensible operating point gains five AI documents and wrongly flags five human ones. Held out: +0.562 points of detection for +0.106 points of false positives — beaten on both axes by a candidate already measured and not yet spent.

The reason this does not become a shipping decision is arithmetic, not doubt about the phenomenon.

SYNTHETIC-CADENCE.md, the measurement's own closing line. The rhythm now powers a visible evidence tell instead — an explanation, never a verdict.

02 / The instrument

An ear, rewritten as rules a reader can check

No language model classifies the sentences. Each sentence gets one of eight approximate roles — claim, contrast, example, instruction, recommendation, consequence, question, aside — assigned from word lists, sentence-initial shape, mood and punctuation, in fixed precedence. The approximation is deliberately crude: a rule the reader can check against a paragraph by eye is worth more here than an accurate one nobody can audit.

On the two passages the founder picked out of his own humanised article, the roles land exactly where he said they would. Both are 41 words. Both score 8 at a gate of 4.

Claim. A business owner may build a website to explain an offer, collect enquiries, manage appointments or sell online. Instruction. Write the priority customer tasks into the brief before asking for web design prices. Consequence. Otherwise, each web designer may price a different solution.

Compressed triple, micro-paragraph, 41 words, roles C I S — quoted from the article the model scores 80.8%, below its flag point.

And the false positive this family of rules would produce is his own writing. His closest human paragraph — claim, consequence, instruction, from an article about the Panda update — scores 3. A single parenthetical aside is the only thing keeping it under the gate. That is worth publishing precisely because it is the shape of the mistake.

Figure 1 The founder's nine documents, strongest paragraph score One paragraph-cadence score per document: the highest-scoring paragraph, at a gate of 4. The two AI documents the model reads weakest score 7 and 8; all six genuine human articles stay under the gate; the hand-rewritten AI piece scores alongside the human writing, which matches his description of it as substantially rewritten.
Strongest paragraph-cadence score for each of the founder's nine documents, against the gate of 4 0 2 4 6 8 Humanised AI article (scored 80.8%) AI, the target case 8 Unedited AI article AI 7 Hand-rewritten AI article AI, substantially rewritten 3 Panda-penguin article human 3 Mobile algorithm article human 1 Social objectives article human 0 Esports article human 0 Facebook article human 0 EU ranking article human 0 the gate, 4
Strongest paragraph-cadence score per document
DocumentWhat it isStrongest paragraph
Humanised AI article (scored 80.8%)AI, the target case8
Unedited AI articleAI7
Hand-rewritten AI articleAI, substantially rewritten3
Panda-penguin articlehuman3
Mobile algorithm articlehuman1
Social objectives articlehuman0
Esports articlehuman0
Facebook articlehuman0
EU ranking articlehuman0

SYNTHETIC-CADENCE.md §7. Model probabilities re-derived through the fp32 harness, model SHA e313ab00de1fffd2…; the smallest denominator in this paper, and labelled as such.

03 / The conditional result

It fires where the model fails — the first signal in this project that does

The only question that matters for a candidate signal is not whether it separates AI from human — the model already does that at AUROC 0.997 on this corpus — but whether it fires on the documents the model gets wrong. Among the 45 AI documents the model missed at the retired 0.984 threshold, against the 4,580 human documents it correctly cleared:

Measured at a rule that no longer ships

Corpus
5,558 long-form documents; conditional population 45 misses against 4,580 correctly cleared humans. Bootstrap intervals from 2,000 resamples, because n = 45.
Retired flag point
0.984 single threshold, segments-v2
Detector
tier3-cycle5-full-e5small-fp32.onnx, SHA-256 45e00978b10d1df6…
Runtime
Feature pass is pure local computation — no model, no network, 26 seconds over the corpus. The document scores it conditions on are the fp32 cycle-2 checkpoint's, reproduced (877/922, 56/4,636) before anything was computed.
Measured
30 August 2026
Also
Signal gates are the 95th percentile of the human distribution, in sample, so every human fire rate is about 5% by construction and the likelihood ratio reads directly.

The rule that ships today is margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562). The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.

Figure 2 Fire rate on the model's misses, as a multiple of the rate on cleared humans Log-scale reading not needed: bars are the ratio itself. Requiring two of the three signals reaches 15.4× on a wide interval whose lower bound is 3.51. The seven previously measured rhythm rules — the earlier attempt at exactly this — sit at 1.2 with an interval that includes 1, which is why this is a new result and not a repeat.
Conditional likelihood ratios of the cadence signals on the model's misses 12× 16× 95% interval Two of the three signals 5/45 misses · 33/4,580 cleared humans 15.42× 3.51–33.93 Compressed sentence triples 11/45 · 213/4,580 5.26× 2.73–8.12 Repeated paragraph shapes 7/45 · 160/4,580 4.45× 1.72–7.97 Any of the three 20/45 · 571/4,580 3.56× 2.40–4.75 Balanced two-part constructions 7/45 · 232/4,580 3.07× 1.22–5.28 The seven previous rhythm rules 2/45 · 168/4,580 — the earlier attempt 1.2× interval includes 1 no information, 1×
Conditional fire rates and ratios
SignalOn the 45 missesOn the 4,580 clearedRatio [95% CI]
Compressed sentence triples11 = 24.4%213 = 4.7%5.26 [2.73, 8.12]
Repeated paragraph shapes7 = 15.6%160 = 3.5%4.45 [1.72, 7.97]
Balanced constructions7 = 15.6%232 = 5.1%3.07 [1.22, 5.28]
Any of the three20 = 44.4%571 = 12.5%3.56 [2.40, 4.75]
Two of the three5 = 11.1%33 = 0.7%15.42 [3.51, 33.93]
The seven previous rhythm rules2 = 4.4%168 = 3.7%1.2 — interval includes 1

SYNTHETIC-CADENCE.md §4. The 95% intervals are printed beside each bar; the two-of-three interval is wide because five documents of forty-five fired it.

One caution the headline hides. The two strongest signals only work on long documents: compressed triples read at chance below 1,200 words (AUROC 0.460 under 800 words against 0.703 above 1,600). The founder's own humanised article is 512 words — in the band where the corpus rule fails — which is exactly why the corpus rule misses his target case even though the paragraph rule reads it perfectly.

04 / The refusal

Five catches for five false accusations — the price of letting the ear vote

An escalation rule was priced rather than assumed: flag when the model's probability sits in a band below the threshold and the cadence rule fires. Every setting was costed against the retired 0.984 threshold's 877/922 detection and 56/4,636 false positives — not the shipped pair, and not the current cycle-5 rule.

Escalation settings and their measured costs
RuleBandAI gainedHumans wrongly flagged
Any one of the three signals0.80–0.984+20+117
Any one of the three0.95–0.984+18+34
Two of the three0.80–0.984+5+5
All threeany00

The two-of-three rule is the only line worth reading twice, and cross-validated over 200 split-halves it gains +0.562 points of detection for +0.106 points of human false positives. That survives out of sample — and it is still refused, for three reasons stated in the measurement itself:

It is beaten. A different candidate, already measured and needing no new tier and no policy change, buys +0.82 points for +0.077. Spending the false-positive budget on the worse deal because its story is better would be exactly backwards.

The looser settings are forbidden by name. The single-signal rule at band 0.95 takes human fiction from 11.15% wrongly flagged to 15.38%, and fiction is a register this programme's objectives protect explicitly.

And the worst register cannot be tested at all. 0 of 260 human fiction documents in the corpus have paragraph breaks — the scrapers stripped them — so every paragraph-shape signal has no value at all on the register with the worst false-positive rate. Any claim that this rule leaves fiction alone would be untested, not verified. A rule that cannot be tested where it is most dangerous does not get to vote.

05 / A hypothesis refuted

The most promising lead measured nothing, and that is published too

The strongest a-priori bet was the discourse markers: the founder's house prompt bans connectives like "furthermore", "moreover" and "in conclusion", so their absence should fingerprint anti-AI-instructed text. Measured: 44.25% of AI documents against 44.44% of human ones fire the low-marker gate — likelihood ratio 1.00. The instruction does change the text (median density falls from 0.835 per 1,000 words on plain prompts to 0.411 under the house brief), but human published prose sits at 0.742, between the two, so the measure cannot separate either group.

Conditionally it is worse than useless: on the documents the model misses, the wider marker inventory reads AUROC 0.379 — reversed. The AI the model fails on uses more of these connectives than the human text it clears. Four other candidate signals also point the wrong way, including paragraph-opening diversity: human documents in this corpus repeat their paragraph openings twelve times more often than AI documents do. "Every paragraph starts the same way" is a human tell here, not a machine one.

06 / What ships

An explanation, not a verdict: the compressed-rhythm evidence tell

The measurement's own recommendation was followed to the letter: not detection, possibly editorial. Since 31 August 2026 the checker's "Why it reads this way" card quotes the single highest-scoring paragraph of an AI-leaning draft — "this paragraph reads as a claim, an instruction, then a consequence — 41 words with the conversational mess removed" — with the measured rates as fine print: this compressed cadence appears in 28.9% of AI paragraph-bearing documents (266 of 921) and 14.0% of human ones (483 of 3,451).

That 14% is why the copy reads "over-structured" and never "AI". One human document in seven carries the rhythm; the founder's own writing carries it at score 3. The tell explains the model's reading. It has no vote in it, and the separation is enforced in code, not in intentions.

The scorer shipped with its probe: the founder's quoted passages must produce the role sequences he described and clear their floors, no paragraph of his six human articles may reach the gate, and four mutation tests break the detectors one by one and assert the probe then fails — because this project once shipped a dead control that passed its tests, and a test that cannot fail is not a test.

07 / Limits

What this page does not prove

The fiction gap is the disqualifying one.

Every paragraph-shape figure on this page is measured on a corpus whose human fiction — the tool's worst false-positive register — has no paragraph breaks at all (0 of 260). The signals' behaviour there is unknown, not clean.

The conditional population is 45 documents. Every ratio in Figure 2 carries a bootstrap interval for that reason, and the strongest (15.4×) rests on five firings. The founder's nine documents are the most encouraging result here and the smallest denominator in the paper.

The gates are in-sample except where the split-half cross-validation is quoted, the operating point is the retired 0.984 single threshold rather than the shipped pair, and the role approximation has never been checked against a human annotator — it agrees with the founder's reading of three passages, which is not an evaluation.

One true target case exists. The corpus's anti-AI-instructed analogues are milder than the founder's own instructions; his 512-word humanised article is the only genuine example, and n = 1 is a case study, not a rate. What would change the answer is a corpus of short, anti-AI-instructed articles with known provenance — which is now being gathered.

Apply the method

Run a draft and see the rhythm quoted back, never scored.