Get in Touch

Threshold study · 18 settings

The price of strictness

Eighteen operating points show the missed machines and wrongly flagged people behind each strictness setting.

On this page
  1. The finding
  2. The eighteen notches
  3. How to read it
  4. If a control ever ships
Published
31 August 2026
Measured
30 August 2026, full-precision re-score of the whole corpus at eighteen operating points
Corpus
922 AI and 4,636 human long-form documents
Superseded 1 September 2026
Every notch below, including the default marked "shipped", was priced against tier3-cycle2-v1, retired the day tier3-cycle5-v1 replaced it with a margin-space rule. This curve has not been re-run at the pair that ships today.

01 / The finding

The default sits where the trade turns, and the bottom of the dial is pure cost

Every notch below was priced against tier3-cycle2-v1, retired 1 September 2026. tier3-cycle5-v1 now ships at a margin-space rule rather than a probability pair, and this eighteen-notch curve has not been re-run against it — the shape of the trade-off is the finding, not the specific pair marked "shipped".

One notch looser than the default2× the false positives

0.9855 → 0.9820 buys 0.9 points of detection and takes wrongly-flagged human documents from 45 to 91 of 4,636 — 0.97% to 1.96%. The recurring complaint about a 0.98 document that did not flag is answered by this row: catching it costs a doubling.

Detection left on the table by tightening1 in 7

At 0.9950 the tool wrongly flags 13 human documents in 4,636 — and misses one AI document in seven (85.1%). Tightening is cheap in accusations and expensive in catches.

Below 0.96Pure cost

Detection is already 99.5% and only the false positives climb — to more than a fifth of all human documents at 0.80. A control offering those settings should say so, not present them as a neutral preference.

02 / The curve

Eighteen operating points, full precision, both denominators on every row

Measured at a rule that no longer ships

Corpus
922 AI and 4,636 human long-form documents, the whole corpus at every notch.
Retired flag point
0.9855 / 0.9763
Detector
tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d (cycle 2, superseded 1 September 2026)
Runtime
tier3-cycle2-e5small-fp32.onnx, SHA e313ab00de1fffd2…, fp32 server path, segments-v3, full precision — the 4 dp segment store rounds 884/922 where the truth is 883/922, and a curve built on it would be wrong by a document at every notch.
Measured
30 August 2026
Also
Each notch keeps the shipped 0.0092 gap between the two arms of the minimum-evidence rule, so every row is the same rule at a different strictness rather than a different rule. The harness reproduced the published 883/922 and 45/4,636 before any new notch was taken.

The rule that ships today is margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562). The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.

Detection and false positives at eighteen operating points
Primary / secondaryAI detectedHuman false positives
0.9985 / 0.9893562/922 (61.0%)3/4,636 (0.06%)
0.9970 / 0.9878718/922 (77.9%)6/4,636 (0.13%)
0.9950 / 0.9858785/922 (85.1%)13/4,636 (0.28%)
0.9920 / 0.9828828/922 (89.8%)20/4,636 (0.43%)
0.9895 / 0.9803851/922 (92.3%)26/4,636 (0.56%)
0.9855 / 0.9763 — tier3-cycle2's shipped default (retired 1 September 2026)883/922 (95.8%)45/4,636 (0.97%)
0.9820 / 0.9728892/922 (96.7%)91/4,636 (1.96%)
0.9780 / 0.9688900/922 (97.6%)126/4,636 (2.72%)
0.9730 / 0.9638912/922 (98.9%)177/4,636 (3.82%)
0.9670 / 0.9578915/922 (99.2%)225/4,636 (4.85%)
0.9600 / 0.9508917/922 (99.5%)288/4,636 (6.21%)
0.9500 / 0.9408918/922 (99.6%)358/4,636 (7.72%)
0.9400 / 0.9308918/922 (99.6%)425/4,636 (9.17%)
0.9200 / 0.9108919/922 (99.7%)554/4,636 (11.95%)
0.9000 / 0.8908919/922 (99.7%)661/4,636 (14.26%)
0.8700 / 0.8608921/922 (99.9%)798/4,636 (17.21%)
0.8400 / 0.8308922/922 (100.0%)907/4,636 (19.56%)
0.8000 / 0.7908922/922 (100.0%)1,041/4,636 (22.45%)

03 / Reading it

Three sentences the table compresses

Loosening from the default to 0.9730 buys 3.1 points of detection and quadruples the wrongly-flagged humans to 3.82%. Tightening to 0.9950 spares 32 human documents and abandons 98 machine-written ones. And the cost is not evenly spread across kinds of writing: human fiction carries a disproportionate share of the false positives at every notch, so a dial would punish novelists first — the per-register breakdown ships with the raw output.

04 / If it ships

The design rules a control would have to obey

Nothing here is a commitment to build the dial. If one ships: every notch prints its own measured false-positive rate with denominators; the default is the measured shipped point and is labelled as such; the bottom of the range carries a warning in words ("about one human document in five would be flagged at this setting"), not just a number; the browser runtime gets its own separately measured curve, because the two runtimes already need separately fitted flag points; and none of these values is hard-coded — they are data, re-measured whenever the model or operating point moves.

Apply the method

Check a document at the published flag point.