01 / The finding
The default sits where the trade turns, and the bottom of the dial is pure cost
Every notch below was priced against tier3-cycle2-v1, retired 1 September 2026. tier3-cycle5-v1 now ships at a margin-space rule rather than a probability pair, and this eighteen-notch curve has not been re-run against it — the shape of the trade-off is the finding, not the specific pair marked "shipped".
0.9855 → 0.9820 buys 0.9 points of detection and takes wrongly-flagged human documents from 45 to 91 of 4,636 — 0.97% to 1.96%. The recurring complaint about a 0.98 document that did not flag is answered by this row: catching it costs a doubling.
At 0.9950 the tool wrongly flags 13 human documents in 4,636 — and misses one AI document in seven (85.1%). Tightening is cheap in accusations and expensive in catches.
Detection is already 99.5% and only the false positives climb — to more than a fifth of all human documents at 0.80. A control offering those settings should say so, not present them as a neutral preference.
02 / The curve
Eighteen operating points, full precision, both denominators on every row
Measured at a rule that no longer ships
- Corpus
- 922 AI and 4,636 human long-form documents, the whole corpus at every notch.
- Retired flag point
- 0.9855 / 0.9763
- Detector
- tier3-cycle2-e5small-fp32.onnx, SHA-256 e313ab00de1fffd2…4d2788d (cycle 2, superseded 1 September 2026)
- Runtime
- tier3-cycle2-e5small-fp32.onnx, SHA e313ab00de1fffd2…, fp32 server path, segments-v3, full precision — the 4 dp segment store rounds 884/922 where the truth is 883/922, and a curve built on it would be wrong by a document at every notch.
- Measured
- 30 August 2026
- Also
- Each notch keeps the shipped 0.0092 gap between the two arms of the minimum-evidence rule, so every row is the same rule at a different strictness rather than a different rule. The harness reproduced the published 883/922 and 45/4,636 before any new notch was taken.
The rule that ships today is margin 3.570935 / gap 0.34 (display 0.9679 / 0.9562). The figures under this stamp answer the same question at a flag point this tool no longer uses, and they are not a description of the tool as it runs now.
| Primary / secondary | AI detected | Human false positives |
|---|---|---|
| 0.9985 / 0.9893 | 562/922 (61.0%) | 3/4,636 (0.06%) |
| 0.9970 / 0.9878 | 718/922 (77.9%) | 6/4,636 (0.13%) |
| 0.9950 / 0.9858 | 785/922 (85.1%) | 13/4,636 (0.28%) |
| 0.9920 / 0.9828 | 828/922 (89.8%) | 20/4,636 (0.43%) |
| 0.9895 / 0.9803 | 851/922 (92.3%) | 26/4,636 (0.56%) |
| 0.9855 / 0.9763 — tier3-cycle2's shipped default (retired 1 September 2026) | 883/922 (95.8%) | 45/4,636 (0.97%) |
| 0.9820 / 0.9728 | 892/922 (96.7%) | 91/4,636 (1.96%) |
| 0.9780 / 0.9688 | 900/922 (97.6%) | 126/4,636 (2.72%) |
| 0.9730 / 0.9638 | 912/922 (98.9%) | 177/4,636 (3.82%) |
| 0.9670 / 0.9578 | 915/922 (99.2%) | 225/4,636 (4.85%) |
| 0.9600 / 0.9508 | 917/922 (99.5%) | 288/4,636 (6.21%) |
| 0.9500 / 0.9408 | 918/922 (99.6%) | 358/4,636 (7.72%) |
| 0.9400 / 0.9308 | 918/922 (99.6%) | 425/4,636 (9.17%) |
| 0.9200 / 0.9108 | 919/922 (99.7%) | 554/4,636 (11.95%) |
| 0.9000 / 0.8908 | 919/922 (99.7%) | 661/4,636 (14.26%) |
| 0.8700 / 0.8608 | 921/922 (99.9%) | 798/4,636 (17.21%) |
| 0.8400 / 0.8308 | 922/922 (100.0%) | 907/4,636 (19.56%) |
| 0.8000 / 0.7908 | 922/922 (100.0%) | 1,041/4,636 (22.45%) |
03 / Reading it
Three sentences the table compresses
Loosening from the default to 0.9730 buys 3.1 points of detection and quadruples the wrongly-flagged humans to 3.82%. Tightening to 0.9950 spares 32 human documents and abandons 98 machine-written ones. And the cost is not evenly spread across kinds of writing: human fiction carries a disproportionate share of the false positives at every notch, so a dial would punish novelists first — the per-register breakdown ships with the raw output.
04 / If it ships
The design rules a control would have to obey
Nothing here is a commitment to build the dial. If one ships: every notch prints its own measured false-positive rate with denominators; the default is the measured shipped point and is labelled as such; the bottom of the range carries a warning in words ("about one human document in five would be flagged at this setting"), not just a number; the browser runtime gets its own separately measured curve, because the two runtimes already need separately fitted flag points; and none of these values is hard-coded — they are data, re-measured whenever the model or operating point moves.