calibration receipts rendered from committed export data (gaige-receipt-export/1). every number below carries the instrument fingerprint that produced it and a reproduce command a stranger can run. measurements, never verdicts.
3 receipts.
instrument: model tiiuae/falcon-7b · quant 4bit · device cuda · Linux/x86_64 · receipt 20260722-163959-fast-detect-gpt (2026-07-22T21:40:31+00:00, gaige 0.0.2) · corpus sha256 7d2819d3e83bd10d…
AUROC 0.9720 [0.9458–0.9938] (bootstrap, n_boot=1000)
| target FPR | threshold | achieved FPR | TPR | TPR 95% CI |
|---|---|---|---|---|
| 1.0% | 2.1229 | 1.0% | 86.0% | [79.0%–92.0%] |
| 5.0% | 1.8319 | 5.0% | 91.0% | [85.0%–96.0%] |
conformal (distribution-free FPR bound):
| alpha | threshold | TPR | guarantee |
|---|---|---|---|
| 0.05 | 1.8468 | 90.0% | P(human flagged) <= 0.05, marginal over calibration draws, under exchangeability (split conformal) |
| 0.01 | 2.4446 | 76.0% | P(human flagged) <= 0.01, marginal over calibration draws, under exchangeability (split conformal) |
| 0.005 | refused: alpha=0.005 needs >= 199 human calibration samples, got 100. A tighter guarantee than your data supports is not a guarantee. | ||
at target FPR 1.0% (threshold 2.1229):
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 42 | withheld: n below the honesty floor (floor 20/stratum) | 7 | withheld | ||
| 100-250w | 39 | 0.0% | [0.0%–0.0%] | 81 | 85.2% | [77.7%–92.6%] |
| 250-500w | 19 | withheld: n below the honesty floor (floor 20/stratum) | 12 | withheld | ||
max FPR disparity across length_bucket: 0.0% (all reported groups equal)
at target FPR 5.0% (threshold 1.8319):
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 42 | withheld: n below the honesty floor (floor 20/stratum) | 7 | withheld | ||
| 100-250w | 39 | 2.6% | [0.0%–7.7%] | 81 | 88.9% | [81.5%–95.1%] |
| 250-500w | 19 | withheld: n below the honesty floor (floor 20/stratum) | 12 | withheld | ||
max FPR disparity across length_bucket: 2.2% (worst 0-100w 4.8% vs best 100-250w 2.6%)
reproduce:
gaige run --corpus hc3-mini --n 100 --seed 17 --detector fast-detect-gpt --model tiiuae/falcon-7b --quant 4bit --device cuda --max-tokens 1024
instrument: model tiiuae/falcon-7b + tiiuae/falcon-7b-instruct · quant 4bit · device cuda · Linux/x86_64 · receipt 20260722-213952-binoculars (2026-07-23T02:40:53+00:00, gaige 0.0.2) · corpus sha256 7d2819d3e83bd10d…
AUROC 0.9992 [0.9974–1.0000] (bootstrap, n_boot=1000)
| target FPR | threshold | achieved FPR | TPR | TPR 95% CI |
|---|---|---|---|---|
| 1.0% | -0.7829 | 1.0% | 97.0% | [94.0%–100.0%] |
| 5.0% | -0.8706 | 5.0% | 100.0% | [100.0%–100.0%] |
conformal (distribution-free FPR bound):
| alpha | threshold | TPR | guarantee |
|---|---|---|---|
| 0.05 | -0.8706 | 100.0% | P(human flagged) <= 0.05, marginal over calibration draws, under exchangeability (split conformal) |
| 0.01 | -0.7540 | 95.0% | P(human flagged) <= 0.01, marginal over calibration draws, under exchangeability (split conformal) |
| 0.005 | refused: alpha=0.005 needs >= 199 human calibration samples, got 100. A tighter guarantee than your data supports is not a guarantee. | ||
at target FPR 1.0% (threshold -0.7829):
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 42 | withheld: n below the honesty floor (floor 20/stratum) | 7 | withheld | ||
| 100-250w | 39 | 0.0% | [0.0%–0.0%] | 81 | 96.3% | [92.6%–100.0%] |
| 250-500w | 19 | withheld: n below the honesty floor (floor 20/stratum) | 12 | withheld | ||
max FPR disparity across length_bucket: 2.4% (worst 0-100w 2.4% vs best 100-250w 0.0%)
at target FPR 5.0% (threshold -0.8706):
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 42 | withheld: n below the honesty floor (floor 20/stratum) | 7 | withheld | ||
| 100-250w | 39 | 2.6% | [0.0%–7.7%] | 81 | 100.0% | [100.0%–100.0%] |
| 250-500w | 19 | withheld: n below the honesty floor (floor 20/stratum) | 12 | withheld | ||
max FPR disparity across length_bucket: 7.0% (worst 0-100w 9.5% vs best 100-250w 2.6%)
reproduce:
gaige run --corpus hc3-mini --n 100 --seed 17 --detector binoculars --observer tiiuae/falcon-7b --performer tiiuae/falcon-7b-instruct --quant 4bit --device cuda --max-tokens 1024
instrument: model EleutherAI/gpt-neo-1.3B · quant fp32 · device cpu · Linux/x86_64 · receipt 20260727-193225-fast-detect-gpt (2026-07-28T00:44:59+00:00, gaige 0.0.2) · corpus sha256 54b786694f2e4c0d…
AUROC 0.9330 [0.9117–0.9521] (bootstrap, n_boot=1000)
| target FPR | threshold | achieved FPR | TPR | TPR 95% CI |
|---|---|---|---|---|
| 1.0% | 2.3904 | 0.8% | 58.1% | [54.2%–62.7%] |
| 5.0% | 1.4957 | 5.0% | 77.7% | [74.2%–81.5%] |
conformal (distribution-free FPR bound):
| alpha | threshold | TPR | guarantee |
|---|---|---|---|
| 0.05 | 1.5727 | 76.5% | P(human flagged) <= 0.05, marginal over calibration draws, under exchangeability (split conformal) |
| 0.01 | 2.7355 | 49.4% | P(human flagged) <= 0.01, marginal over calibration draws, under exchangeability (split conformal) |
| 0.005 | refused: alpha=0.005 needs >= 199 human calibration samples, got 120. A tighter guarantee than your data supports is not a guarantee. | ||
at target FPR 1.0% (threshold 2.3904):
axis attack
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| none | 120 | 0.8% | [0.0%–2.5%] | 240 | 59.2% | [52.9%–65.0%] |
| paraphrase | 0 | withheld: n below the honesty floor (floor 20/stratum) | 240 | withheld | ||
axis decoding
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| greedy | 0 | withheld: n below the honesty floor (floor 20/stratum) | 218 | withheld | ||
| none | 120 | withheld: n below the honesty floor (floor 20/stratum) | 0 | withheld | ||
| sampling | 0 | withheld: n below the honesty floor (floor 20/stratum) | 262 | withheld | ||
axis domain
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| abstracts | 60 | 1.7% | [0.0%–5.0%] | 240 | 53.3% | [46.7%–59.6%] |
| 60 | 0.0% | [0.0%–0.0%] | 240 | 62.9% | [57.1%–68.3%] |
axis generator
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| gpt4 | 0 | withheld: n below the honesty floor (floor 20/stratum) | 240 | withheld | ||
| human | 120 | withheld: n below the honesty floor (floor 20/stratum) | 0 | withheld | ||
| mistral-chat | 0 | withheld: n below the honesty floor (floor 20/stratum) | 240 | withheld | ||
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 2 | withheld: n below the honesty floor (floor 20/stratum) | 83 | withheld | ||
| 100-250w | 110 | 0.9% | [0.0%–2.7%] | 330 | 61.8% | [56.4%–67.0%] |
| 250-500w | 8 | withheld: n below the honesty floor (floor 20/stratum) | 67 | withheld | ||
max FPR disparity across domain: 1.7% (worst abstracts 1.7% vs best reddit 0.0%)
at target FPR 5.0% (threshold 1.4957):
axis attack
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| none | 120 | 5.0% | [1.7%–9.2%] | 240 | 69.6% | [63.7%–75.4%] |
| paraphrase | 0 | withheld: n below the honesty floor (floor 20/stratum) | 240 | withheld | ||
axis decoding
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| greedy | 0 | withheld: n below the honesty floor (floor 20/stratum) | 218 | withheld | ||
| none | 120 | withheld: n below the honesty floor (floor 20/stratum) | 0 | withheld | ||
| sampling | 0 | withheld: n below the honesty floor (floor 20/stratum) | 262 | withheld | ||
axis domain
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| abstracts | 60 | 5.0% | [0.0%–11.7%] | 240 | 74.2% | [68.8%–79.2%] |
| 60 | 5.0% | [0.0%–11.7%] | 240 | 81.2% | [76.7%–86.2%] |
axis generator
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| gpt4 | 0 | withheld: n below the honesty floor (floor 20/stratum) | 240 | withheld | ||
| human | 120 | withheld: n below the honesty floor (floor 20/stratum) | 0 | withheld | ||
| mistral-chat | 0 | withheld: n below the honesty floor (floor 20/stratum) | 240 | withheld | ||
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 2 | withheld: n below the honesty floor (floor 20/stratum) | 83 | withheld | ||
| 100-250w | 110 | 4.5% | [0.9%–8.2%] | 330 | 79.7% | [75.5%–83.9%] |
| 250-500w | 8 | withheld: n below the honesty floor (floor 20/stratum) | 67 | withheld | ||
max FPR disparity across domain: 0.0% (all reported groups equal)
reproduce:
gaige run --corpus corpora/raid-g2d2a2-n60-s17.jsonl --n 100 --seed 17 --detector fast-detect-gpt --model EleutherAI/gpt-neo-1.3B --quant fp32 --device cpu --max-tokens 1024
This receipt was produced from a locally prepared corpus file. The corpus sha256 above pins the exact bytes; the file itself is not redistributed. RAID slices (Dugan et al., ACL 2024) are prepared from the public dataset with `gaige corpus prepare-raid` and are never redistributed.
rendered from the committed export index; the page recomputes nothing and is rebuilt whole on every refresh.