calibration receipts rendered from committed export data (gaige-receipt-export/1). every number below carries the instrument fingerprint that produced it and a reproduce command a stranger can run. measurements, never verdicts.
5 receipts.
instrument: model tiiuae/falcon-7b · quant 4bit · device cuda · Linux/x86_64 · receipt 20260722-163959-fast-detect-gpt (2026-07-22T21:40:31+00:00, gaige 0.0.1 · export gaige 0.0.4) · corpus sha256 7d2819d3e83bd10d…
AUROC 0.9720 [0.9458–0.9938] (bootstrap, n_boot=1000)
| target FPR | threshold | achieved FPR | TPR | TPR 95% CI |
|---|---|---|---|---|
| 1.0% | 2.1229 | 1.0% | 86.0% | [79.0%–92.0%] |
| 5.0% | 1.8319 | 5.0% | 91.0% | [85.0%–96.0%] |
conformal (distribution-free FPR bound):
| alpha | threshold | TPR | guarantee |
|---|---|---|---|
| 0.05 | 1.8468 | 90.0% | P(human flagged) <= 0.05, marginal over calibration draws, under exchangeability (split conformal) |
| 0.01 | 2.4446 | 76.0% | P(human flagged) <= 0.01, marginal over calibration draws, under exchangeability (split conformal) |
| 0.005 | refused: alpha=0.005 needs >= 199 human calibration samples, got 100. A tighter guarantee than your data supports is not a guarantee. | ||
at target FPR 1.0% (threshold 2.1229):
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 42 | 0.0% | [0.0%–0.0%] | 7 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| 100-250w | 39 | 0.0% | [0.0%–0.0%] | 81 | 85.2% | [77.7%–92.6%] |
| 250-500w | 19 | withheld: n human below the honesty floor (floor 20/stratum) | 12 | withheld: n AI below the honesty floor (floor 20/stratum) | ||
max FPR disparity across length_bucket: 0.0% (all reported groups equal)
at target FPR 5.0% (threshold 1.8319):
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 42 | 4.8% | [0.0%–11.9%] | 7 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| 100-250w | 39 | 2.6% | [0.0%–7.7%] | 81 | 88.9% | [81.5%–95.1%] |
| 250-500w | 19 | withheld: n human below the honesty floor (floor 20/stratum) | 12 | withheld: n AI below the honesty floor (floor 20/stratum) | ||
max FPR disparity across length_bucket: 2.2% (worst 0-100w 4.8% vs best 100-250w 2.6%)
reproduce:
gaige run --corpus hc3-mini --n 100 --seed 17 --detector fast-detect-gpt --model tiiuae/falcon-7b --quant 4bit --device cuda --max-tokens 1024
instrument: model tiiuae/falcon-7b + tiiuae/falcon-7b-instruct · quant 4bit · device cuda · Linux/x86_64 · receipt 20260722-213952-binoculars (2026-07-23T02:40:53+00:00, gaige 0.0.1 · export gaige 0.0.4) · corpus sha256 7d2819d3e83bd10d…
AUROC 0.9992 [0.9974–1.0000] (bootstrap, n_boot=1000)
| target FPR | threshold | achieved FPR | TPR | TPR 95% CI |
|---|---|---|---|---|
| 1.0% | -0.7829 | 1.0% | 97.0% | [94.0%–100.0%] |
| 5.0% | -0.8706 | 5.0% | 100.0% | [100.0%–100.0%] |
conformal (distribution-free FPR bound):
| alpha | threshold | TPR | guarantee |
|---|---|---|---|
| 0.05 | -0.8706 | 100.0% | P(human flagged) <= 0.05, marginal over calibration draws, under exchangeability (split conformal) |
| 0.01 | -0.7540 | 95.0% | P(human flagged) <= 0.01, marginal over calibration draws, under exchangeability (split conformal) |
| 0.005 | refused: alpha=0.005 needs >= 199 human calibration samples, got 100. A tighter guarantee than your data supports is not a guarantee. | ||
at target FPR 1.0% (threshold -0.7829):
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 42 | 2.4% | [0.0%–7.1%] | 7 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| 100-250w | 39 | 0.0% | [0.0%–0.0%] | 81 | 96.3% | [92.6%–100.0%] |
| 250-500w | 19 | withheld: n human below the honesty floor (floor 20/stratum) | 12 | withheld: n AI below the honesty floor (floor 20/stratum) | ||
max FPR disparity across length_bucket: 2.4% (worst 0-100w 2.4% vs best 100-250w 0.0%)
at target FPR 5.0% (threshold -0.8706):
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 42 | 9.5% | [2.4%–19.0%] | 7 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| 100-250w | 39 | 2.6% | [0.0%–7.7%] | 81 | 100.0% | [100.0%–100.0%] |
| 250-500w | 19 | withheld: n human below the honesty floor (floor 20/stratum) | 12 | withheld: n AI below the honesty floor (floor 20/stratum) | ||
max FPR disparity across length_bucket: 7.0% (worst 0-100w 9.5% vs best 100-250w 2.6%)
reproduce:
gaige run --corpus hc3-mini --n 100 --seed 17 --detector binoculars --observer tiiuae/falcon-7b --performer tiiuae/falcon-7b-instruct --quant 4bit --device cuda --max-tokens 1024
instrument: model tiiuae/falcon-7b · quant 4bit · device cuda · Linux/x86_64 · receipt 20260725-035507-fast-detect-gpt (2026-07-25T08:56:30+00:00, gaige 0.0.1 · export gaige 0.0.4) · corpus sha256 54b786694f2e4c0d…
AUROC 0.9285 [0.9061–0.9482] (bootstrap, n_boot=1000)
| target FPR | threshold | achieved FPR | TPR | TPR 95% CI |
|---|---|---|---|---|
| 1.0% | 2.4401 | 0.8% | 61.5% | [57.3%–65.8%] |
| 5.0% | 1.8878 | 5.0% | 73.1% | [69.4%–77.3%] |
conformal (distribution-free FPR bound):
| alpha | threshold | TPR | guarantee |
|---|---|---|---|
| 0.05 | 1.9249 | 72.5% | P(human flagged) <= 0.05, marginal over calibration draws, under exchangeability (split conformal) |
| 0.01 | 2.6423 | 56.2% | P(human flagged) <= 0.01, marginal over calibration draws, under exchangeability (split conformal) |
| 0.005 | refused: alpha=0.005 needs >= 199 human calibration samples, got 120. A tighter guarantee than your data supports is not a guarantee. | ||
at target FPR 1.0% (threshold 2.4401):
axis attack
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| none | 120 | 0.8% | [0.0%–2.5%] | 240 | 59.2% | [52.9%–65.0%] |
| paraphrase | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 63.7% | [57.9%–70.0%] | |
axis decoding
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| greedy | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 218 | 87.6% | [83.0%–91.7%] | |
| none | 120 | 0.8% | [0.0%–2.5%] | 0 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| sampling | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 262 | 39.7% | [34.0%–45.8%] | |
axis domain
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| abstracts | 60 | 1.7% | [0.0%–5.0%] | 240 | 57.9% | [51.2%–64.2%] |
| 60 | 0.0% | [0.0%–0.0%] | 240 | 65.0% | [59.6%–71.2%] |
axis generator
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| gpt4 | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 60.4% | [53.8%–66.2%] | |
| human | 120 | 0.8% | [0.0%–2.5%] | 0 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| mistral-chat | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 62.5% | [56.2%–68.8%] | |
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 2 | withheld: n human below the honesty floor (floor 20/stratum) | 83 | 47.0% | [36.1%–57.8%] | |
| 100-250w | 110 | 0.9% | [0.0%–2.7%] | 330 | 63.9% | [58.5%–68.8%] |
| 250-500w | 8 | withheld: n human below the honesty floor (floor 20/stratum) | 67 | 67.2% | [55.2%–77.6%] | |
max FPR disparity across domain: 1.7% (worst abstracts 1.7% vs best reddit 0.0%)
at target FPR 5.0% (threshold 1.8878):
axis attack
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| none | 120 | 5.0% | [1.7%–9.2%] | 240 | 67.9% | [61.7%–73.8%] |
| paraphrase | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 78.3% | [72.9%–82.9%] | |
axis decoding
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| greedy | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 218 | 94.0% | [90.8%–97.2%] | |
| none | 120 | 5.0% | [1.7%–9.2%] | 0 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| sampling | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 262 | 55.7% | [49.6%–61.5%] | |
axis domain
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| abstracts | 60 | 3.3% | [0.0%–8.3%] | 240 | 68.8% | [62.5%–74.6%] |
| 60 | 6.7% | [1.7%–13.3%] | 240 | 77.5% | [72.5%–82.9%] |
axis generator
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| gpt4 | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 70.8% | [65.0%–76.2%] | |
| human | 120 | 5.0% | [1.7%–9.2%] | 0 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| mistral-chat | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 75.4% | [70.0%–80.8%] | |
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 2 | withheld: n human below the honesty floor (floor 20/stratum) | 83 | 61.4% | [50.6%–71.1%] | |
| 100-250w | 110 | 5.5% | [1.8%–10.0%] | 330 | 75.5% | [71.2%–80.0%] |
| 250-500w | 8 | withheld: n human below the honesty floor (floor 20/stratum) | 67 | 76.1% | [65.7%–86.6%] | |
max FPR disparity across domain: 3.3% (worst reddit 6.7% vs best abstracts 3.3%)
reproduce:
gaige run --corpus corpora/raid-g2d2a2-n60-s17.jsonl --n 100 --seed 17 --detector fast-detect-gpt --model tiiuae/falcon-7b --quant 4bit --device cuda --max-tokens 1024
This receipt was produced from a locally prepared corpus file. The corpus sha256 above pins the exact bytes; the file itself is not redistributed. RAID slices (Dugan et al., ACL 2024) are prepared from the public dataset with `gaige corpus prepare-raid` and are never redistributed.
instrument: model EleutherAI/gpt-neo-1.3B · quant fp32 · device cpu · Linux/x86_64 · receipt 20260727-193225-fast-detect-gpt (2026-07-28T00:44:59+00:00, gaige 0.0.1 · export gaige 0.0.2) · corpus sha256 54b786694f2e4c0d…
AUROC 0.9330 [0.9117–0.9521] (bootstrap, n_boot=1000)
| target FPR | threshold | achieved FPR | TPR | TPR 95% CI |
|---|---|---|---|---|
| 1.0% | 2.3904 | 0.8% | 58.1% | [54.2%–62.7%] |
| 5.0% | 1.4957 | 5.0% | 77.7% | [74.2%–81.5%] |
conformal (distribution-free FPR bound):
| alpha | threshold | TPR | guarantee |
|---|---|---|---|
| 0.05 | 1.5727 | 76.5% | P(human flagged) <= 0.05, marginal over calibration draws, under exchangeability (split conformal) |
| 0.01 | 2.7355 | 49.4% | P(human flagged) <= 0.01, marginal over calibration draws, under exchangeability (split conformal) |
| 0.005 | refused: alpha=0.005 needs >= 199 human calibration samples, got 120. A tighter guarantee than your data supports is not a guarantee. | ||
at target FPR 1.0% (threshold 2.3904):
axis attack
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| none | 120 | 0.8% | [0.0%–2.5%] | 240 | 59.2% | [52.9%–65.0%] |
| paraphrase | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 57.1% | [51.2%–63.3%] | |
axis decoding
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| greedy | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 218 | 83.0% | [78.0%–88.1%] | |
| none | 120 | 0.8% | [0.0%–2.5%] | 0 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| sampling | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 262 | 37.4% | [31.7%–43.9%] | |
axis domain
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| abstracts | 60 | 1.7% | [0.0%–5.0%] | 240 | 53.3% | [46.7%–59.6%] |
| 60 | 0.0% | [0.0%–0.0%] | 240 | 62.9% | [57.1%–68.3%] |
axis generator
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| gpt4 | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 51.2% | [45.0%–57.5%] | |
| human | 120 | 0.8% | [0.0%–2.5%] | 0 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| mistral-chat | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 65.0% | [58.8%–71.3%] | |
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 2 | withheld: n human below the honesty floor (floor 20/stratum) | 83 | 39.8% | [30.1%–50.6%] | |
| 100-250w | 110 | 0.9% | [0.0%–2.7%] | 330 | 61.8% | [56.4%–67.0%] |
| 250-500w | 8 | withheld: n human below the honesty floor (floor 20/stratum) | 67 | 62.7% | [50.7%–74.6%] | |
max FPR disparity across domain: 1.7% (worst abstracts 1.7% vs best reddit 0.0%)
at target FPR 5.0% (threshold 1.4957):
axis attack
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| none | 120 | 5.0% | [1.7%–9.2%] | 240 | 69.6% | [63.7%–75.4%] |
| paraphrase | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 85.8% | [81.2%–90.0%] | |
axis decoding
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| greedy | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 218 | 96.8% | [94.5%–99.1%] | |
| none | 120 | 5.0% | [1.7%–9.2%] | 0 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| sampling | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 262 | 61.8% | [56.1%–67.9%] | |
axis domain
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| abstracts | 60 | 5.0% | [0.0%–11.7%] | 240 | 74.2% | [68.8%–79.2%] |
| 60 | 5.0% | [0.0%–11.7%] | 240 | 81.2% | [76.7%–86.2%] |
axis generator
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| gpt4 | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 73.3% | [67.9%–78.3%] | |
| human | 120 | 5.0% | [1.7%–9.2%] | 0 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| mistral-chat | 0 | withheld: n human below the honesty floor (floor 20/stratum) | 240 | 82.1% | [77.1%–87.1%] | |
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 2 | withheld: n human below the honesty floor (floor 20/stratum) | 83 | 73.5% | [63.9%–81.9%] | |
| 100-250w | 110 | 4.5% | [0.9%–8.2%] | 330 | 79.7% | [75.5%–83.9%] |
| 250-500w | 8 | withheld: n human below the honesty floor (floor 20/stratum) | 67 | 73.1% | [62.7%–83.6%] | |
max FPR disparity across domain: 0.0% (all reported groups equal)
reproduce:
gaige run --corpus corpora/raid-g2d2a2-n60-s17.jsonl --n 100 --seed 17 --detector fast-detect-gpt --model EleutherAI/gpt-neo-1.3B --quant fp32 --device cpu --max-tokens 1024
This receipt was produced from a locally prepared corpus file. The corpus sha256 above pins the exact bytes; the file itself is not redistributed. RAID slices (Dugan et al., ACL 2024) are prepared from the public dataset with `gaige corpus prepare-raid` and are never redistributed.
instrument: model tiiuae/falcon-7b · quant 4bit · device cuda · Linux/x86_64 · receipt 20260810-134151-fast-detect-gpt (2026-08-10T18:42:24+00:00, gaige 0.0.4 · export gaige 0.0.4) · corpus sha256 7d2819d3e83bd10d…
AUROC 0.9720 [0.9448–0.9929] (bootstrap, n_boot=1000)
| target FPR | threshold | achieved FPR | TPR | TPR 95% CI |
|---|---|---|---|---|
| 1.0% | 2.1229 | 1.0% | 86.0% | [79.0%–92.0%] |
| 5.0% | 1.8319 | 5.0% | 91.0% | [85.0%–96.0%] |
conformal (distribution-free FPR bound):
| alpha | threshold | TPR | guarantee |
|---|---|---|---|
| 0.05 | 1.8468 | 90.0% | P(human flagged) <= 0.05, marginal over calibration draws, under exchangeability (split conformal) |
| 0.01 | 2.4446 | 76.0% | P(human flagged) <= 0.01, marginal over calibration draws, under exchangeability (split conformal) |
| 0.005 | refused: alpha=0.005 needs >= 199 human calibration samples, got 100. A tighter guarantee than your data supports is not a guarantee. | ||
at target FPR 1.0% (threshold 2.1229):
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 42 | 0.0% | [0.0%–0.0%] | 7 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| 100-250w | 39 | 0.0% | [0.0%–0.0%] | 81 | 85.2% | [77.7%–92.6%] |
| 250-500w | 19 | withheld: n human below the honesty floor (floor 20/stratum) | 12 | withheld: n AI below the honesty floor (floor 20/stratum) | ||
max FPR disparity across length_bucket: 0.0% (all reported groups equal)
at target FPR 5.0% (threshold 1.8319):
axis length_bucket
| stratum | n human | FPR | FPR CI | n AI | TPR | TPR CI |
|---|---|---|---|---|---|---|
| 0-100w | 42 | 4.8% | [0.0%–11.9%] | 7 | withheld: n AI below the honesty floor (floor 20/stratum) | |
| 100-250w | 39 | 2.6% | [0.0%–7.7%] | 81 | 88.9% | [81.5%–95.1%] |
| 250-500w | 19 | withheld: n human below the honesty floor (floor 20/stratum) | 12 | withheld: n AI below the honesty floor (floor 20/stratum) | ||
max FPR disparity across length_bucket: 2.2% (worst 0-100w 4.8% vs best 100-250w 2.6%)
reproduce:
gaige run --corpus hc3-mini --n 100 --seed 17 --detector fast-detect-gpt --model tiiuae/falcon-7b --quant 4bit --device cuda --max-tokens 1024
rendered from the committed export index; the page recomputes nothing and is rebuilt whole on every refresh.