5 committed receipts · 2 corpora · latest 2026-08-10 · every number below renders from the committed JSON, nothing typed by hand
fast-detect-gpt × hc3-mini(n=100,seed=17)20260722-163959-fast-detect-gpt · gaige 0.0.1
AUROC: 0.9720 [0.9458, 0.9938] (n_boot 1000)
fingerprint · tiiuae/falcon-7b
instrument: tiiuae/falcon-7b
quant: 4bit (verified: 128 modules)
device: cuda · Linux/x86_64
versions: torch 2.13.0+cu130 · transformers 4.49.0
corpus: hc3-mini(n=100,seed=17) · sha256 7d2819d3e83bd10d…
| target FPR | threshold | TPR | TPR 95% CI |
| 1% | 2.1229 | 86.0% | 79.0% to 92.0% |
| 5% | 1.8319 | 91.0% | 85.0% to 96.0% |
alpha=0.005: alpha=0.005 needs >= 199 human calibration samples, got 100. A tighter guarantee than your data supports is not a guarantee.
| conformal alpha | threshold | order stat | TPR at threshold |
| 0.05 | 1.8468 | 96/100 | 90.0% |
| 0.01 | 2.4446 | 100/100 | 76.0% |
max FPR disparity on length_bucket at target FPR 1%: 0.0% (worst 0-100w 0.0%, best 0-100w 0.0%)
$ gaige run --corpus hc3-mini --n 100 --seed 17 --detector fast-detect-gpt --model tiiuae/falcon-7b --quant 4bit --device cuda --max-tokens 1024
binoculars × hc3-mini(n=100,seed=17)20260722-213952-binoculars · gaige 0.0.1
AUROC: 0.9992 [0.9974, 1.0000] (n_boot 1000)
fingerprint · tiiuae/falcon-7b + tiiuae/falcon-7b-instruct
instrument: tiiuae/falcon-7b + tiiuae/falcon-7b-instruct
quant: 4bit (verified: 256 modules)
device: cuda · Linux/x86_64
versions: torch 2.13.0+cu130 · transformers 4.49.0
corpus: hc3-mini(n=100,seed=17) · sha256 7d2819d3e83bd10d…
| target FPR | threshold | TPR | TPR 95% CI |
| 1% | -0.7829 | 97.0% | 94.0% to 100.0% |
| 5% | -0.8706 | 100.0% | 100.0% to 100.0% |
alpha=0.005: alpha=0.005 needs >= 199 human calibration samples, got 100. A tighter guarantee than your data supports is not a guarantee.
| conformal alpha | threshold | order stat | TPR at threshold |
| 0.05 | -0.8706 | 96/100 | 100.0% |
| 0.01 | -0.7540 | 100/100 | 95.0% |
max FPR disparity on length_bucket at target FPR 1%: 2.4% (worst 0-100w 2.4%, best 100-250w 0.0%)
$ gaige run --corpus hc3-mini --n 100 --seed 17 --detector binoculars --observer tiiuae/falcon-7b --performer tiiuae/falcon-7b-instruct --quant 4bit --device cuda --max-tokens 1024
fast-detect-gpt × raid-g2d2a2-n60-s1720260725-035507-fast-detect-gpt · gaige 0.0.1
AUROC: 0.9285 [0.9061, 0.9482] (n_boot 1000)
fingerprint · tiiuae/falcon-7b
instrument: tiiuae/falcon-7b
quant: 4bit (verified: 128 modules)
device: cuda · Linux/x86_64
versions: torch 2.13.0+cu130 · transformers 4.49.0
corpus: raid-g2d2a2-n60-s17 · sha256 54b786694f2e4c0d…
| target FPR | threshold | TPR | TPR 95% CI |
| 1% | 2.4401 | 61.5% | 57.3% to 65.8% |
| 5% | 1.8878 | 73.1% | 69.4% to 77.3% |
alpha=0.005: alpha=0.005 needs >= 199 human calibration samples, got 120. A tighter guarantee than your data supports is not a guarantee.
| conformal alpha | threshold | order stat | TPR at threshold |
| 0.05 | 1.9249 | 115/120 | 72.5% |
| 0.01 | 2.6423 | 120/120 | 56.3% |
max FPR disparity on domain at target FPR 1%: 1.7% (worst abstracts 1.7%, best reddit 0.0%)
$ gaige run --corpus corpora/raid-g2d2a2-n60-s17.jsonl --n 100 --seed 17 --detector fast-detect-gpt --model tiiuae/falcon-7b --quant 4bit --device cuda --max-tokens 1024
raw JSON: data/receipts/20260725-035507-fast-detect-gpt.jsonThis receipt was produced from a locally prepared corpus file. The corpus sha256 above pins the exact bytes; the file itself is not redistributed. RAID slices (Dugan et al., ACL 2024) are prepared from the public dataset with `gaige corpus prepare-raid` and are never redistributed.
fast-detect-gpt × raid-g2d2a2-n60-s1720260727-193225-fast-detect-gpt · gaige 0.0.1
AUROC: 0.9330 [0.9117, 0.9521] (n_boot 1000)
fingerprint · EleutherAI/gpt-neo-1.3B
instrument: EleutherAI/gpt-neo-1.3B
quant: fp32
device: cpu · Linux/x86_64
versions: torch 2.13.0+cu130 · transformers 4.49.0
corpus: raid-g2d2a2-n60-s17 · sha256 54b786694f2e4c0d…
| target FPR | threshold | TPR | TPR 95% CI |
| 1% | 2.3904 | 58.1% | 54.2% to 62.7% |
| 5% | 1.4957 | 77.7% | 74.2% to 81.5% |
alpha=0.005: alpha=0.005 needs >= 199 human calibration samples, got 120. A tighter guarantee than your data supports is not a guarantee.
| conformal alpha | threshold | order stat | TPR at threshold |
| 0.05 | 1.5727 | 115/120 | 76.5% |
| 0.01 | 2.7355 | 120/120 | 49.4% |
max FPR disparity on domain at target FPR 1%: 1.7% (worst abstracts 1.7%, best reddit 0.0%)
$ gaige run --corpus corpora/raid-g2d2a2-n60-s17.jsonl --n 100 --seed 17 --detector fast-detect-gpt --model EleutherAI/gpt-neo-1.3B --quant fp32 --device cpu --max-tokens 1024
raw JSON: data/receipts/20260727-193225-fast-detect-gpt.jsonThis receipt was produced from a locally prepared corpus file. The corpus sha256 above pins the exact bytes; the file itself is not redistributed. RAID slices (Dugan et al., ACL 2024) are prepared from the public dataset with `gaige corpus prepare-raid` and are never redistributed.
fast-detect-gpt × hc3-mini(n=100,seed=17)20260810-134151-fast-detect-gpt · gaige 0.0.4
AUROC: 0.9720 [0.9448, 0.9929] (n_boot 1000)
fingerprint · tiiuae/falcon-7b
instrument: tiiuae/falcon-7b
quant: 4bit (verified: 128 modules)
device: cuda · Linux/x86_64
versions: torch 2.13.0+cu130 · transformers 4.49.0
corpus: hc3-mini(n=100,seed=17) · sha256 7d2819d3e83bd10d…
| target FPR | threshold | TPR | TPR 95% CI |
| 1% | 2.1229 | 86.0% | 79.0% to 92.0% |
| 5% | 1.8319 | 91.0% | 85.0% to 96.0% |
alpha=0.005: alpha=0.005 needs >= 199 human calibration samples, got 100. A tighter guarantee than your data supports is not a guarantee.
| conformal alpha | threshold | order stat | TPR at threshold |
| 0.05 | 1.8468 | 96/100 | 90.0% |
| 0.01 | 2.4446 | 100/100 | 76.0% |
max FPR disparity on length_bucket at target FPR 1%: 0.0% (worst 0-100w 0.0%, best 0-100w 0.0%)
$ gaige run --corpus hc3-mini --n 100 --seed 17 --detector fast-detect-gpt --model tiiuae/falcon-7b --quant 4bit --device cuda --max-tokens 1024