R10-1193 arm A3 — statistics computed under BASELINE-REGISTRY.md §5, generated by this script.
A2 (published comparator) detected 27 of 30; median AATpI over its detected scenes = 0.0 s

=== SCENE DETECTION @ imgsz 640 (the PRE-REGISTERED run) ===
arm            det/30   Wilson 95%         McNemar vs A2 (b = A2 only, c = A3 only)
A3_seed0         9/30   [0.167, 0.479]     b=18 c= 0  p=7.629e-06
A3_seed1        16/30   [0.361, 0.698]     b=11 c= 0  p=9.766e-04
A3_seed2        16/30   [0.361, 0.698]     b=12 c= 1  p=3.418e-03
A3_seed3         8/30   [0.142, 0.444]     b=19 c= 0  p=3.815e-06
A3_seed4        12/30   [0.246, 0.577]     b=15 c= 0  p=6.104e-05

median across the five seeds: 12/30   range 8-16   pooled 61/150 = 40.7%
pooled Wilson 95%: [0.331, 0.487]
A2 27/30 = 90.0%  Wilson 95% [0.744, 0.965]
pooled A3 - A2 = -0.493, 95% CI [-0.593, -0.318] -> EXCLUDES zero: the comparator BEATS us


=== AATpI, paired on the scenes BOTH arms detected ===
(AATpI is time-to-first-alarm; a positive difference means A3 alarms LATER than the comparator.)
A3_seed0     n= 9  median A3 0.600s  median A2 0.000s  median diff +0.400s  boot 95% [+0.000, +1.400]  wilcoxon W=5.0 p=0.1719  -> TIE (crosses 0)
A3_seed1     n=16  median A3 0.400s  median A2 0.000s  median diff +0.000s  boot 95% [+0.000, +0.400]  wilcoxon W=6.0 p=0.1094  -> TIE (crosses 0)
A3_seed2     n=15  median A3 0.200s  median A2 0.000s  median diff +0.000s  boot 95% [+0.000, +0.800]  wilcoxon W=9.0 p=0.1289  -> TIE (crosses 0)
A3_seed3     n= 8  median A3 0.600s  median A2 0.000s  median diff +0.600s  boot 95% [+0.000, +1.000]  wilcoxon W=2.0 p=0.0469  -> TIE (crosses 0)
A3_seed4     n=12  median A3 0.400s  median A2 0.000s  median diff +0.200s  boot 95% [+0.000, +0.400]  wilcoxon W=5.5 p=0.2031  -> TIE (crosses 0)

=== FRAME-LEVEL FALSE ALARM, 5,000 COCO val2017 images (no firearm present) ===
A1            2689/5000 = 53.8%  Wilson [0.524, 0.552]   nuisance 507/817 = 62.1%  [0.587, 0.653]
A3_seed0      1087/5000 = 21.7%  Wilson [0.206, 0.229]   nuisance 165/817 = 20.2%  [0.176, 0.231]
A3_seed1      1406/5000 = 28.1%  Wilson [0.269, 0.294]   nuisance 251/817 = 30.7%  [0.277, 0.340]
A3_seed2      1090/5000 = 21.8%  Wilson [0.207, 0.230]   nuisance 184/817 = 22.5%  [0.198, 0.255]
A3_seed3      1136/5000 = 22.7%  Wilson [0.216, 0.239]   nuisance 166/817 = 20.3%  [0.177, 0.232]
A3_seed4      1064/5000 = 21.3%  Wilson [0.202, 0.224]   nuisance 173/817 = 21.2%  [0.185, 0.241]

=== BASELINE-REGISTRY.md §6 — WHAT WOULD FALSIFY OUR CLAIM ===
(1) 'fewer than 24 of 30 scenes'  -> A3 = 9, 16, 16, 8, 12; every seed below 24 ....... MET
(2) 'median AATpI worse by >0.2 s with an interval excluding zero' -> see the table above
(3) 'false-alarm rate on 5,000 COCO images exceeds 5%' -> A3 = 21.7%-28.1% ........... MET

Per §6 the Appendix A performance claim for arm A3 is WITHDRAWN. Appendix A reports this
negative result and what we would change. The comparator wins scene detection outright.

=== POST-HOC SENSITIVITY: same protocol at imgsz 1024 (registry §11, declared before the run) ===
seed       640    1024   1024 Wilson 95%     
0          9/30    15/30   [0.332, 0.668]
1         16/30    17/30   [0.392, 0.726]
2         16/30    14/30   [0.302, 0.639]
3          8/30     9/30   [0.167, 0.479]
4         12/30     8/30   [0.142, 0.444]

median 640 = 12/30    median 1024 = 14/30
pooled 640  = 61/150 = 40.7%  Wilson [0.331, 0.487]
pooled 1024 = 63/150 = 42.0%  Wilson [0.344, 0.500]
1024 vs the comparator's 27/30: difference -0.480, 95% CI [-0.580, -0.304] -> EXCLUDES zero: the comparator still beats us

BOTH image sizes fall far below the pre-registered floor of 24/30 (best single seed = 17).
The larger image size does NOT rescue the claim. The §6 withdrawal stands on either run.
