Region 10 ESC RFP R10-1193 — the reproduction package and the AI system card.
Every performance number in our response to this solicitation was produced under a protocol fixed in writing before any result existed. The protocol, the scripts, every seed, every denominator and the raw per-image results are on this page. Not on request. Not under NDA.
1. What this page is
This is the public evidence page for Precision Federal's response to Region 10 Education Service Center RFP R10-1193 — Threat Detection, Screening and Emergency Response Solutions, solicited on behalf of Equalis Group. It publishes two documents: the reproduction package behind Appendix A of that response, and the AI system card that is Appendix C.
Precision Federal has responded to this solicitation. No contract has been awarded to us under it, and nothing on this page describes a contract, a price or an ordering process. If Region 10 ESC awards a contract to Precision Federal, the contract price list, the ordering instructions and the Member documentation will be published on this same page within ten business days of that award — that is the commitment made in our response, and this page is where it lands.
What is here today. The measurement protocol as written, the results including the ones that cost us a claim, the scripts and raw results as a downloadable archive with checksums, and the full AI system card. What is not here today: any contract, price list, ordering instruction or Member document. Those appear only after award.
2. The measured protocol, and what it measured
This section states the protocol and the results exactly as Appendix A of the response states them. The comparison arms, the data, the metric definitions, the statistics and the criteria for withdrawing the claim were all fixed in writing before any result existed, in an append-only baseline registry dated 2026-09-08. That registry is the first file in the archive below.
2.1 What was fixed in advance
| Fixed in advance | What was fixed |
|---|---|
| The arms | Four, including the arm a district already has and a published third-party comparator — chosen before we knew how any of them would score |
| The data | One training set, two held-out sets, named with their licences and sizes; the held-out sets are never seen in training and never used to choose a checkpoint, a threshold or a seed |
| The metrics | Scene detection, alarm activation time, frame-level false-alarm rate, nuisance-alarm rate and throughput, each defined to the decimal so none of them can be redefined afterwards to fit a result |
| The statistics | Wilson intervals for proportions; exact McNemar for the paired 30-scene outcome; Wilcoxon signed-rank with a bootstrap median-difference interval for paired activation times |
| The tie rule | A 95% interval that crosses zero is written as a tie, never as a win |
| The drop rule | No arm and no seed is removed after the fact; all of them appear whatever they show |
| The falsification criteria | The specific results that would make us withdraw the claim rather than explain it (section 2.5 below) |
2.2 The data
| Set | What it is | Size | Role | Licence / source |
|---|---|---|---|---|
| TRAIN | fcakyon/gun-object-detection (Hugging Face), COCO format, 416×416 | 3,761 training images / 4,494 annotations; 905 validation images / 1,062 annotations; classes grenade, knife, pistol, rifle | Training and checkpoint selection only | CC BY 4.0 |
| HELD-OUT A | The 30 published test scenes of Olmos, Tabik and Herrera, from the authors' own public repository | 30 scenes, 661 frames, 1000×1000, source video at 5.000 fps | Scene-level detection and alarm activation time | Public repository of the Neurocomputing 275 (2018) 66–72 paper |
| HELD-OUT B | COCO val2017 with instances_val2017.json | 5,000 images | False-alarm and nuisance-alarm rate | COCO Terms of Use |
HELD-OUT A and HELD-OUT B are never seen during training and never used to select a checkpoint, a confidence threshold or a seed. The operating threshold is chosen once, on the training set's own validation split, and then frozen for every held-out run.
2.3 The four arms
- A0 — DO NOTHING. No automated detection: the status quo in most K-12 buildings that already own cameras. Zero scenes detected and zero false alarms, both by construction.
- A1 — THE ANALYTIC A DISTRICT ALREADY HAS. Stock YOLO11s on COCO weights, unmodified, alarming on its person class — the COCO label set it ships with contains no firearm class at all — at the stock default confidence of 0.25. The threshold rule for this arm was written into the registry before A1 was run, precisely so that A1 could not be made to look bad by handing it a threshold tuned for our own detector.
- A2 — THE PUBLISHED COMPARATOR. Olmos, Tabik and Herrera, Automatic Handgun Detection Alarm in Videos Using Deep Learning, Neurocomputing 275 (2018) 66–72 — Faster R-CNN with VGG-16, evaluated by its own authors on HELD-OUT A. We did not run this arm and we do not need to: the authors published their per-scene result in the file name of every scene archive, so the comparator's numbers come from the comparator, not from us.
- A3 — PRECISION FEDERAL. YOLO11s fine-tuned on TRAIN, pistol and rifle merged into a single firearm class, knife and grenade kept as separate classes and excluded from the firearm alarm. Five random seeds, all five reported.
2.4 The metric definitions, and the hardware
- Scene detection (binary, per scene). A scene counts as detected if the alarm rule fires anywhere in it. The alarm rule is five successive frames each containing at least one qualifying box at or above the frozen threshold — the same k=5 rule the published comparator uses, so the two are comparable.
- AATpI — alarm activation time per interval, in seconds. The index of the first frame of the first run of five successive detections, divided by 5.000, the measured frame rate of the seven source videos.
- False-alarm rate (frame level). The share of the 5,000 HELD-OUT B images on which one or more qualifying boxes fire at the frozen threshold. Every firing image is written out by file name, with its confidence and box count, into the results file — COCO has no firearm category, so we do not assume the negative set is free of firearms, and the count can be audited image by image.
- Nuisance-alarm rate. The same measurement restricted to a nuisance subset declared before the run: COCO val2017 images whose annotations include cell phone, remote, scissors, knife, hair drier or bottle. That subset is 817 of the 5,000 images.
- Throughput. Milliseconds per frame at p50 and p95 — never a single average — measured at 640×640 and at 1000×1000, after 20 warm-up frames and over 200 timed frames, on named hardware.
Hardware and software, named. Apple M1 Max, 10 CPU cores, 64 GB unified memory, macOS 26.6.1; PyTorch 2.14.0 on the MPS backend; Ultralytics 8.4.103; Python 3.14.7. One machine, no cluster, no quantisation, batch size 1 at inference.
2.5 What would make us withdraw the claim
Written into the registry before the first run, and reproduced here without softening: If A3 detects fewer than 24 of 30 scenes, or its median AATpI is worse than the comparator's by more than 0.2 s with an interval excluding zero, or its false-alarm rate on 5,000 COCO images exceeds 5%, the Appendix A claim is withdrawn and the appendix reports the negative result and what we would change.
2.6 Result — arm A1, the analytic a district already owns
Measured 2026-09-08 on the frozen protocol above.
| Measurement | Arm A1 result | 95% confidence interval |
|---|---|---|
| Scenes detected, HELD-OUT A | 30 of 30 | 88.6% – 100.0% |
| Alarm activation time, all 30 scenes | 0.0 s in every scene | 0.0 s – 0.0 s — degenerate by construction |
| False-alarm rate, 5,000 ordinary photographs (HELD-OUT B) | 2,689 of 5,000 — 53.8% | 52.4% – 55.2% |
| Nuisance-alarm rate, 817 hand-held-object images | 507 of 817 — 62.1% | 58.7% – 65.3% |
Read those four rows together, because separately each one is misleading. On the detection benchmark this arm looks flawless: it fires in all 30 scenes, instantly. On 5,000 ordinary photographs the same arm fires more than half the time, and on images where a person is merely holding a phone, a bottle or a pair of scissors it fires five times in eight. It is not a broken model — it is an excellent model answering a different question from the one a school is asking. Detection without specificity is not detection. It is an alarm a campus learns to ignore by the second week, and an alarm that is ignored has a detection rate of zero on the day it matters.
2.7 The published comparator — arm A2
Taken from the authors' own published archive names in Scene-N-PF-AATpI=Xs.zip, not from any run of ours:
| Measurement | Published result |
|---|---|
| Scenes detected, HELD-OUT A | 27 of 30 — scenes 3, 23 and 26 NotDetected |
| Alarm activation time | 0 s in 21 scenes, 0.2 s in 1, 0.4 s in 1, 0.6 s in 2, 1.2 s in 1, 2.4 s in 1 |
| Throughput, as published | 5.3 fps at 1000×1000 on an NVIDIA Titan X |
| Frame-level false positives, as published | 57 on the authors' own 304-frame negative set (18.75%) |
The comparator's false-positive figure is context, never a beat. It was measured on a different negative set from ours, and the registry forbids comparing the two — a rule written before we knew which way the comparison would fall. The scene-detection and activation-time columns are comparable, because the scenes, the frames and the k=5 alarm rule are identical.
2.8 Result — arm A3, Precision Federal, and the determination
Every seed we trained is in this table. None was dropped, and the threshold was frozen before any of it ran.
| Measurement | Arm A3 result, five seeds | 95% confidence interval |
|---|---|---|
| Scenes detected, HELD-OUT A | 9, 16, 16, 8, 12 of 30 — median 12, pooled 61 of 150 = 40.7% | 33.1% – 48.7% |
| Scenes detected, same protocol at 1024-pixel input | 15, 17, 14, 9, 8 of 30 — median 14, pooled 63 of 150 = 42.0% | 34.4% – 50.0% |
| Alarm activation time, scenes both arms detected | median difference +0.0 s to +0.6 s vs the comparator — every interval touches zero | tie on all five seeds |
| False-alarm rate, 5,000 ordinary photographs | 1,064 / 1,087 / 1,090 / 1,136 / 1,406 of 5,000 — 21.3% to 28.1% | e.g. seed 0: 20.6% – 22.9% |
| Nuisance-alarm rate, 817 hand-held-object images | 20.2% to 30.7% | e.g. seed 0: 17.6% – 23.1% |
| Throughput, p50 / p95 ms per frame | 23.9 / 29.1 at 640 px · 29.8 / 35.5 at 1000 px on the laptop GPU; 60.5 / 67.1 and 186.3 / 194.1 CPU-only | 200 timed frames per configuration |
The determination. Against the published comparator's 27 of 30, the pooled difference is −0.493 (95% CI −0.593 to −0.318) at 640 pixels and −0.480 (−0.580 to −0.304) at 1024. The interval excludes zero, so under our own separation rule this is a loss, not a tie; exact McNemar on the paired 30-scene outcome rejects at every seed. Two of the three criteria above are met — fewer than 24 of 30 scenes, and a false-alarm rate above 5% — so the claim is withdrawn. The activation-time criterion was not met: when this detector alarms at all, it is not measurably slower than the comparator.
One row is worth a district's attention even so. The same table that costs us the detection claim carries the false-alarm rate: 21% against 54% for the analytic a district already owns, on identical data, and 20% against 62% on nuisance. The 1024-pixel row exists because we suspected, after seeing the first result, that the input size was shrinking a small handgun below what the detector could resolve. We wrote that suspicion into the registry, ran it, and it moved 40.7% to 42.0% — it did not rescue the claim, and we print it rather than quietly keeping the better half.
A school-safety performance claim a district cannot check is worth nothing to that district, and a claim a competitor cannot attack has not been tested. That is why this page exists, and why the unflattering rows are on it.
3. The reproduction package
Everything needed to re-run the measurement above, at no cost and under no NDA. The three datasets are public and fetched by the script in the archive; they are roughly 2.2 GB and are not redistributed here.
Download the reproduction package (ZIP, 97.6 KB) File manifest (JSON, with SHA-256 for every file)
| Archive | Bytes | SHA-256 |
|---|---|---|
| r10-1193-reproduction-package.zip | 99,936 | d597da492fe6efd43001003ba484eabed3ccb206023d22537a2f569660b50b7a |
38 files. Verify your download with shasum -a 256 r10-1193-reproduction-package.zip. The per-file byte counts and digests are in manifest.json.
3.1 What is in it
| Path | What it does |
|---|---|
| README.md | How to run it, the hardware used, the expected outputs, and the weight digests |
| BASELINE-REGISTRY.md | The append-only registry, with its dated corrections |
| scripts/fetch_data.sh | Fetch the three public datasets |
| scripts/prepare_data.py | Build the splits and the held-out sets |
| scripts/train.py · scripts/train_all.sh | Train one seed / all five seeds |
| scripts/pick_threshold.py | Choose the operating threshold on the training split only, and freeze it |
| scripts/eval_scenes.py | Scene detection and activation time on HELD-OUT A |
| scripts/eval_negatives.py | False-alarm and nuisance-alarm rate on HELD-OUT B, writing out every firing image by name |
| scripts/bench_throughput.py | p50 and p95 milliseconds per frame at both image sizes |
| scripts/statlib.py | Wilson intervals, exact McNemar, Wilcoxon signed-rank, bootstrap median difference |
| scripts/run_A3.sh · scripts/run_A3_1024.sh | The two end-to-end evaluation runs |
| results/*.json | The raw per-scene and per-image results behind every number above |
| results/A3-STATISTICS.txt | The computed statistics, printed, every seed |
3.2 How to run it
3.3 The trained weights
The five best.pt checkpoints are 19,173,146 bytes each and are not shipped in the archive, to keep it small enough to download over a school district's connection. They are provided to any evaluator or Member on request, at no cost and under no NDA — write to bo@precisionfederal.com. Verify any copy you receive from us against these SHA-256 digests:
The starting checkpoint is the public Ultralytics yolo11s.pt, downloaded by fetch_data.sh.
4. AI system card
Structured on the NIST Artificial Intelligence Risk Management Framework (NIST AI 100-1) functions — Govern, Map, Measure, Manage. This card is written so a district's technology director, its counsel and its board can each read the part they need without asking us for a briefing. It is a public document, and it is updated with every model release.
4.1 MAP — what the system is, and what it is not
| Question | Answer |
|---|---|
| What does it do? | Detects the visual presence of a firearm, knife or grenade in a video frame from a camera the district already owns, and raises an alarm when the object is present in five successive frames. |
| What decision does it make? | One: whether to place an alert in front of a human. It makes no other decision. |
| What does it not do? | It does not identify, recognize or name any person. It does not use facial recognition, gait recognition, iris, voice or any other biometric identifier. It does not infer age, race, sex, emotion, intent or affiliation. It does not track individuals between cameras. It does not read license plates. It does not listen. |
| What is out of scope? | Concealed-weapon detection through clothing, bags or walls; behavioural threat prediction; anything that would require identifying a person. If a Member needs those, we say so rather than implying our system provides them. |
| Who is the user? | A trained district or campus staff member, or a monitoring-center operator designated by the district. |
| What is the deployment context? | Existing fixed cameras on a school, campus or public building, indoors and outdoors, at the resolutions and frame rates the district already runs. |
The design decision that drives everything else. The system was deliberately built as an object detector, not a person-recognition system. That choice costs us capability a competitor might advertise, and it buys the district three things: the biometric-privacy statutes do not attach, the failure modes are bounded and testable, and the question "what does it know about my students?" has a one-word answer — nothing.
4.2 MAP — data
| Data | What it is | Provenance and licence |
|---|---|---|
| Training data | 3,761 images with 4,494 annotated objects, split into three classes: firearm (pistols and rifles merged, 2,236 instances), knife (845), grenade (1,413) | fcakyon/gun-object-detection, published under CC BY 4.0 |
| Validation data | 905 images with 1,062 annotated objects, from the same source, held apart from training | CC BY 4.0 |
| Held-out benchmark — detection | 30 published scenes, 661 frames, from the peer-reviewed firearm-detection benchmark of Olmos, Tabik and Herrera, Neurocomputing 275 (2018) 66–72 | Published academic benchmark, third-party |
| Held-out benchmark — false alarms | 5,000 images, COCO val2017, containing everyday objects including the nuisance classes most likely to be mistaken for a firearm — cell phone, remote control, scissors, knife, hair dryer, bottle | Public research dataset |
| Member data | Video from the Member's own cameras, processed to produce alerts | The Member's property. Never used to train, tune or evaluate any model. |
We do not train on Member data — ever, and not with permission either. A district's hallway footage does not become our training set, is not pooled with other districts, and is not used to improve a model that we then sell to somebody else. If a Member ever wants us to tune specifically for their site, that is a separate, written, paid engagement with the tuned model belonging to them.
4.3 MAP — the model and the decision logic
| Element | Specification |
|---|---|
| Architecture | Single-stage convolutional object detector (YOLO11s), trained at 640-pixel input |
| Output per frame | Zero or more bounding boxes, each with a class and a confidence score |
| Operating threshold | A single confidence threshold, selected to maximize F1 on the training and validation split only, then frozen before any held-out data is touched |
| Alarm rule | An alarm is raised at the first frame of the first run of five successive frames in which the object is detected. Alarm Activation Time is reported as that frame index divided by the source frame rate |
| Why five frames | A single-frame detection is a coin flip against motion blur and compression artifacts. Requiring five consecutive frames converts a noisy per-frame signal into a stable event and is the single largest lever on the false-alarm rate. It costs one second of latency at 5 fps and buys an order of magnitude in nuisance suppression |
| Reproducibility | Trained across five random seeds with deterministic settings; all five are reported, not the best |
| Hardware measured | Apple M1 Max, 10 CPU cores, 64 GB unified memory, macOS 26.6.1; PyTorch with the MPS backend; Ultralytics 8.4.103; Python 3.14.7 |
4.4 MEASURE — how performance is established
The numbers are in section 2 above. What belongs here is the discipline that produced them, because a number without its method is a marketing claim.
- The arms were declared before the run. Which variants, on which data, against which published figures, went into a baseline registry before any held-out data was touched — with the decision rules, including that an interval crossing zero is written as a tie and no arm is dropped after the fact.
- The threshold was frozen on training data. It was never adjusted to make a held-out result look better.
- Every seed is reported. Five seeds; we publish the distribution, not the best one.
- No firing image is assumed to be a false alarm. COCO val2017 has no firearm category, and a benchmark set can contain the thing it is supposed to lack. Every image the alarm fires on is written out by file name, with its confidence and box count, into the results file, so the count is auditable image by image rather than taken on our word.
- Falsification criteria were written in advance. The registry states the results that would cause the claim to be withdrawn rather than explained — and on this measurement, two of the three were met and the claim was withdrawn.
- Throughput is reported at p50 and p95 on named hardware, not as a single average on unnamed hardware.
- A district can run it. The datasets are public and the scripts, thresholds and environment are on this page, so the result can be reproduced or refuted.
The claim's edges. It covers the object classes, cameras and lighting conditions actually tested, and it is never converted into a promise about a specific district — which is why every deployment ends with a site acceptance test on the Member's own cameras, and why that number, not ours, is written into the Member's record.
4.5 MANAGE — human oversight
| Control | How it works |
|---|---|
| No automated action | The system does not lock a door, dispatch a responder, notify a parent, page a PSAP or trigger a lockdown on its own. Every consequential action is taken by a human who has seen the image. |
| The alert carries its evidence | Each alert presents the still frame, the bounding box, the confidence, the camera, the timestamp and a short clip — so the reviewer decides from the image, not from a score. |
| Disposition is mandatory and recorded | Every alert is dispositioned by a named human as confirmed, unfounded or unclear. The disposition record is the Member's, exportable, and the input to tuning. |
| Escalation is the district's design | Who is notified, in what order, on what device, and what happens if nobody acknowledges, is configured with the district during implementation and rehearsed in the site acceptance test. |
| Reviewer training | Every designated reviewer is trained before go-live and re-trained annually; training covers what the system can and cannot see, and how to handle an unclear image. |
4.6 MANAGE — known limitations, stated plainly
- Occlusion. An object substantially hidden by a body, a bag, clothing or furniture may not be detected. The system sees what a camera sees.
- Distance and resolution. Detection degrades as the object's pixel size falls. Cameras positioned for wide-area coverage will detect later, or not at all, compared with cameras positioned for a doorway.
- Lighting and motion. Low light, strong backlight, heavy compression and fast motion all reduce detection.
- Look-alikes. Tools, toys, sports equipment and some handheld electronics can produce a false alarm. The five-frame rule suppresses most, not all.
- Camera angle. A camera looking down at a steep angle sees a different object profile than the training distribution.
- It is not a screening system. It does not replace a walkthrough detector, a bag search or a school resource officer, and we will not let a district believe otherwise.
- It cannot see what is not on camera. Coverage is a function of where the district's cameras point, which is why our site assessment produces a written coverage map and an explicit list of what is not covered.
4.7 GOVERN — model lifecycle and change control
| Control | Commitment |
|---|---|
| Versioning | Every model has a version number, a training-data manifest, a threshold, and a benchmark result recorded at release. |
| Release cadence | Model updates are released no more than quarterly and no less than annually. Members are notified 30 days before a release. |
| Regression gate | No model ships unless it matches or beats the previous version on the held-out benchmark and does not increase the false-alarm rate. Both numbers are published in the release note. |
| Member choice | A Member may pin a version and defer an update for up to two release cycles. An update is never pushed silently. |
| Rollback | Any release can be rolled back to the prior version within one business day, on request, without charge. |
| Drift monitoring | Per-camera alarm rates are baselined and monitored; a statistically significant shift raises an internal ticket and appears in the Member's monthly report. |
| Annual re-benchmark | Every Member's site acceptance test is re-run annually on their own cameras, and the result is written into their record whether it improved or not. |
| Incident process | A missed detection or a cluster of false alarms is a recorded incident with a root cause, a corrective action, and a written closure note to the Member. |
| Ownership | Model governance is owned by Bo Peng and Reid McKinley; privacy and security governance by Keith Parker. Named people, not a committee. |
4.8 GOVERN — the standards this system is built to
| Standard | How it is applied |
|---|---|
| NIST AI 100-1 (AI RMF 1.0) | The structure of this card; Govern/Map/Measure/Manage applied to a single narrow model |
| NIST SP 800-171 | Control baseline for the platform handling Member data |
| NIST SP 800-53 families AC, AU, IR, SC, SI | Access control, audit, incident response, system and communications protection, system and information integrity |
| FERPA (20 U.S.C. § 1232g; 34 CFR Part 99) | Data-handling posture and the school-official designation |
| Section 508 / WCAG 2.1 AA | The operator console and every report |
| 2 CFR 200.216 / Section 889 | We provide no covered telecommunications or video surveillance equipment, and we screen the Member's camera inventory for it at assessment |
4.9 The one-page summary for a school board
Precision Federal's system looks at video from cameras the district already owns and answers one question: is there a firearm, knife or grenade visible right now? It does not know who anyone is. It cannot recognize a face. It never acts on its own — it puts a picture in front of a trained person, who decides. Its accuracy is reported with the data and the method, and the code is on this page. Before it goes live, it is tested on the district's own cameras and that number is written down. Every alert the district's staff handle is recorded, and every month the district gets a report of what the system did.
A board member who reads nothing else should read this: the district is not asked to take our word for anything. The benchmark, the thresholds, the scripts and the per-image results are in section 2 and in the archive; what the system cannot see is in section 4.6. Both were written before the district signs, not after something goes wrong.
5. What appears on this page after award
Our response to RFP R10-1193 commits that, within ten business days of an award to Precision Federal, this page will also carry: the contract price list row for row, the ordering instructions, the contract documents, the W-9, and a Member FAQ — each linked from the site's primary navigation. Until an award is made, none of that exists and none of it is on this page.
6. Contact
Questions about the measurement, requests for the trained weights, or anything else on this page go to Bo Peng, Founder and Chief Technology Officer: bo@precisionfederal.com. We answer within one business day.