Two KIOS detectors
320px YOLO11n baseline versus 480px YOLO11n with image augmentations. The source’s “Phase 22 Baseline” is a detector, separate from the Phase 22 simulation model.
Clean and blur improved. Occlusion and mixed stress got worse.
Illustrative views from the published set. The aggregate table below summarizes all 86 protected frames.






These are condition examples, not detector-output overlays or six independent test sets. Published metrics are condition-level aggregates over the same 86 protected frames.
Deltas are percentage points (pp).
Swipe sideways for every metric →
| Condition | Baseline mAP50 | Phase 23 mAP50 | mAP50 Δ | Baseline recall | Phase 23 recall | Recall Δ |
|---|---|---|---|---|---|---|
| Clean | 42.7% | 55.5% | +12.9 pp | 38.4% | 59.3% | +20.9 pp |
| Blur | 41.3% | 58.6% | +17.3 pp | 37.8% | 57.0% | +19.2 pp |
| Low light | 41.7% | 42.4% | +0.7 pp | 37.2% | 44.2% | +7.0 pp |
| Noise | 43.3% | 59.6% | +16.4 pp | 39.8% | 54.7% | +14.9 pp |
| Occlusion | 34.8% | 13.8% | −21.0 pp | 32.6% | 24.4% | −8.1 pp |
| Mixed stress | 18.5% | 12.9% | −5.5 pp | 18.6% | 8.1% | −10.5 pp |
Phase 23 trained the detector; Phase 24 reanalyzed the results.
320px YOLO11n baseline versus 480px YOLO11n with image augmentations. The source’s “Phase 22 Baseline” is a detector, separate from the Phase 22 simulation model.
252 train, 64 validation, 20 embargoed, and 86 protected test frames. Conditions reuse the test frames.
The original best.pt checkpoint was recovered without substitution. SHA-256: 43240be969708c32b3e11340c9846b3a588baf361378398677c794d16a455310. Fresh inference matched all 24 published aggregate cells with maximum delta 0.0, then regenerated four prediction tables byte-for-byte. This was a pinned original-model replay, not a newly trained model or a complete clean-room reconstruction of every original dependency.
The recovered lock matched Python 3.13.2, Ultralytics 8.4.160, Torch 2.7.0+cu118, NumPy 2.2.3 and Pillow 11.0.0. Phase 23 retained CUDA device 0 and 480 px; the baseline retained its separate 320 px CPU settings.
Checkpoint packaging alone does not rerun inference. The historical bundler remains available as py scripts\bundle_phase23_checkpoint.py; use the current replay commands and receipt to check outcomes.
The 516 cross-model views and 860 within-model clean-versus-stressed pairs answer different questions. Neither is independent validation. Read the paired research record →