AEGISLANDMODEL VASE
Phase 23 · KIOS protected detector benchmark

The detector gained recall.

Clean and blur improved. Occlusion and mixed stress got worse.

Detector comparison. 86 protected real-video frames; separate from the frozen Phase 1–22 simulation record.
59.3%Clean recall · +20.9 percentage points vs detector baseline
55.5%Clean mAP50 · +12.9 percentage points
13.8%Occlusion mAP50 · −21.0 percentage points
8.1%Mixed-stress recall · −10.5 percentage points
Mixed result: clean precision fell 79.6% → 71.7%; mAP50–95 fell 27.2% → 22.3%.

Six condition examples

Illustrative views from the published set. The aggregate table below summarizes all 86 protected frames.

What changed by condition

Deltas are percentage points (pp).

Swipe sideways for every metric →

ConditionBaseline mAP50Phase 23 mAP50mAP50 ΔBaseline recallPhase 23 recallRecall Δ
Clean42.7%55.5%+12.9 pp38.4%59.3%+20.9 pp
Blur41.3%58.6%+17.3 pp37.8%57.0%+19.2 pp
Low light41.7%42.4%+0.7 pp37.2%44.2%+7.0 pp
Noise43.3%59.6%+16.4 pp39.8%54.7%+14.9 pp
Occlusion34.8%13.8%−21.0 pp32.6%24.4%−8.1 pp
Mixed stress18.5%12.9%−5.5 pp18.6%8.1%−10.5 pp
The weak tail: both mAP50 and recall fell under occlusion and mixed stress. See the Phase 24 charts.

Method and boundary

Phase 23 trained the detector; Phase 24 reanalyzed the results.

Two KIOS detectors

320px YOLO11n baseline versus 480px YOLO11n with image augmentations. The source’s “Phase 22 Baseline” is a detector, separate from the Phase 22 simulation model.

Temporal split

252 train, 64 validation, 20 embargoed, and 86 protected test frames. Conditions reuse the test frames.

Original Phase 23 replay verified: 24 / 24 metric cells matched

The original best.pt checkpoint was recovered without substitution. SHA-256: 43240be969708c32b3e11340c9846b3a588baf361378398677c794d16a455310. Fresh inference matched all 24 published aggregate cells with maximum delta 0.0, then regenerated four prediction tables byte-for-byte. This was a pinned original-model replay, not a newly trained model or a complete clean-room reconstruction of every original dependency.

The recovered lock matched Python 3.13.2, Ultralytics 8.4.160, Torch 2.7.0+cu118, NumPy 2.2.3 and Pillow 11.0.0. Phase 23 retained CUDA device 0 and 480 px; the baseline retained its separate 320 px CPU settings.

Checkpoint packaging alone does not rerun inference. The historical bundler remains available as py scripts\bundle_phase23_checkpoint.py; use the current replay commands and receipt to check outcomes.

The 516 cross-model views and 860 within-model clean-versus-stressed pairs answer different questions. Neither is independent validation. Read the paired research record →