PHASE 22 / SEPARATE DETECTOR FOLLOW-UPS

Representation first. Occlusion second.

The exact recovered detector was evaluated on 86 protected frames from two videos. These retrospective detector studies have their own records; the frozen Phase 22 simulation verdict remains separate.

ORIGINAL RAW-SOURCE SWEEPINCONCLUSIVE

Clean gate FAILED: raw mAP50 0.423908 versus historical 0.426627, beyond tolerance 0.001. No nonzero treatment inference ran.

STAGE A / REPRESENTATIONPASS

Q95 reproduced all four historical clean metrics. 86 × five representations = 430 cases. Raw and Q95 had zero binary success transitions; raw and lossless PNG predictions were identical.

SEPARATE OCCLUSION V2COMPLETE · CLEAN GATE PASS

86 × six frozen centered-mask doses = 516 cases. Masks use Q95 canonical pixels; nonzero treatments are lossless PNGs. This result does not change the original failed sweep.

DESCRIPTIVE DOSE RESPONSE

53 to 48 detected targets across the grid.

Object recall changed from 61.6% at zero dose to 55.8% at nominal 75% occlusion. Best target IoU changed from 0.609 to 0.554; best-overlap confidence from 0.312 to 0.065. Each dose contains the same 86 source frames.

Four descriptive curves showing object recall, all-target frame success, best target IoU, and best-overlap confidence against achieved visible annotation-box fraction.
Annotation-box visibility is a pixel proxy. Lines connect the six frozen interventions; they do not identify a population or safety threshold.
Requested occlusionMean box visibilityObject recallFrame successFP/frameBest IoUBest-overlap confidence
0.0%100.0%53/86 · 61.6%53/86275.300.6090.312
15.0%85.0%52/86 · 60.5%52/86279.170.6040.249
30.0%69.8%50/86 · 58.1%50/86282.090.5980.178
45.0%55.4%51/86 · 59.3%51/86283.570.5710.101
60.0%40.1%50/86 · 58.1%50/86283.140.5610.068
75.0%25.4%48/86 · 55.8%48/86281.200.5540.065

Frame success requires all targets to match and permits false positives. This recall uses the frozen confidence floor 0.001 and IoU ≥ 0.50. Aggregate clean-gate recall uses the official evaluator's operating point and is a different measure.

V2 retains the original runner's per-dose prediction path; Stage A exported a validation pass. Zero-dose frame success is 53/86 in v2 and 51/86 in Stage A's Q95 validation pass. These different execution passes retain their original preprocessing defaults; this count difference is not an additional representation transition.

Where do the observed losses begin?

33 of the 86 frames already fail at zero added dose. Among the 53 clean successes, first sampled new losses occur for two frames at requested occlusion 15%, three at 30% and two at 75%; 46 have no sampled new loss. One frame recovers after an earlier loss.

The confidence and localization summaries decline across doses, while recall is modestly and nonmonotonically lower. No preregistered breakpoint exists, and these six points do not establish a universal abrupt threshold.

Scene-specific vulnerability is concentrated in land_pad: 33/66 to 28/66 matched frames. land_pad2 retains 20/20 at every sampled dose. There is only one annotated object class; this does not establish a causal object-size effect or generalize to other scenes.

Verified saved-data curves: matched-frame fraction changes from 53 of 86 to 48 of 86, while mean best-overlap confidence falls from 0.312 to 0.065 as measured annotation-box occlusion increases.
The 1.0.0rc1 summary figure is copied byte-for-byte from the committed analysis. Saved detections were replay-verified; no new controlled-occlusion detector inference ran.
EVIDENCE BOUNDARY

86 paired sources. Two dependent videos.

516 cases are repeated views of 86 source frames. The original plan supports descriptive paired changes from dose zero and adjacent-dose transitions, with no population significance tests or confidence intervals. No preregistered threshold criterion exists: threshold finding is NOT APPLICABLE.

Best-overlap confidence is the score of the prediction with greatest target overlap. Matched-TP IoU and confidence are survivor-only summaries in the downloadable tables. Confidence is a model score, not a probability of a safe landing. Annotation-box visibility is not measured physical pad surface visibility. Previously inspected frames and this single synthetic intervention do not establish flight safety or independent-session generalization.

PROVENANCE / REPLAY

Check the frozen records.

Recovery source: f090da03d20b2425c5addccb4d19117c8991bcc1. Only its ten research files were imported; the old website branch was not merged.

Original protocol SHA:
e169e16ea40ca44ee60ca8b6074d054c33aeeaf1a3d24e55fa4cd3b8581882d7

Separate v2 protocol SHA:
7d9229050192a500c42daf1d05b4afad8cdefd41f9f408d203248d14635a9600

The lost original implementation/runtime receipt remains unavailable. The recovered publication runner reproduces both complete historical geometry manifests. V2 records its own Stage A CPU runtime and clean gate; it is not relabeled as the Phase 23 CUDA runtime. The latest saved-artifact check replayed 145,169 predictions, 516 frame rows and 516 target rows.