The champion put to the product’s own test: a virtual fixed-wing UAV that must fly from a to b over Berlin with the model’s vision fixes as its only position source — dead reckoning in between, wind and gyro bias it cannot see, false fixes accepted like any other. Arrival is declared on the aircraft’s own estimate and scored against the truth.
100 simulated flights, random start and target ≥2 km apart, random 2–6 m/s wind. Each vision fix is the real exported champion (berlin.onnx) run on a real orthophoto crop at the aircraft’s true position and heading, classified with the frozen scorer’s own thresholds.
Corner to corner — 9.2 km in ten minutes, seven false fixes survived, arrival at 25.5 m.
North-west corner to south-east corner, 9 km. At t = 85 s a false fix — barely confident, 6.5 km wrong (the red line points to its claim, marked ×) — hijacks the navigator’s belief. The aircraft never flies to the claimed spot: only its estimate teleports there, so the dashed believed path simply breaks off and rejoins once usable fixes haul the belief back within 300 m of the truth. A late burst of false fixes on final approach does little harm: by then the estimate is fresh, so wrong answers get little weight. The flight lands 25.5 m from the target.
Seed 6, disclosed: three of six seeds on this route overshot the corner out of the mapped box and were lost — beyond the edge there is no imagery and no way back. The 100-flight record below uses random routes, failures included.
—
Bottom camera — what the model saw
—
128 × 128 px · 1 m/px · ~100 m AGL
How far the navigator's belief is from the aircraft's true position, second by second.
Between fixes the error grows — wind the navigator cannot see. Each usable fix snaps it back down. The false fix at t = 85 s is the spike: one confident wrong answer worth over 5 km of injected error — the reason the mission score prices false fixes so high. The late cluster of false fixes barely registers: a fresh estimate gives a wrong answer little weight. Square-root vertical scale, so the 20–60 m cruise detail and the kilometre spike fit one honest axis; the dashed line is the 100 m usable radius.
Every Monte-Carlo flight's miss distance — arrivals, overshoots, and the three catastrophes.
Each dot is one flight’s true distance from target at the moment it declared arrival (hollow = timed out, plotted at final distance). 91 land inside the 100 m line. Six miss moderately (120–314 m). Three fail catastrophically — and all three had their target on or beside the never-trained holdout ground: the ~3% of Berlin that the evaluation harness deliberately withholds from training as a diagnostic. Over that ground the model hallucinates confidently — 167 of 207 false fixes across all 100 flights stand on it. Excluding it, the in-flight false-fix rate is 0.77%, consistent with the 0.5% measured by the frozen eval. A deployed model would train on 100% of its box, so this exact failure is an artifact of simulating over the research configuration — reported as flown, not filtered out.
| flight | ended | true miss | false fixes | target on never-trained ground |
|---|---|---|---|---|
| 39 | declared arrival | 120 m | 3 (3 on never-trained ground) | |
| 80 | declared arrival | 125 m | 2 (2 on never-trained ground) | yes |
| 15 | declared arrival | 137 m | 0 | yes |
| 49 | declared arrival | 262 m | 2 | |
| 92 | declared arrival | 288 m | 1 | yes |
| 89 | declared arrival | 314 m | 0 | |
| 95 | timed out | 400 m | 72 (72 on never-trained ground) | yes |
| 57 | declared arrival | 1,397 m | 25 (24 on never-trained ground) | yes |
| 27 | declared arrival | 7,362 m | 10 (10 on never-trained ground) | yes |
The test is honest only if its seams are visible.
The bottom line
Asked to actually fly, the champion delivers 91 flights in 100 to within 100 m — and every failure is on the page, three of them explained by ground the research harness deliberately never taught it.