Where we are: The experiment record

Every experiment the autonomous loop has run, across all 5 evaluation eras — kept and discarded. Each row was pre-registered before training (hypothesis, method, expected outcome, architecture figure), then trained on a single GPU and scored. Every score on this page is measured by one frozen ruler: the worst mission score on held-out viewpoints — (1 − usable-fix rate) + false-fix rate, the product requirement rather than an error statistic (currently 3,007 KB, hard limit 4 MiB). The earlier eras were judged by different rulers at the time, and where those two disagree the row says so. One agent designs, one implements; failures stay on the record. New here? Start with the overview.

Where we are right now

Goal reached: mission score 0.040 — very nearly every frame yields a usable fix.

5 experiments in this era, 4 kept — and 76 before it, across 4 earlier evaluation eras. The chart below shows all 81. Cost of the whole search: $364.64 — $342.51 of agent tokens (equivalent API cost; the work ran on a Max subscription) plus $22.13 of rented GPU, recorded for 61 of 81 experiments.

Kept improvement Discarded Running best Goal (score 0) ×Gated fail Holdout check Pivot-directed Not on this ruler Evaluation era updated 2026-08-01 14:57 UTC
Mission score per experiment, all 81 of them — (1 − usable-fix rate) + false-fix rate, lower is better
bootstrap 4 areas · 6 lighting buckets · region holdout berlin only viewpoint holdout mission score 0 0.25 0.5 0.75 1 1.5 2 score 1 — no better than saying nothing 0 — every frame a usable fix × × × × × × × × × × × × × × × × × 1 4 1 4 7 10 13 16 19 22 25 28 31 34 37 40 43 46 49 52 55 61 1 1 1 4 experiment № — restarting at 1 with each evaluation era
#Experiment Category Mission score Time Cost Status
5.5 Decode-consistent training targets: bilinear tent cell labels + residual offset against the model's own pooled centroid 7a1314f9
new best: 0.090 → 0.040
loss 0.040 22 m 58 s $3.90 kept
5.4 Local belief pooling decode: confidence = softmax mass on the 3x3 cell neighborhood, fix = its mass-weighted centroid 7958b03d
new best: 0.183 → 0.090
architecture 0.090 20 m 36 s $2.75 kept
5.3 Full-lattice coverage training for the map-cell classifier: every epoch is a complete pass over all 72,712 vantages 014a6196
new best: 2.001 → 0.183
training 0.183 22 m 28 s $3.61 kept
5.2 Map-cell classification decode: pick 1 of ~3,000 128 m tiles, confidence = softmax mass 9011afe6
abstained on too many frames in a cell
architecture gated fail 9 m 24 s $4.25 gated fail
5.1 Baseline seed: current model/ code, no design agent 92d1f12e
starting line: mission score 2.001
other 2.001 47 s kept
4.4 Re-implementation of lost experiment 2 — classify-then-refine: 32x32 map-cell softmax + within-cell offset replaces direct coordinate regression c0f3cd4e
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.840 — worse than silence.
architecture 1.840 3 m 36 s KEPT
4.3 Coverage over repetition: spend the same 48k-step budget on ~48k fresh-rotated distinct vantages instead of 6000 static crops replayed 8x f74a8666
The loop discarded this — no improvement under the metric then in force.
training 1.995 6 m 36 s DISCARDED
4.2 Classify-then-refine: 32x32 map-cell softmax + within-cell offset replaces direct coordinate regression 4affa746
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.840 — worse than silence.
architecture 1.840 6 m 20 s KEPT
4.1 Baseline seed on the corrected viewpoint-holdout split 8bdcc185
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 2.000 — worse than silence.
other 2.000 44 s KEPT
3.5 Per-epoch position resampling: expose all ~45k Berlin train vantages instead of a frozen 6k subset 0ede8d45
The loop kept this — it beat the running best under the metric then in force. It cannot be placed on today's ruler at all — see the row for why.
training 16 m 23 s NO SCORE
3.4 Hierarchical decode: 16x16 coarse-cell classification + per-cell continuous offset regression from a layout-preserving descriptor f625d8d7
The loop discarded this — no improvement under the metric then in force.
architecture 1.653 17 m 29 s $3.59 DISCARDED
3.3 Re-land the validated 64x64 grid-classification head with a convergence-scale training budget (x12 epochs, cosine LR, fresh per-epoch rotations) bea8a408
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.556 — worse than silence.
training 1.556 17 m 28 s $3.70 KEPT
3.2 Reparameterize localization as 64x64 map-grid classification with in-graph soft-argmax d9805aa2
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
architecture 6 m 37 s $3.15 GATED FAIL
3.1 Baseline seed: original TinyLocNet, fresh berlin-slim lineage 28f2ebe1
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.996 — worse than silence.
other 1.996 16 s KEPT
2.65 Harness smoke test — SKIP_AGENT loop.sh iteration, 200 crops, 1 epoch 5e9419c3
The loop discarded this — no improvement under the metric then in force.
other 1.406 28 s DISCARDED
2.64 Harness smoke test — SKIP_AGENT loop.sh iteration, 200 crops, 1 epoch 13657130
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
other 34 s GATED FAIL
2.63 From-scratch trunk + domain-contrastive pretraining + fixed retinex channel + D4 symmetry + top-K-masked decode (forced pivot, mobilenet_v3_small banned) c4616251
The loop discarded this — no improvement under the metric then in force.
architecture 1.563 64 m 52 s $9.41 DISCARDED
2.62 Fire-module SqueezeNet1.1 trunk + FiLM-conditioned single field + D4 symmetry + adaptive-sharpening decode (forced pivot, mobilenet_v3_small banned) 0041192d
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
architecture 26 m 11 s $6.86 GATED FAIL
2.61 Quantization-freed capacity: from-scratch depthwise-separable trunk + unified single-head field, weights trained for int8 export 9421cc39
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
architecture 45 m 05 s $7.88 GATED FAIL
2.60 Corrected exp-38 retry: fixed illumination-invariant channel + from-scratch trunk + cross-lighting consistency loss, gate deleted 4add0860
The loop discarded this — no improvement under the metric then in force.
architecture 1.681 86 m 54 s $6.23 DISCARDED
2.59 Independently-trained day/night specialist twins with a trunk-free brightness dispatcher efa8f69b
The loop discarded this — no improvement under the metric then in force.
architecture 1.581 80 m 09 s $6.15 DISCARDED
2.55 Dense scene-coordinate voting with a differentiable IRLS robust consensus (from-scratch trunk) c998b079
The loop discarded this — no improvement under the metric then in force.
architecture 1.761 21 m 51 s $3.41 DISCARDED
2.54 Memory-safe flying delighting translator + from-scratch trunk: the never-fairly-tested canonicalization pivot, corrected d06f7d8d
The loop discarded this — no improvement under the metric then in force.
relighting 1.496 19 m 42 s $2.88 DISCARDED
2.53 ShuffleNetV2 trunk with FiLM-conditioned single field and top-K decode replaces the two-expert MobileNet blend 01ab7892
The loop discarded this — no improvement under the metric then in force.
architecture 1.733 21 m 36 s $2.96 DISCARDED
2.52 Continuous Gaussian-mixture location regression replaces the shared map-cell classification field 5361f941
The loop discarded this — no improvement under the metric then in force.
architecture 1.476 21 m 32 s $3.90 DISCARDED
2.51 Residual vector-quantized location codebook: hard nearest-neighbor coding replaces the shared-softmax cell field 4752c6c2
The loop discarded this — no improvement under the metric then in force.
quantization 1.646 28 m 14 s $3.73 DISCARDED
2.50 Spatial mixture-of-experts: a coarse macro-region classifier dispatches to disjoint local fine-cell experts e4a5f33d
The loop discarded this — no improvement under the metric then in force.
architecture 1.408 64 m 50 s $4.72 DISCARDED
2.49 Dense per-patch coordinate consensus (re-run after ENOSPC): from-scratch trunk, no map-cell field, error-driven sampler, disagreement-based confidence 7d7bea5d
The loop discarded this — no improvement under the metric then in force.
architecture 1.311 54 m 57 s $3.52 DISCARDED
2.48 Dense per-patch coordinate consensus (corrected re-attempt): from-scratch trunk, no map-cell field, error-driven sampler, disagreement-based confidence aedcce64
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
architecture 14 m 08 s $3.25 GATED FAIL
2.47 Dense per-patch coordinate regression with learned-consensus pooling replaces the map-cell classification field aedcce64
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
architecture 7 m 27 s $0.09 GATED FAIL
2.46 Two-expert lighting dispatcher with from-scratch trunks and factorized row/column coordinate fields f42a8309
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
architecture 40 m 55 s $3.09 GATED FAIL
2.45 Retrieve-then-rerank: a coarse field proposes a shortlist, a pairwise comparator decides among it 0d258f6f
The loop discarded this — no improvement under the metric then in force.
architecture 1.591 60 m 53 s $4.24 DISCARDED
2.44 Sequential delighting translator + from-scratch trunk: canonicalize the pixels first, localize second f61b14fa
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
relighting 75 m 59 s $5.09 GATED FAIL
2.43 Joint reconstruction-canonicalization trunk: a from-scratch encoder that must repaint the clean daytime crop, trained alongside localization every step 498f9a89
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
architecture 77 m 03 s $8.41 GATED FAIL
2.42 Domain-native contrastive pretraining replaces the ImageNet trunk (corrected re-attempt) fed7c60a
The loop discarded this — no improvement under the metric then in force.
architecture 1.836 69 m 24 s $4.13 DISCARDED
2.41 Domain-native contrastive pretraining replaces the ImageNet trunk c0fda62e
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
architecture 15 m 38 s $0.18 GATED FAIL
2.40 Coarse-to-fine geolocalization: a shared conv field ranks the region, a shared local head regresses the offset 9880938d
The loop discarded this — no improvement under the metric then in force.
architecture 1.738 81 m 04 s $8.40 DISCARDED
2.39 Coordinate-generated map field: replace 1,024 free-parameter cell templates with a shared position-conditioned embedding generator 898d8cb5
The loop discarded this — no improvement under the metric then in force.
architecture 1.503 71 m 34 s $7.28 DISCARDED
2.38 Illumination-invariant retinex channel joins raw RGB from the pixel level through the trunk 898d8cb5
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
architecture 20 m 32 s $0.24 GATED FAIL
2.37 Multi-hypothesis coordinate regression replaces the 1024-cell field 5a6771be
The loop discarded this — no improvement under the metric then in force.
architecture 1.643 66 m 02 s $4.59 DISCARDED
2.36 Hardest-impostor margin hinge, retested on the confusability-weighted champion 3ee64311
The loop discarded this — no improvement under the metric then in force.
loss 1.601 59 m 54 s $3.14 DISCARDED
2.35 Confusability-weighted location sampling: oversample look-alike cell pairs, not just more places 3750ddb6
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.588 — worse than silence.
training 1.588 62 m 59 s $4.20 KEPT
2.34 Hardest-impostor margin loss: the field must outrank its best lookalike, not just light up the truth bfbca30d
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
loss 20 m 01 s $8.03 GATED FAIL
2.33 Sensor-sharpness nuisance randomization: per-crop PSF/resampling jitter breaks the blur-texture shortcut 0aaab0f5
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
augmentation 59 m 18 s $7.04 GATED FAIL
2.32 Risk-controlled abstention: six per-lighting-regime thresholds calibrated on never-trained fenced blocks 4871a8c7
The loop discarded this — no improvement under the metric then in force.
training 1.438 108 m 15 s $19.07 DISCARDED
2.31 Neural-map correlation field: localization becomes sliding-window matching of a crop fingerprint against a learned 32 m map raster 1ee201ff
The loop discarded this — no improvement under the metric then in force.
architecture 1.771 73 m 46 s $12.11 DISCARDED
2.30 Supervised lighting dispatcher: day/dusk/night specialist field heads replace the two-expert blend, int8-paid 92404766
The loop discarded this — no improvement under the metric then in force.
architecture 1.696 73 m 20 s $11.12 DISCARDED
2.29 Stride-8 fine-layout tap: the field scorer reads an 8-m-resolution mid-trunk snapshot, int8-paid 7e6fb3ea
The loop discarded this — no improvement under the metric then in force.
architecture 1.618 70 m 11 s $11.05 DISCARDED
2.28 Rerun, never scored: unseen-ground confidence calibration + reinstated 3x trunk (iter 7 died on a full disk before training) 83bf925a
The loop discarded this — no improvement under the metric then in force.
training 1.568 57 m 43 s $5.51 DISCARDED
2.27 Unseen-ground confidence calibration unlocks the reinstated 3×-capacity trunk cc3faa2c
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
training 14 m 55 s $6.71 GATED FAIL
2.26 Deployment-envelope capacity scaling: 3× pretrained trunk depth paid for by int8 FC-head storage 4fb65e2a
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
architecture 69 m 41 s $10.18 GATED FAIL
2.25 Convergence-scaled training: 3× optimizer steps with cosine LR decay under the kept fresh-draw sampler 7f8a2986
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.623 — worse than silence.
training 1.623 63 m 08 s $9.15 KEPT
2.24 Daytime-redraw auxiliary, rerun: exp 23 was OOM-killed mid-harness, never scored — retest it with a memory-lean epoch sampler 5794e840
The loop discarded this — no improvement under the metric then in force.
loss 1.678 42 m 59 s $7.56 DISCARDED
2.23 Daytime-redraw auxiliary: a throwaway decoder must reconstruct the clean daytime crop from the encoder's features adfc8b91
Failed a deployment gate: too large, too slow, or it abstained on too many frames to be scoreable.
loss 42 m 08 s $6.88 GATED FAIL
2.22 Train-only per-patch place supervision: every feature cell places its own patch, the flying decode is untouched 9df5c839
The loop discarded this — no improvement under the metric then in force.
loss 1.696 40 m 44 s $8.67 DISCARDED
2.21 Periodic blind holdout check (hamburg) cac6ec61
Blind Hamburg holdout check. Logged for honesty, never allowed to drive keep/revert.
other HOLDOUT
2.20 Per-epoch training-set resampling: a fresh 6,000-location draw per bucket every epoch replaces the one frozen 36k-crop tensor cac6ec61
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.683 — worse than silence.
training 1.683 35 m 33 s $7.80 KEPT
2.19 Off-site distractor patching: half the training crops carry a pasted block of terrain from elsewhere, labels unchanged d5080d5d
The loop discarded this — no improvement under the metric then in force.
augmentation 1.633 42 m 51 s $9.87 DISCARDED
2.18 Dense per-patch field voting: 64 local place votes with learned fusion replace the global 1024-way template head 48e17da6
The loop discarded this — no improvement under the metric then in force.
architecture 1.623 45 m 05 s $10.24 DISCARDED
2.17 Nuisance-randomized training renders: each bucket's training crops drawn from three seeded realizations of the frozen relighting sim 95b4c3a0
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.753 — worse than silence.
relighting 1.753 39 m 10 s $8.60 KEPT
2.16 C4 rotation-vote field: average field logits over the crop's four 90° turns 5c5373df
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.701 — worse than silence.
architecture 1.701 36 m 53 s $7.58 KEPT
2.15 Selective prediction: field-shape confidence head with per-bucket calibrated abstention f3b87247
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.681 — worse than silence.
architecture 1.681 40 m 01 s $11.51 KEPT
2.14 Peak-commit decode: β-sharpened softmax replaces mean-of-map soft-argmax d2a0a2b9
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.971 — worse than silence.
architecture 1.971 31 m 13 s $7.30 KEPT
2.13 Periodic blind holdout check (hamburg) fe745eb4
Blind Hamburg holdout check. Logged for honesty, never allowed to drive keep/revert.
other HOLDOUT
2.12 Luminance-gated dark-expert field head blended with the existing layout head fe745eb4
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.996 — worse than silence.
architecture 1.996 42 m 10 s $4.95 KEPT
2.11 ImageNet-pretrained MobileNetV3-Small trunk replaces the from-scratch encoder d69862d9
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.986 — worse than silence.
architecture 1.986 31 m 35 s $2.64 KEPT
2.10 Layout-aware field head: squeezed 8×8 spatial code ⊕ GAP replaces GAP-only head input ff271753
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.996 — worse than silence.
architecture 1.996 38 m 08 s $4.27 KEPT
2.9 Cross-lighting contrastive pairs: NT-Xent metric learning on the place descriptor 76c45e3c
The loop discarded this — no improvement under the metric then in force.
loss 1.996 23 m 41 s $4.93 DISCARDED
2.8 Deployment-envelope residual encoder: 4.2× capacity (232k → 973k params), 11 conv layers with skips 76c45e3c
The loop discarded this — no improvement under the metric then in force.
architecture 1.961 48 m 45 s $5.33 DISCARDED
2.7 Scale training coverage 7.5x: 6,000 of ~45k available train locations per bucket (was 800) 76c45e3c
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.986 — worse than silence.
training 1.986 21 m 44 s $3.83 KEPT
2.6 Per-epoch rotation resampling replaces the one-shot frozen training tensor de614d47
The loop discarded this — no improvement under the metric then in force.
augmentation 1.981 6 m 52 s $1.67 DISCARDED
2.5 Hierarchical coarse-to-fine supervision of the probability field via probability pooling 7d02e02d
The loop discarded this — no improvement under the metric then in force.
loss 1.991 6 m 20 s $1.91 DISCARDED
2.4 ACE-style dense per-patch scene-coordinate regression replaces global map-cell probability field 7d02e02d
The loop discarded this — no improvement under the metric then in force.
architecture 2.001 9 m 11 s $2.92 DISCARDED
2.3 Argmax-anchored local soft-argmax decode replaces global expected-coordinate ff65590e
The loop discarded this — no improvement under the metric then in force.
architecture 1.971 4 m 58 s $1.25 DISCARDED
2.2 DSNT-style spatial probability field over the map replaces direct (u,v) regression ff65590e
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 2.001 — worse than silence.
architecture 2.001 7 m 08 s $2.06 KEPT
2.1 Starting baseline — naive TinyLocNet, from-scratch, frozen pipeline v1 b47d44a8
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.991 — worse than silence.
architecture 1.991 KEPT
1.7 Relight v1 — ambient-compensated auto-exposure + golden/blue-hour grading (eval-set reset) 678dd1e7
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.991 — worse than silence.
relighting 1.991 KEPT
1.6 Per-epoch crop resampling — train on fresh positions/rotations every epoch instead of a fixed 4,800-crop snapshot 4959f8fb
The loop discarded this — no improvement under the metric then in force.
training 1.991 DISCARDED
1.5 Grid-cell classification + within-cell offset regression replaces direct (u,v) MSE regression 7b7b491a
The loop discarded this — no improvement under the metric then in force.
loss 1.976 DISCARDED
1.4 Split rebalance — 360 m blocks so the no-leakage buffer stops eating the train area 2e5b3020
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 1.996 — worse than silence.
other 1.996 KEPT
1.3 Data v2 baseline — same TinyLocNet on 1 m/px DOP imagery, all 4 dev areas c0a1bdfe
The loop kept this — it beat the running best under the metric then in force. On today's ruler it scores 2.001 — worse than silence.
other 2.001 KEPT
1.2 Harness smoke test (SKIP_AGENT loop.sh iteration, 1 epoch) 0687f20e
The loop kept this — it beat the running best under the metric then in force. It cannot be placed on today's ruler at all — see the row for why.
other NO SCORE
1.1 Bootstrap baseline — naive TinyLocNet direct-coordinate regressor 15b43a25
The loop kept this — it beat the running best under the metric then in force. It cannot be placed on today's ruler at all — see the row for why.
architecture NO SCORE
What do the columns and marks mean?
Mission score
The single number the loop optimizes, and deliberately the product requirement rather than a statistic about errors. The aircraft takes a vision fix every 5–10 s and it is its only drift correction, so per frame exactly three things can happen: it is confident and within 100 m (a usable fix — the product), not confident (it abstains — safe, it waits), or confident and wrong (a false fix — dangerous, it feeds a wrong position into navigation). The score is (1 − usable-fix rate) + false-fix rate: 0 means every frame is a usable fix, 1.0 means it abstains on everything, 2.0 means it is confidently wrong on everything — strictly worse than silence, which is the point. Error statistics (median, geometric mean, p10, p25) are all still recorded, but none is optimized: each was tried as the target and each rewarded something the aircraft does not want.
Region-holdout (diagnostic)
A small 1-in-32-block slice of Berlin is kept genuinely out of training and scored separately. It is logged and never optimized — it cannot move keep/revert. Its only job is to show whether the model has real spatial structure or is just a lookup table over memorized vantages, so expect it to read far worse than the headline number. That gap is information, not failure.
gated fail (×)
The §6 score also enforces the aircraft's hard limits: the exported model must fit the ESP32-P4 flight computer (≤ 4 MiB) and answer within the latency budget (≤ 250 ms host proxy), and it may not dodge hard cases by refusing to answer (a cell where it abstains on > 80% of frames counts as failed). Any violation scores the whole experiment as failed regardless of accuracy.
rejected
Never trained at all. After a losing streak, the loop demands a genuine pivot — every non-frozen part of the design must change, not just the one piece that's gone stalest. If the proposed design doesn't clear that bar (checked against the actual code diff, not just what the agent claims it changed), it's rejected before spending any GPU time on it.
cov (coverage)
Share of test frames the model was confident enough to answer at all. Abstaining honestly on bad frames is allowed — down to the 20% floor above.
Kept / Discarded
Karpathy-style loop discipline: branch from the best code, run one focused experiment, keep the change (git commit) only if the worst-case error improves — otherwise revert it. The step line in the chart is the running best.
Eval-set reset (┆)
During the bootstrap phase the frozen evaluation data itself was revised twice (10 m satellite → 1 m orthophotos; then a split rebalance). Scores on different eval sets are measurements on different rulers and must not be compared — the dashed vertical rule marks the break, and the running-best line restarts there instead of pretending continuity. From Phase 2 on, the eval set does not change.
Holdout (○)
Hamburg is the blind fifth area: structurally different (port, river, spread-out), never seen by the loop, scored only as a periodic read-only check. If its error diverges from the four development areas, the pipeline has learned their quirks rather than a general method. Its result never influences keep/revert.
Pivot-directed (▽)
After 4 consecutive experiments in a row fail to beat the running best, the harness injects a mandatory pivot preamble into the next design prompt: do not refine the champion's current mechanism again — propose from a design family absent from the recent history. Marked ones ran under that directive; the streak resets every time an experiment is kept.
Cost
Estimated cloud-GPU spend for that one experiment: its measured wall time × a $0.69/hr rate, billed continuously by the clock (not per phase), so it covers the whole experiment — design, implementation, training, scoring, publishing — not GPU time alone. Shown only for the middle, rented-GPU window; the bootstrap and later local eras run at ~$0 marginal compute and show —.
Category
Which lever the experiment pulls: architecture, loss, augmentation, relighting, training procedure, or quantization.
Where the deployment numbers went
Exported model size, weight initialization and single-frame latency each used to have a column. All three were dropped: nearly every row read the same, and the size and latency gates are already folded into the mission score — a model that misses either scores as a gated fail, so the column was restating the Status column. All three remain in the expanded record, and init also on the model designs page.
Training set
Each area offers ~45,000 distinct training positions (1 m/px, 128 m crops, times random rotation); how many crops an experiment actually samples per lighting condition is its own choice and is shown in the detail view. Training crops never overlap the eval blocks — enforced by the frozen split, not by convention.
Agent prompt
Every loop experiment records the exact prompt its headless agent received — expandable in the detail view, so any experiment can be re-run or audited later.
In plain words
Each experiment's own jargon-free explanation of what it tried, pre-registered alongside the technical design — read this first if the title looks like alphabet soup.
One real test
The figure at the top of each detail view is not an illustration: the same held-out Berlin viewpoint is fed to that experiment's actual exported model, and the red glow on the map is the probability field recovered from the deployed ONNX artifact itself — with the true location (○), the model's answer (×), and the real miss distance. Because every experiment sees the identical crop, any difference between two figures is the mechanism change, not the example. The pipeline sentence beneath names the stages; red marks what that experiment changed (a “+” stage is training-only and never flies).
Click any row — or any point in the chart — to see the experiment's pre-registered hypothesis, method and expected outcome, the measured result, the per-area × lighting scoreboard, and what the model actually looked at.

Every experiment the project has run, across 5 evaluation eras. The lineage was wiped at each era boundary, so the numbering restarts inside every band — but the vertical scale is one ruler throughout. 53 of these were re‑measured for this chart by running their exported model through today’s frozen scorer on today’s held‑out viewpoints; 4 are native to the current era; 4 had their run artifacts deleted and have their usable- and false-fix rates recovered arithmetically from the logged record; 17 failed a deployment gate and had no working model to score, then or now. Nothing was converted by a fudge factor. Black means the loop kept it — judged at the time, under whatever metric was then in force, which is why kept experiments sit high in the early bands: the old metric was rewarding the wrong thing. The step line is the best mission score anyone had actually achieved, and it is unbroken across the boundaries because the ruler no longer changes at them.

scroll to zoom · drag to pan · double-click to reset
view on GitHub