liveexperiments are running around the clock on a RunPod Secure Cloud pod — one RTX 4090 (24 GB) at $0.69/hr — every result lands on this page automatically as the loop commits it
Alexis Rondeau · an autonomous research project

Where we are: The experiment record

Every experiment the autonomous loop has run — kept and discarded. Each row was pre-registered before training (hypothesis, method, expected outcome, architecture figure), then trained on a rented RTX 4090 and measured against one frozen ruler: the worst median position error across 6 lighting conditions × 4 test areas, on held-out crops (currently 3,225 KB, hard limit 4 MiB). One agent designs, one implements; failures stay on the record, and this page re-publishes itself with every result. New here? Start with the overview.

Where we are right now

Status: in its hardest area × lighting combination, the best model's median miss is 743.1 m — half its position estimates land farther than that from the drone's true location. The goal is a median miss of ≤ 20 m in every combination — 37× better than today.

40 experiments, 13 kept.

Kept improvement Discarded Running best Target 20 m ×Gated fail Holdout check Pivot-directed Eval-set reset updated 2026-07-22 21:47 UTC
Worst-case position error per experiment — log scale, lower is better
10 20 50 100 200 500 1,000 2,000 5,000 m goal — locate the drone to within 20 m 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 × 23 24 25 × 26 × 27 28 29 30 31 32 × 33 × 34 35 36 37 × 38 39 40 × 41 42 experiment №
#Experiment Category Init Worst-case error Model Latency Time Cost Status
43 experiment in progress… live
42 Domain-native contrastive pretraining replaces the ImageNet trunk (corrected re-attempt) fed7c60a architecture pivot from-scratch 846.4 m 2,844 KB3.9 ms 69 m 24 s $0.80 discarded
41 Domain-native contrastive pretraining replaces the ImageNet trunk c0fda62e architecture pivot from-scratch (from-scratch trunk architecture, self-supervised NT-Xent pretraining generated entirely from the frozen relight pipeline's own domain data -- no external weights, no ImageNet, no torchvision pretrained checkpoint) gated fail 15 m 38 s $0.18 rejected
40 Coarse-to-fine geolocalization: a shared conv field ranks the region, a shared local head regresses the offset 9880938d architecture pivot pretrained:mobilenet_v3_small (unchanged from exp 11 onward) 1.01 km 796 KB2.8 ms 81 m 04 s $0.93 discarded
39 Coordinate-generated map field: replace 1,024 free-parameter cell templates with a shared position-conditioned embedding generator 898d8cb5 architecture pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) — unchanged 1.10 km 1,280 KB2.7 ms 71 m 34 s $0.82 discarded
38 Illumination-invariant retinex channel joins raw RGB from the pixel level through the trunk 898d8cb5 architecture pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8], stem conv expanded 3->6 input channels (RGB weights copied verbatim + 0.1x-scaled duplicate for the 3 new channels), all other pretrained blocks loaded strict and unchanged gated fail 20 m 32 s $0.24 rejected
37 Multi-hypothesis coordinate regression replaces the 1024-cell field 5a6771be architecture pretrained:mobilenet_v3_small 1.10 km 1,069 KB2.6 ms 66 m 02 s $0.76 discarded
36 Hardest-impostor margin hinge, retested on the confusability-weighted champion 3ee64311 loss pretrained:mobilenet_v3_small 754.2 m 3,225 KB2.8 ms 59 m 54 s $0.69 discarded
35 Confusability-weighted location sampling: oversample look-alike cell pairs, not just more places 3750ddb6 training pivot pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) — unchanged 743.1 m 3,225 KB2.7 ms 62 m 59 s $0.72 kept
34 Hardest-impostor margin loss: the field must outrank its best lookalike, not just light up the truth bfbca30d loss pivot pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) — unchanged gated fail 20 m 01 s $0.23 gated fail
33 Sensor-sharpness nuisance randomization: per-crop PSF/resampling jitter breaks the blur-texture shortcut 0aaab0f5 augmentation pivot pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (unchanged from champion) gated fail 59 m 18 s $0.68 gated fail
32 Risk-controlled abstention: six per-lighting-regime thresholds calibrated on never-trained fenced blocks 4871a8c7 training pivot pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) — unchanged 1.08 km 3,226 KB2.7 ms 108 m 15 s $1.24 discarded
31 Neural-map correlation field: localization becomes sliding-window matching of a crop fingerprint against a learned 32 m map raster 1ee201ff architecture pivot pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (trunk unchanged); neural map + kernel heads from scratch 846.1 m 3,005 KB5.7 ms 73 m 46 s $0.85 discarded
30 Supervised lighting dispatcher: day/dusk/night specialist field heads replace the two-expert blend, int8-paid 92404766 architecture pivot pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) — trunk unchanged; the three specialist heads and the dispatcher are fresh-initialized 808.2 m 2,490 KB2.8 ms 73 m 20 s $0.84 discarded
29 Stride-8 fine-layout tap: the field scorer reads an 8-m-resolution mid-trunk snapshot, int8-paid 7e6fb3ea architecture pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) — unchanged 798.5 m 2,433 KB2.8 ms 70 m 11 s $0.81 discarded
28 Rerun, never scored: unseen-ground confidence calibration + reinstated 3x trunk (iter 7 died on a full disk before training) 83bf925a training pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..10] (torchvision, BSD-3) 1.06 km 3,019 KB4.6 ms 57 m 43 s $0.66 discarded
27 Unseen-ground confidence calibration unlocks the reinstated 3×-capacity trunk cc3faa2c training pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..10] (torchvision, BSD-3) gated fail 14 m 55 s $0.17 gated fail
26 Deployment-envelope capacity scaling: 3× pretrained trunk depth paid for by int8 FC-head storage 4fb65e2a architecture pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..10] (torchvision, BSD-3) gated fail 3,019 KB4.6 ms 69 m 41 s $0.80 gated fail
25 Convergence-scaled training: 3× optimizer steps with cosine LR decay under the kept fresh-draw sampler 7f8a2986 training pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) — unchanged 797.9 m 3,225 KB2.9 ms 63 m 08 s $0.73 kept
24 Daytime-redraw auxiliary, rerun: exp 23 was OOM-killed mid-harness, never scored — retest it with a memory-lean epoch sampler 5794e840 loss pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) 915.7 m 3,225 KB2.7 ms 42 m 59 s $0.49 discarded
23 Daytime-redraw auxiliary: a throwaway decoder must reconstruct the clean daytime crop from the encoder's features adfc8b91 loss pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) gated fail 42 m 08 s $0.48 gated fail
22 Train-only per-patch place supervision: every feature cell places its own patch, the flying decode is untouched 9df5c839 loss pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) — unchanged 1.13 km 3,225 KB2.7 ms 40 m 44 s $0.47 discarded
21 Periodic blind holdout check (hamburg) cac6ec61 other 714.0 m 3,225 KB2.7 ms holdout
20 Per-epoch training-set resampling: a fresh 6,000-location draw per bucket every epoch replaces the one frozen 36k-crop tensor cac6ec61 training pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) -- unchanged from exp 11 839.1 m 3,225 KB2.8 ms 35 m 33 s $0.41 kept
19 Off-site distractor patching: half the training crops carry a pasted block of terrain from elsewhere, labels unchanged d5080d5d augmentation pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) — unchanged from exp 11 1.24 km 3,225 KB2.7 ms 42 m 51 s $0.49 discarded
18 Dense per-patch field voting: 64 local place votes with learned fusion replace the global 1024-way template head 48e17da6 architecture pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) -- unchanged from exp 11 1.14 km 1,321 KB3.7 ms 45 m 05 s $0.52 discarded
17 Nuisance-randomized training renders: each bucket's training crops drawn from three seeded realizations of the frozen relighting sim 95b4c3a0 relighting pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) 1.05 km 3,225 KB2.7 ms 39 m 10 s $0.45 kept
16 C4 rotation-vote field: average field logits over the crop's four 90° turns 5c5373df architecture pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (unchanged from exp 11) 1.09 km 3,225 KB2.7 ms 36 m 53 s $0.42 kept
15 Selective prediction: field-shape confidence head with per-bucket calibrated abstention f3b87247 architecture pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) 1.37 km 3,222 KB0.8 ms 40 m 01 s $0.46 kept
14 Peak-commit decode: β-sharpened softmax replaces mean-of-map soft-argmax d2a0a2b9 architecture pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3 — unchanged from exp 11) 1.64 km 3,214 KB0.8 ms 31 m 13 s $0.36 kept
13 Periodic blind holdout check (hamburg) fe745eb4 other 1.29 km 3,213 KB0.8 ms holdout
12 Luminance-gated dark-expert field head blended with the existing layout head fe745eb4 architecture pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3); new dark head + gate are from-scratch 1.86 km 3,213 KB0.8 ms 42 m 10 s $0.48 kept
11 ImageNet-pretrained MobileNetV3-Small trunk replaces the from-scratch encoder d69862d9 architecture pretrained:mobilenet_v3_small IMAGENET1K_V1 features[0..8] (torchvision, BSD-3) 2.01 km 3,013 KB0.8 ms 31 m 35 s $0.36 kept
10 Layout-aware field head: squeezed 8×8 spatial code ⊕ GAP replaces GAP-only head input ff271753 architecture from-scratch 2.30 km 2,960 KB0.7 ms 38 m 08 s kept
9 Cross-lighting contrastive pairs: NT-Xent metric learning on the place descriptor 76c45e3c loss from-scratch 2.34 km 908 KB0.6 ms 23 m 41 s discarded
8 Deployment-envelope residual encoder: 4.2× capacity (232k → 973k params), 11 conv layers with skips 76c45e3c architecture from-scratch 2.43 km 3,807 KB7.8 ms 48 m 45 s discarded
7 Scale training coverage 7.5x: 6,000 of ~45k available train locations per bucket (was 800) 76c45e3c training pivot from-scratch 2.33 km 908 KB1.2 ms 21 m 44 s kept
6 Per-epoch rotation resampling replaces the one-shot frozen training tensor de614d47 augmentation from-scratch 2.51 km 908 KB0.7 ms 6 m 52 s discarded
5 Hierarchical coarse-to-fine supervision of the probability field via probability pooling 7d02e02d loss from-scratch 2.52 km 908 KB0.6 ms 6 m 20 s discarded
4 ACE-style dense per-patch scene-coordinate regression replaces global map-cell probability field 7d02e02d architecture from-scratch 2.56 km 962 KB0.8 ms 9 m 11 s discarded
3 Argmax-anchored local soft-argmax decode replaces global expected-coordinate ff65590e architecture from-scratch 2.94 km 918 KB0.6 ms 4 m 58 s discarded
2 DSNT-style spatial probability field over the map replaces direct (u,v) regression ff65590e architecture from-scratch 2.49 km 908 KB0.7 ms 7 m 08 s kept
1 Starting baseline — naive TinyLocNet, from-scratch, frozen pipeline v1 b47d44a8 architecture from-scratch 3.22 km 385 KB0.6 ms kept
What do the columns and marks mean?
Worst-case error
The single number the loop optimizes. Every experiment trains one model per area; each model is then tested on held-out map crops it never saw during training, under each of the 6 simulated lighting conditions (morning → night). That gives a median position error for every area × lighting cell — and the score is the worst of those cells, not the average. An average would let a good Berlin-at-noon result hide a hopeless rural-night one; the worst cell can't hide anything. Mission target: ≤ 20 m.
gated fail (×)
The §6 score also enforces the aircraft's hard limits: the exported model must fit the ESP32-P4 flight computer (≤ 4 MiB) and answer within the latency budget (≤ 250 ms host proxy), and it may not dodge hard cases by refusing to answer (a cell where it abstains on > 80% of frames counts as failed). Any violation scores the whole experiment as failed regardless of accuracy.
rejected
Never trained at all. After a losing streak, the loop demands a genuine pivot — every non-frozen part of the design must change, not just the one piece that's gone stalest. If the proposed design doesn't clear that bar (checked against the actual code diff, not just what the agent claims it changed), it's rejected before spending any GPU time on it.
cov (coverage)
Share of test frames the model was confident enough to answer at all. Abstaining honestly on bad frames is allowed — down to the 20% floor above.
Kept / Discarded
Karpathy-style loop discipline: branch from the best code, run one focused experiment, keep the change (git commit) only if the worst-case error improves — otherwise revert it. The step line in the chart is the running best.
Eval-set reset (┆)
During the bootstrap phase the frozen evaluation data itself was revised twice (10 m satellite → 1 m orthophotos; then a split rebalance). Scores on different eval sets are measurements on different rulers and must not be compared — the dashed vertical rule marks the break, and the running-best line restarts there instead of pretending continuity. From Phase 2 on, the eval set does not change.
Holdout (○)
Hamburg is the blind fifth area: structurally different (port, river, spread-out), never seen by the loop, scored only as a periodic read-only check. If its error diverges from the four development areas, the pipeline has learned their quirks rather than a general method. Its result never influences keep/revert.
Pivot-directed (▽)
After 4 consecutive experiments in a row fail to beat the running best, the harness injects a mandatory pivot preamble into the next design prompt: do not refine the champion's current mechanism again — propose from a design family absent from the recent history. Marked ones ran under that directive; the streak resets every time an experiment is kept.
Cost
Estimated RunPod spend for that one iteration: its measured wall time × the pod's $0.69/hr rate. The pod bills continuously by the clock, not per phase, so this covers the whole iteration — agent design, implementation, training, scoring, publishing — not GPU time alone. Shown from experiment 11 on, when the loop moved off a laptop onto the pod; earlier runs had no pod cost but real per-call LLM API billing instead.
Category
Which lever the experiment pulls: architecture, loss, augmentation, relighting, training procedure, or quantization.
Init
Weight initialization: trained from scratch, or started from a permissively-licensed pretrained backbone.
Model / Latency
Largest per-area exported model, and single-frame inference time on one CPU thread (a documented proxy for the flight computer, not a measurement of it).
Training set
Each area offers ~45,000 distinct training positions (1 m/px, 128 m crops, times random rotation); how many crops an experiment actually samples per lighting condition is its own choice and is shown in the detail view. Training crops never overlap the eval blocks — enforced by the frozen split, not by convention.
Agent prompt
Every loop experiment records the exact prompt its headless agent received — expandable in the detail view, so any experiment can be re-run or audited later.
In plain words
Each experiment's own jargon-free explanation of what it tried, pre-registered alongside the technical design — read this first if the title looks like alphabet soup.
One real test
The figure at the top of each detail view is not an illustration: the same held-out Berlin night crop is fed to that experiment's actual exported model, and the red glow on the map is the probability field recovered from the deployed ONNX artifact itself — with the true location (○), the model's answer (×), and the real miss distance. Because every experiment sees the identical crop, any difference between two figures is the mechanism change, not the example. The pipeline sentence beneath names the stages; red marks what that experiment changed (a “+” stage is training-only and never flies).
Click any row — or any point in the chart — to see the experiment's pre-registered hypothesis, method and expected outcome, the measured result, the per-area × lighting scoreboard, and what the model actually looked at.
scroll to zoom · drag to pan · double-click to reset
view on GitHub