Research evolution: the shape of the search

The experiment lineage shows what the loop tried. This shows what happened to the research itself. Time runs left to right. The dark line is the path that survived; above it are branches that left the trunk — grey and ×-ended where they stopped; everything below ran as the mainline for a while before being replaced — including which machine the work ran on. Hover anything for what happened and why.

the trunk — what the project still does todaytried, then abandoned — the × is where it stoppedran as the mainline, then replaced — the ◯ is the moment it was retiredan insight that merged in — the arrow is where it changed the projectincident — something brokebootstrap4 areas · 6 lighting buckets · region holdoutberlin onlyviewpoint holdoutmission score

scroll sideways → twelve days, 20 July – 1 August

bootstrap4 areas · 6 lighting buckets · region holdoutberlin onlyviewpoint holdoutmission score21 Jul22 Jul23 Jul24 Jul25 Jul26 Jul27 Jul28 Jul29 Jul30 Jul31 Jul1 Augsix days of silencemain — unchanged sinceSentinel-2 10 m/pxdropped in 20 minPrompt-only pivot rulesnever worked3D shadow relightingparkedRetrieval-index probeparked, not disprovenPrignitz probe · 0.113era opened, then closedRegion-holdout evaluation→ held-out viewpointsMedian error as the score→ mission scoreRunPod - rented 4090 pod→ ModalModal A100 - serverlesscredits out → local M1Berlin-only scope cut→ 4 areas restoredGeometric mean→ mission scorepod disk fullFable hits its capCI out of disktwo writers on mainpod won't resumeloop dies unnoticedcredits exhaustedenforce it in the source, not the promptthe problem is two writers, not the platformstop guessing, profile itit is never shown the places it is tested onthe score must BE the product requirementit was starved, not refutedBootstrap1 m/px orthophotosRelighting rebuiltGoes publicPivot enforced in codeOne git writerPNG decode fixberlin-slim branch0.040 - 96.5% usableback to berlin-slim

The two long red lines are the story: for ten of the project’s twelve days the evaluation was asking a question the model could not answer, and the score preferred guessing to learning. Every result measured against them had to be thrown away.

view on GitHub