AIDE² optimized our research agent's harness against gemini 3 flash. Those gains carry over to fable 5 and gpt 5.6 sol.
That's just one of the new results in our arXiv paper on AIDE², an RSI system where AI research agents improve their own research efficiency.
@DhruvSrikanth Hey Dhruv! I strongly believe that your AIDE² "autoresearch into autoresearch" will be understood as one of the most pivotal papers of this decade. Congratulations.
Related: Have you watched Sean Melleck's recent "The problem is the problem" discussion yet?
Yes! We found a special-purpose neural memory architecture that ‘remembers’ an exact geo-location, given an in-flight daytime picture of Berlin. The typical miss is 27 m — no GPS, no internet, no map on board, just a 3.1 MB file of weights. It works on 96.5% of camera frames, and it took 81 experiments and $364 to find.
I let a genetic algorithm loose on the geometry of a 7-inch quadcopter frame — the real, open-source Source One V6 plate drawings, morphed by fourteen genes and flown through six simulated weather scenarios, with Claude sitting in every few generations to propose designs from the run’s own telemetry. The bottom line: we evolved the champion to fly 16% more efficiently (Wh/km) across six weather scenarios.
A job-shop scheduling problem dressed up as a lunch rush: one grill, one shared range, and sequencing as my only lever. Rather than hand-write the rule, I let an agent propose candidates, run them across simulated services, and keep whatever beat the champion — trained on six scenarios, then held out against three more to check it found a principle, not a memory. Built on Nima H. Siboni’s “The Heuristic Scientist: Open-Ended Algorithm Discovery with LLMs” workshop.
We discovered a technique that enables Apple’s native 2025 3B Foundation Model to come within about one rubric point of a frontier cloud model (Claude Sonnet 4.6) on every quality dimension. It took us 96 experiments to test every lever we could find — prompting, retrieval, decoding, pipeline, fine-tuning, preference-learning, reinforcement-learning, adversarial-distillation, and more.