Alexis Rondeau

#discovery-loops

7 dispatches tagged discovery-loops

26 views

Oh perfect! Thank you for sharing the slides as well. I watched your talk on youtube yesterday and have been telling all my IRL friends about it :)

I love that your system is explicitly designed to hand back the recommendation to a human, not yet another loop.

  • ProbXiv keeps throwing errors for the past hour or so. Any idea when it'll be back up?
  • What do you think of Zhengyao Jiang and team's new AIDE² paper?
Quoting @zhengyaojiang · Sep 23, 2026

We're releasing the arXiv paper on AIDE².

The RSI system where AI research agents improve their own research efficiency.

It includes new results on transfer across models and comparisons with more AI research agents: [1/4] t.co/P8pfOiiO55 t.co/EHeHtqS4mi

Image from the post
2 views
In reply to @DhruvSrikanth · Sep 23, 2026

Excited to see this paper out!

AIDE² optimized our research agent's harness against gemini 3 flash. Those gains carry over to fable 5 and gpt 5.6 sol.

That's just one of the new results in our arXiv paper on AIDE², an RSI system where AI research agents improve their own research efficiency.

@DhruvSrikanth Hey Dhruv! I strongly believe that your AIDE² "autoresearch into autoresearch" will be understood as one of the most pivotal papers of this decade. Congratulations.

Related: Have you watched Sean Melleck's recent "The problem is the problem" discussion yet?

"Not all who wander are lost": Can a UAV learn a city by heart — no GPS, no map on board, just a $4 flight computer?

Yes! We found a special-purpose neural memory architecture that ‘remembers’ an exact geo-location, given an in-flight daytime picture of Berlin. The typical miss is 27 m — no GPS, no internet, no map on board, just a 3.1 MB file of weights. It works on 96.5% of camera frames, and it took 81 experiments and $364 to find.

Open →

"The snuggle is real": Evolving quadcopter frame geometry for Wh/km, with Claude as an occasional co-designer

I let a genetic algorithm loose on the geometry of a 7-inch quadcopter frame — the real, open-source Source One V6 plate drawings, morphed by fourteen genes and flown through six simulated weather scenarios, with Claude sitting in every few generations to propose designs from the run’s own telemetry. The bottom line: we evolved the champion to fly 16% more efficiently (Wh/km) across six weather scenarios.

Open →

"We don't write the rules. We just discover them. With AI.": How an autonomous agent searches for the groundrule, and what makes the loop different

A job-shop scheduling problem dressed up as a lunch rush: one grill, one shared range, and sequencing as my only lever. Rather than hand-write the rule, I let an agent propose candidates, run them across simulated services, and keep whatever beat the champion — trained on six scenarios, then held out against three more to check it found a principle, not a memory. Built on Nima H. Siboni’s “The Heuristic Scientist: Open-Ended Algorithm Discovery with LLMs” workshop.

Open →

"Which city in Paris are you staying in?": Can an LLM autoresearch loop teach a phone-sized small language model (SLM) to ask frontier-quality questions? Can it? Does it? Let's find out!

We discovered a technique that enables Apple’s native 2025 3B Foundation Model to come within about one rubric point of a frontier cloud model (Claude Sonnet 4.6) on every quality dimension. It took us 96 experiments to test every lever we could find — prompting, retrieval, decoding, pipeline, fine-tuning, preference-learning, reinforcement-learning, adversarial-distillation, and more.

Open →