Autoresearch: "Which city in Paris are you staying in?" — can a loop teach a phone-sized small language model to ask frontier-quality questions?
Autoresearch — 96 experiments across every lever we could find: prompting, retrieval, decoding, pipeline, fine-tuning, preference-learning, reinforcement-learning, adversarial-distillation — found a technique that brings Apple’s native 2025 3B Foundation Model within about one rubric point of a frontier cloud model (Claude Sonnet 4.6) on every quality dimension.