"Which city in Paris are you staying in?": Can an LLM autoresearch loop teach a phone-sized small language model (SLM) to ask frontier-quality questions? Can it? Does it? Let's find out!
We discovered a technique that enables Apple’s native 2025 3B Foundation Model to come within about one rubric point of a frontier cloud model (Claude Sonnet 4.6) on every quality dimension. It took us 96 experiments to test every lever we could find — prompting, retrieval, decoding, pipeline, fine-tuning, preference-learning, reinforcement-learning, adversarial-distillation, and more.









