In reply to @SpringStreetNYC · May 29, 2026I'm unironically on a quest to figure out how to get local Apple's Foundation Models to respond at the quality of recent Sonnet across five rubriks that I care for.
Alright!
The autoresearch agent came up with an unexpected way to push quality to from 40% to 44%:
"First, generate one more option than we actually need, then throw away the one we can most afford to lose, keeping the rest."
Suprises:
1. It's much simpler than the LoRA adapter route
2. It produced four responses that were significantly better than their eval targets
3. I hit an Anthropic API usage limit for the first time (Yay!)
And, man, I have to say: As an experimenter autoresearch is really, really fun to work with, guide and learn from. I wish more people would have this experience in their domains.

