The autoresearch agent came up with an unexpected way to push quality to from 40% to 44%:
"First, generate one more option than we actually need, then throw away the one we can most afford to lose, keeping the rest."
Suprises:
1. It's much simpler than the LoRA adapter route
2. It produced four responses that were significantly better than their eval targets
3. I hit an Anthropic API usage limit for the first time (Yay!)
And, man, I have to say: As an experimenter autoresearch is really, really fun to work with, guide and learn from. I wish more people would have this experience in their domains.
Ah, scratch the "Doesn't use LoRA" surprise. Experiment 056 is built ON-TOP of 046. I forgot to make the diagram-generator aware of this kind of experiment heritage.