The autoresearch agent came up with an unexpected way to push quality to from 40% to 44%:
"First, generate one more option than we actually need, then throw away the one we can most afford to lose, keeping the rest."
Suprises:
1. It's much simpler than the LoRA adapter route
2. It produced four responses that were significantly better than their eval targets
3. I hit an Anthropic API usage limit for the first time (Yay!)
And, man, I have to say: As an experimenter autoresearch is really, really fun to work with, guide and learn from. I wish more people would have this experience in their domains.
Ah, scratch the "Doesn't use LoRA" surprise. Experiment 056 is built ON-TOP of 046. I forgot to make the diagram-generator aware of this kind of experiment heritage.
I'm unironically on a quest to figure out how to get local Apple's Foundation Models to respond at the quality of recent Sonnet across five rubriks that I care for.
Alright!
The autoresearch agent came up with an unexpected way to push quality to from 40% to 44%:
"First, generate one more option than we actually need, then throw away the one we can most afford to lose, keeping the rest."
Suprises:
1. It's much simpler than the LoRA adapter route
2. It produced four responses that were significantly better than their eval targets
3. I hit an Anthropic API usage limit for the first time (Yay!)
And, man, I have to say: As an experimenter autoresearch is really, really fun to work with, guide and learn from. I wish more people would have this experience in their domains.
I'm unironically on a quest to figure out how to get local Apple's Foundation Models to respond at the quality of recent Sonnet across five rubriks that I care for.
But basically I'm using frontier models to train a local ("homestead"?) model as to impart some of that good brain stuff but minus the overwhelm. Like a teacher, hopefully.
I'm unironically on a quest to figure out how to get local Apple's Foundation Models to respond at the quality of recent Sonnet across five rubriks that I care for.
LOL, yeah, this one also not a keeper.
An old friend of mine shared this nugget of wisdom from art school with me:
I'm unironically on a quest to figure out how to get local Apple's Foundation Models to respond at the quality of recent Sonnet across five rubriks that I care for.
Okay! Adding a domain-specific LoRA adapter moved the needle the most.
I bootstrapped an autoresearch loop that initially went through a bunch of experiments that tested a bunch of prompt/sequence/selection/critique and reflexion combos. Those didn't really result in much lift.
PS: If you haven't seen this yet, Apple provides the "Foundation Models Adapter Training Toolkit" over at
I'm unironically on a quest to figure out how to get local Apple's Foundation Models to respond at the quality of recent Sonnet across five rubriks that I care for.
Infermation (/ɪnfərˈmeɪʃən/; portmanteau of inference and information) is knowledge produced by inferential processes rather than direct observation, measurement, or first-hand transmission.
Unlike conventional information, which is grounded in a source, infermation is derived — synthesized from patterns, priors, and partial evidence.
Okay, with a bit of fiddling yesterday, I've now got a serious day-to-day contender to Claude Code. Pi(.)dev + unsloth/qwen3.6-35b-a3b-ud-mlx served via @lmstudio.
The video below is not sped up. It shows the agent wrapping up bootstrapping and running a small suite of UX tests, making a commit and doing some house cleaning.
Honestly: It's fast. It's "Good Enough™️". And most importantly: It's predictable.
After spending most of March and April outright fighting with Opus every single day, I'll take predictable over powerful any day now.
And I want to point out: Running it requires 20G of RAM. A clean 20G. Not less. But also NOT more.
@lmstudio PS: Thank you @badlogicgames and the team for building Pi. Feels good to work with a clean, lean agent harness again.
Okay, with a bit of fiddling yesterday, I've now got a serious day-to-day contender to Claude Code. Pi(.)dev + unsloth/qwen3.6-35b-a3b-ud-mlx served via @lmstudio.
The video below is not sped up. It shows the agent wrapping up bootstrapping and running a small suite of UX tests, making a commit and doing some house cleaning.
Honestly: It's fast. It's "Good Enough™️". And most importantly: It's predictable.
After spending most of March and April outright fighting with Opus every single day, I'll take predictable over powerful any day now.
And I want to point out: Running it requires 20G of RAM. A clean 20G. Not less. But also NOT more.