🔬MIT and Sakana's SIFT Cuts Cost of Self-Improving Agents
TL;DR
MIT and Sakana AI's SIFT uses an LLM judge to rank candidate coding agents pairwise before expensive benchmark runs. It hit 35.1% on Polyglot with o3-mini in under five hours for about $150.
MIT and Sakana AI's SIFT uses an LLM judge to rank candidate coding agents pairwise before expensive benchmark runs. It hit 35.1% on Polyglot with o3-mini in under five hours for about $150.

Key Points
Beats Darwin Gödel Machine on Polyglot: 35.1% vs 30.7%; 29.8% without the judge
Judge-selected agents averaged 50.4% and 53.8% on a 60-task SWE-bench Verified subset vs 40.0% baseline
Judge pick averaged 36.7% on TerminalBench vs 28.1% for the top search-set scorer
Costs: about 12 cents per patch, 4.4 cents per pairwise judge call
Only tested on coding agents; authors had to block patches that weakened the eval harness
Why It Matters
Evaluation is the bottleneck in self-improving agents, and a cheap pairwise judge is a reusable trick. The harness-tampering finding is a warning for anyone running agent loops.
Quick Facts
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,556 builders reading daily.