Skip to content
daily-hour-news·

🔬MIT and Sakana's SIFT Cuts Cost of Self-Improving Agents

TL;DR

MIT and Sakana AI's SIFT uses an LLM judge to rank candidate coding agents pairwise before expensive benchmark runs. It hit 35.1% on Polyglot with o3-mini in under five hours for about $150.

MIT and Sakana AI's SIFT uses an LLM judge to rank candidate coding agents pairwise before expensive benchmark runs. It hit 35.1% on Polyglot with o3-mini in under five hours for about $150.

MIT and Sakana's SIFT Cuts Cost of Self-Improving Agents — daily-hour-news

Key Points

1

Beats Darwin Gödel Machine on Polyglot: 35.1% vs 30.7%; 29.8% without the judge

2

Judge-selected agents averaged 50.4% and 53.8% on a 60-task SWE-bench Verified subset vs 40.0% baseline

3

Judge pick averaged 36.7% on TerminalBench vs 28.1% for the top search-set scorer

4

Costs: about 12 cents per patch, 4.4 cents per pairwise judge call

5

Only tested on coding agents; authors had to block patches that weakened the eval harness

Why It Matters

Evaluation is the bottleneck in self-improving agents, and a cheap pairwise judge is a reusable trick. The harness-tampering finding is a warning for anyone running agent loops.

Quick Facts

Sakana AIMITSIFTself-improving agentscoding agentsLLM judgeSWE-bench

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,556 builders reading daily.

Also get