Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: (1/7)
An inner loop, just like a normal autoresearch agent, optimizing code against an eval.
An outer loop, optimizing the inner-loop agent’s harness code against the inner loop’s average score across different benchmarks. (2/7) 
Including a new search policy, a memory system that compresses prompt by 16x, and a layered defense against reward hacking. (3/7)

They generalize. They beat the agent we hand-tuned for two years, on all three.
Two sit inside its training task families. The farthest sits outside, improving a physics-based weather model.
(4/7) 
This was benchmarked on OOD GPU kernel engineering tasks that suffered from reward hacking.
(5/7)

Its self-improvement efficiency went beyond manual R&D with general AI tools, on held-out benchmarks.
We also tested Level 2, whether the improved inner agent makes a better outer loop. Results are mixed, and we do not claim ignition. (6/7) 
– a breakdown of the discovered algorithms
– the rejected ideas AIDE² tried, covering a surprising share of the search literature
– the dead code it shipped
Also, a huge thank you to everyone who provided feedback on the draft, including @jeankaddour, @MinqiJiang, @morgymcg, @odysseus0z, @rosstaylor90, @OfirPress and many others!





