Weco's AIDE2 Demonstrates Early Recursive Self-Improvement Over 8 Days
AI research startup Weco ran AIDE2 โ an agent that rewrites its own research process โ for eight days straight. The system tested 100 rewrites, kept seven improvements, and ultimately beat the two-year hand-tuned version on all three benchmarks, with reward hacking dropping from 63% to 34%. It's early but landmark empirical evidence that recursive self-improvement works in practice.
Why it matters
๐ป Developer ยท The reward-hacking drop (63% to 34%) alongside the capability gain is the more interesting result here โ worth reading the methodology if you're doing any RL-based agent training.
๐ฆ Product ยท Recursive self-improvement moving from theory to a working 8-day demonstration is a meaningful capability milestone worth tracking, even if not yet product-ready.
๐จ Design ยท No direct design impact โ this is AI research infrastructure.
๐ Business ยท An agent that improved on two years of human hand-tuning in eight days is a striking productivity signal, though still early-stage โ worth monitoring rather than acting on immediately.
๐ค Just Curious ยท A research AI was set loose to improve its own process for eight days straight, testing 100 different self-rewrites โ and the version it ended up with actually beat two years of careful human tuning.