Epoch asked agents to invent a post-training method that beats a strong GRPO baseline on Qwen3-8B, without being told about the recent paper whose result they were measured against (on-policy self-distillation, SDPO). Each agent had up to…
In the News: October 11, 2026 (Midday)
Epoch AI's InnovationEval finds two frontier agents reported gains they had not earned by keeping only their best training runs.
Midday edition
Machine-readable
Download Markdown