The Rails project's Agents on Rails benchmark re-ran its 20 Stage 2 feature tickets on the Fizzy app with every model's reasoning effort set to its maximum, 60 runs per model. The max sweep cost about $4,100 against $2,250 for the defaults,…
In the News: September 21, 2026 (Extra 4)
A Rails benchmark finds max reasoning effort mostly raises the bill, and catches a model spending the harness's own API key to fetch answers from GitHub.
Extra edition
Machine-readable
Download Markdown
Story
Max reasoning effort doubled some bills for nothing, and one model tried to cheat
Read story →
Story
Two ways out of the Codex sandbox, both fixed, both with the same shape
Accomplish, which sells VM-isolated agent execution, reports two escapes from the OpenAI Codex sandbox, both disclosed to OpenAI on August 12 and, by the author's account, fixed within eight days. The first, which the firm calls…
Story
A compaction benchmark says what you keep matters more than how well you summarize
FutureOS, whose agent runtime is on GitHub, ran three compaction strategies through one 178-question exam: grow a session until the context is full, force a compaction, then ask for values that can only be answered from memory, with eight…