ARC-AGI-3 asks an agent to work out how an unfamiliar 2D game works without being told. GPT-5.6 Sol scored 7.8% on it, and OpenAI went looking for why.
In the News: August 1, 2026, Evening
OpenAI found two harness settings tripled GPT-5.6 Sol's ARC-AGI-3 score and cut its output tokens sixfold.
Evening edition
Machine-readable
Download Markdown
Story
Two API settings tripled a frontier model's score on a benchmark it was failing
Read story →
Story
Six Claude Code changelog entries, most of them about what an agent may do to your machine
The changelog section spanning versions 2.1.215 to 2.1.220 focused largely on containment. The costliest single line is in 2.1.219: Claude Opus 5 became the default Opus model, with a 1M context window and fast mode at $10 and $50 per…
Story
A harness claimed 99% on ARC-AGI-3. The verified board's best number is 30.2%
On 16 July an anonymous team published "Frontier Models with Our Harness Achieve ~99% on ARC-AGI-3 Public," reporting 98.98% against a 42.83% Claude Code baseline on the same models. The Dark Factory sweep read that essay end to end on 17…