← In the News

SWE-2 scores within one point of Fable 5.1 on Cognition's benchmark

Introducing SWE-2: Pushing the Pareto Frontier · The Cognition Team · Cognition, September 10, 2026

Machine-readable Download Markdown

Cognition released SWE-2, a coding model post-trained from Moonshot's 2.8-trillion-parameter Kimi K3. On the company's FrontierCode 1.1 Main benchmark, SWE-2 scores 50.0 percent. Anthropic's Fable 5.1 scores 50.9 percent, while SWE-2 costs 64 percent less per task.

Cognition credits a reinforcement learning method that trains the medium, high, and max effort levels in one run instead of using separate models. At medium effort, SWE-2 beats SWE-1.7's score on the same benchmark while taking 58 percent fewer turns and costing 81 percent less on average. It reaches its first real code edit after a median of 18 steps, compared with 48 for SWE-1.7.

SWE-2 is available today in Devin Desktop and Devin CLI, with Devin Web and Fusion rolling out. Cognition has not published a standalone API, pricing, or open weights.

Why it matters: Teams comparing their own harness with Devin now have vendor-reported cost and turn-count figures for the model behind Devin's latest release. Cognition measured the accuracy and cost numbers on its own benchmark, so an independent test still needs to confirm the comparison.