Cognition releases SWE-2 coding model, matching rivals at a fraction of the cost

Cognition says its new SWE-2 model scores within a point of Fable 5.1 on a key coding benchmark while costing 64% less to run, and is now live in Devin.

AIVIO News Desk 2 min read
AIVIO News cover illustration for: Cognition releases SWE-2 coding model, matching rivals at a fraction of the cost

Cognition released SWE-2, a new coding model the company says pushes its cost-to-performance tradeoff closer to the top models on the market. The model scored 50.0% on FrontierCode 1.1 Main, a benchmark that tests how well an AI agent solves real software engineering tasks, putting it within one point of Fable 5.1’s 50.9% score while, according to Cognition, costing 64% less to run.

On a second benchmark, DeepSWE 1.1, SWE-2 reached 73.0%, ahead of Kimi K3 (68.5%), Grok 4.6 (67.5%) and Fable 5.1 (67.4%), and close behind GPT-5.6 Sol (72.7%) and GPT-6 Astra (74.1%), the two top scorers in Cognition’s comparison table. Cognition says SWE-2 beats its previous model, SWE-1.7, and Grok 4.6 on both score and cost, and comes within a few points of GPT-6 Astra at a quarter of that model’s price. The gap is not uniform: on Terminal-Bench 4, a separate test, SWE-2 scored 27.3%, well behind Fable 5.1’s 55.8% and GPT-6 Astra’s 57.9%.

SWE-2 is post-trained from Kimi K3, a 2.8-trillion-parameter base model, using infrastructure Cognition built for its earlier SWE-1.7 release. The company says the main technical addition is a reinforcement learning algorithm that trains every reasoning-effort level, the settings that trade speed for thoroughness, in a single training run rather than separately. Cognition describes tuning a cost penalty for each effort level to match the slope of the base model’s existing cost-performance curve, a method it says lets the model improve without changing the shape of that tradeoff.

Fewer steps, faster edits

Cognition says the efficiency gains come mostly from better judgment about which parts of a codebase matter for a given task. On FrontierCode 1.1 Main, the company reports that SWE-2 at medium effort scored higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average, and that it made its first real code edit after a median of 18 steps, compared with 48 for SWE-1.7. Cognition also says internal testing showed SWE-2 writing more thorough end-to-end tests, finding alternate routes when a tool or integration was unavailable, and re-checking its own conclusions rather than simply agreeing with a user’s hypothesis. These are the company’s own observations from internal use, not independently verified results.

SWE-2 is available starting today in Devin Desktop and Devin CLI, Cognition’s coding-agent products, with a rollout to Devin Web and Fusion underway. All of the benchmark figures above come from Cognition’s own blog post; independent testing of SWE-2 against Fable 5.1, GPT-6 Astra and other rivals has not yet been published. Watch for third-party benchmark runs and for how SWE-2 performs on Terminal-Bench 4, the one test in Cognition’s own numbers where it trails the top models by a wide margin.

Sources

  1. Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra Hacker News

AI-generated · AIVIO News Desk