Cognition shipped SWE-2 on 10 September 2026, its most advanced coding model, built on Moonshot AI’s 2.8-trillion-parameter Kimi K3 base and trained with a novel multi-effort reinforcement learning method. On the FrontierCode 1.1 Main benchmark — which measures whether AI-generated pull requests would actually be merged — SWE-2 scores 50.0%, placing it within one point of Anthropic’s Fable 5.1 (51%) and within two points of OpenAI’s GPT-6 Astra (52%), at 64% lower cost than those frontier models to run.
What SWE-2 Scores on Coding Benchmarks
FrontierCode 1.1 Main measures real-world PR merge rate, making it a closer proxy for production engineering value than synthetic pass-rate benchmarks. SWE-2 scores 50.0% on this benchmark, outperforming its predecessor SWE-1.7 and xAI’s Grok 4.6. At 64% lower cost than frontier models for the same FrontierCode result, SWE-2 removes the primary objection that has stalled enterprise AI coding deployments: cost.
On Terminal-Bench 4, which tests harder agentic and multi-step engineering tasks, the gap widens significantly. SWE-2 scores 27.3%, compared with 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra — a 28-to-30-point gap that reflects SWE-2’s current limits on complex, deeply sequenced engineering work.
How Cognition Trained SWE-2
Cognition built SWE-2 on top of Kimi K3, an open-weight model from Moonshot AI with 2.8 trillion parameters — Cognition did not train the base model from scratch. The novel element is the training method: a single reinforcement learning run that simultaneously trains the model across medium, high, and max effort levels, rather than training separate models for each difficulty tier. This approach allows SWE-2 to allocate more compute to harder tasks at inference time.
Where SWE-2 Is Available
SWE-2 is live across all four Devin surfaces as of 10 September 2026: Devin Desktop (native app), Devin CLI (terminal integration), Devin Web (browser interface), and Devin Fusion (combined environment). Engineering teams already using Devin receive access to SWE-2 through their existing subscriptions without a separate onboarding step.
What This Means for Engineering Teams Evaluating AI Coding Agents
SWE-2 establishes a meaningful tier split in the AI coding agent market. For teams with bounded, routine coding tasks — code review, standard PR generation, test writing — SWE-2 delivers near-frontier merge-rate performance at 64% lower cost than running Fable 5.1 or GPT-6 Astra. For teams requiring multi-step agentic reasoning on complex engineering problems, the 27% Terminal-Bench 4 score versus 55–57% for frontier models is the honest limiting factor.
Cognition competes directly with GitHub Copilot Workspace, Cursor, and Windsurf for enterprise engineering budgets. The Kimi K3 base, built by a Chinese lab under open weights, demonstrates that frontier-competitive coding performance is now accessible to specialized labs applying targeted RL, rather than requiring foundational model training from scratch. For teams tracking the best AI agents for business tasks, SWE-2 is the first evidence that frontier-grade coding agent quality is about to fall sharply in price.
For Context: The demand SWE-2 targets is well-documented: 8 in 10 engineers now use AI agents daily, according to Temporal’s 2026 developer report, with coding assistance as the dominant use case. The developer community’s interest in SWE-2 is high — the announcement reached an HN score of 392, concentrated among the engineers and CTOs making or influencing AI tooling decisions.
Our Take: SWE-2 is the first credible signal that the cost of frontier-quality coding agents is about to fall sharply. At 50% on FrontierCode 1.1 for 36% of the price of Fable 5.1, it makes the “too expensive for production” objection obsolete for standard PR work. The gap on hard agentic tasks — 27% versus 56% — is the honest asterisk, and teams should weight it according to whether their workload is bounded or open-ended.

