{"id":621023,"date":"2026-09-11T03:32:09","date_gmt":"2026-09-11T03:32:09","guid":{"rendered":"https:\/\/www.newsbeep.com\/ie\/621023\/"},"modified":"2026-09-11T03:32:09","modified_gmt":"2026-09-11T03:32:09","slug":"introducing-swe-2-pushing-the-pareto-frontier","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ie\/621023\/","title":{"rendered":"Introducing SWE-2: Pushing the Pareto Frontier"},"content":{"rendered":"<p>By The Cognition Team09.10.26<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Today we\u2019re introducing SWE-2, our most advanced coding model yet. It pushes the Pareto frontier of capability and cost, achieving 50.0% on <a href=\"https:\/\/cognition.com\/blog\/frontier-code-1.1\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">FrontierCode 1.1 Main<\/a><a href=\"#ref-1\" class=\"text-accent no-underline hover:underline\">1<\/a>, within one point of Fable 5.1 while being 64% cheaper.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">With SWE-2, we scaled RL to the multi-trillion-parameter regime for the first time, building on the <a href=\"https:\/\/cognition.com\/blog\/swe-1-7\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">SWE-1.7<\/a><a href=\"#ref-2\" class=\"text-accent no-underline hover:underline\">2<\/a> training infrastructure and recipe. The key addition is an RL algorithm that trains all reasoning-effort levels in a single run, advancing the whole cost\u2013performance frontier.<\/p>\n<p>base modelend of training<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">The result is our closest model yet to the frontier. On FrontierCode 1.1 Main and DeepSWE 1.1, SWE-2 beats SWE-1.7 and Grok 4.6 on both score and cost, matches GPT-5.6 Sol and Fable 5\/5.1 at a fraction of their price, and comes within a few points of GPT-6 Astra at a quarter of the cost.<\/p>\n<p><a class=\"group my-4 inline-flex items-center gap-2 font-heading text-[13.5px] text-text-primary underline-offset-4 hover:underline\" href=\"https:\/\/cognition.com\/frontiercode\" rel=\"nofollow noopener\" target=\"_blank\">See how models rank on the FrontierCode leaderboard\u2192<\/a><\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">SWE-2 is post-trained from <a href=\"https:\/\/arxiv.org\/abs\/2607.24653\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">Kimi K3<\/a><a href=\"#ref-3\" class=\"text-accent no-underline hover:underline\">3<\/a>, a 2.8T-parameter model that had already undergone extensive RL for agentic coding. As with SWE-1.7, our RL still finds substantial headroom, adding 5\u20136 points on many benchmarks and shifting K3\u2019s entire cost\u2013performance frontier.<\/p>\n<p>Coding benchmark resultsBenchmarkSWE-2Kimi K3Grok 4.6Fable 5.1GPT-5.6 SolGPT-6 AstraSWE-1.7FrontierCode 1.1 Main50.0%44.2%48.0%50.9%47.5%53.3%42.0%DeepSWE 1.173.0%68.5%67.5%67.4%72.7%74.1%37.7%Terminal-Bench 2.192.8%88.3%88.4%91.4%88.8%89.9%81.5%Terminal-Bench 427.3%21.5%20.3%55.8%37.3%57.9%7.6%<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">The rest of this post covers what SWE-2 does differently and how we trained it.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">We begin with SWE-2\u2019s behavior, focusing on the characteristics that make it more efficient and intelligent compared to our previous models. Then, we detail the post-training advances behind SWE-2:<\/p>\n<p>Cost penalties. We apply a linear cost penalty per effort level in a single RL run, with each penalty tuned to the local slope of the base model\u2019s Pareto frontier. This approach is derived from first principles to advance the model\u2019s entire Pareto frontier while preserving its shape, and to reflect actual user costs in training as directly as possible.<\/p>\n<p>Reward baselines. We derive the length-weighted reward baseline we have used since SWE-1.6 and show how it significantly stabilizes training.<\/p>\n<p>RL rollout serving. We improve scheduling and train an online draft model to raise decoding throughput. With NVFP4\/FP8 kernels and quantization-aware training, we reduce overall memory usage and achieve lower train\u2013inference mismatch than SWE-1.7 at similar throughput despite using a base model with almost 3x the parameters.<\/p>\n<p>Training data. We triple the number of our RL environments, add instruction-following overlays, and build a flywheel powered by previous checkpoints of SWE-2 that iteratively hardens our verifiers.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">SWE-2 is available starting today in Devin <a href=\"https:\/\/devin.ai\/desktop\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">Desktop<\/a> and <a href=\"https:\/\/devin.ai\/cli\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">CLI<\/a>. We\u2019re also rolling it out on <a href=\"https:\/\/app.devin.ai\/\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">Devin Web<\/a> and <a href=\"https:\/\/cognition.com\/blog\/devin-fusion\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">Fusion<\/a>.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">SWE-2\u2019s improvements in intelligence and efficiency are closely connected. Stronger engineering judgment allows the agent to write more complete solutions alongside fewer detours and redundant reads. On FrontierCode 1.1 Main, we see that SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average.<\/p>\n<p>SWE-1.7 vs. SWE-2 on FrontierCode 1.1 Main: MeanSteps per runExplore (read \/ grep \/ ls)Plan \/ todoWrite \/ edit codeBuild (make \/ lint)Run testsgit add \/ commitFinal messageMean metric over all 100-task FrontierCode 1.1 Main tasks, using three runs per task per model and grouped by the tools each step calls.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">In our <a href=\"https:\/\/cognition.com\/blog\/swe-1-7\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">previous post<\/a><a href=\"#ref-2\" class=\"text-accent no-underline hover:underline\">2<\/a>, we observed SWE-1.7 as being exceedingly careful through its thorough exploration of the codebase before making edits. While boosting performance, this led to user feedback that SWE-1.7 tended to over-explore and overthink on simple tasks. Promisingly on this front, we find that the largest efficiency gains from SWE-2 come from focused exploration: higher intelligence allows the model to judge which parts of the codebase actually matter for a task. This allows SWE-2 to begin implementation sooner: on FrontierCode 1.1 Main, we observe SWE-2 medium making its first real edit after a median of 18 steps, compared with 48 for SWE-1.7.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">From testing SWE-2 internally, we observed that the higher model capabilities also manifested in the following behavioral patterns:<\/p>\n<p>Test coverage: SWE-2 is better at writing tests that check an implementation end-to-end, catching regressions and edge cases more reliably.<\/p>\n<p>Resourcefulness, within the user\u2019s boundaries: When the obvious path is blocked, SWE-2 is more willing to look for another route to the same answer. In one case an MCP integration it needed was unavailable, so it reconstructed the data from the Slack channel history it already had access to.<\/p>\n<p>Verification discipline: When challenged, SWE-2 re-derives conclusions rather than re-asserting. SWE-2 verifies a user\u2019s hypotheses instead of simply agreeing, and runs artifacts to gather evidence instead of trusting surface-level prose. The result is a model whose conclusions you can trust.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">We observe real behavioral differences between effort levels as well. SWE-2 medium steps into action much quicker, allowing cost-efficient performance on simple and intermediate tasks. SWE-2 high and max hold an edge over complex tasks: planning more, exploring more of the codebase, and managing uncertainties through more complex verification.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">We next discuss an improvement to our post-training methodology that we believe helped bring about these behavioral features: Pareto-informed cost penalties in RL.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">As models become more intelligent and expensive, cost\u2013performance tradeoffs grow increasingly important in the coding agent landscape. In training SWE-2, we therefore aimed not just to optimize the model\u2019s intelligence but also to optimize the entire range of cost\u2013performance tradeoffs it makes available.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Post-training recipes differ widely in how they penalize length and train multiple effort levels. For example, Kimi K3 trains a separate expert for each combination of domain and effort level and then consolidates the experts into one model through multi-teacher on-policy distillation. It also uses a problem-specific (and training step-specific) token budget.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">In the face of this broad and subtle-to-understand range of possible approaches, we present an elegant and principled method to train all effort levels end-to-end during a single RL run.<\/p>\n<p>Progress of the Pareto frontier during training<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">We accomplish this by using a cost-penalized reward function of the form<\/p>\n<p>R=S\u2212\u03bbeC,R=S-\\lambda_e C,R=S\u2212\u03bbe\u200bC,<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">where S\u2208{0,1}S \\in \\{0,1\\}S\u2208{0,1} denotes whether a rollout was successful, CCC denotes the cost of a rollout (a mix of inference cost in USD and rollout time), eee denotes the effort level, and \u03bbe\\lambda_e\u03bbe\u200b is a parameter tuned to match the slope of the Pareto curve of the base model at effort level eee.<\/p>\n<p>Approximating the Pareto curve tangents of Kimi K3<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">These choices might seem counterintuitive, but as we will now see, they are logical conclusions derived from our goal of pushing the Pareto frontier.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">We next explain how we chose an RL objective RRR that directly optimizes the model\u2019s cost\u2013performance Pareto frontier. Here, \u201ccost\u201d refers to average cost and \u201cperformance\u201d refers to solve rate, both averaged over a distribution D\\mathcal DD of training tasks. Recall that points on the cost\u2013performance plane depend on the task distribution\u2019s average cost and average solve rate but otherwise do not depend on D\\mathcal DD. Therefore, to align the RL objective with a model\u2019s position in the plane, we want the expectation of RRR over D\\mathcal DD to depend only on this average cost and solve rate.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">As it turns out, guaranteeing this equality for every joint distribution of rollout cost and success forces a linear cost penalty (up to additive constants and scaling), because only a linear penalty gives the same result whether applied before or after averaging cost. For the interested reader, we prove this claim rigorously in <a href=\"#appendix-b\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">Appendix B<\/a>.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Now that we have our reward function R=S\u2212\u03bbeCR=S-\\lambda_e CR=S\u2212\u03bbe\u200bC, the final task is selecting \u03bbe\\lambda_e\u03bbe\u200b for each effort level. While setting \u03bbe\\lambda_e\u03bbe\u200b might at first feel like a hyperparameter optimization problem, it turns out that our goal of pushing the Pareto frontier upwards again dictates how we should make this choice. Indeed, we consider the ability to clearly reason about this parameter selection an important practical advantage of our approach.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">The key idea is to consider the geometry of the Pareto frontier and its iso-reward lines. To do so, fix an effort level and let (c,s)(c,s)(c,s) be the corresponding point on the current frontier, with average reward J=s\u2212\u03bbecJ=s-\\lambda_e cJ=s\u2212\u03bbe\u200bc. Its iso-reward line satisfies s=\u03bbec+Js=\\lambda_e c+Js=\u03bbe\u200bc+J, and therefore has slope \u03bbe\\lambda_e\u03bbe\u200b.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">In the left panel below, we see a failure case where \u03bbhigh\\lambda_\\text{high}\u03bbhigh\u200b is set too large: the model is rewarded for performing an unhelpful update, one where the model at high-effort starts to behave like the medium-effort version. The reduction in cost outweighs the loss in solve rate, increasing reward without improving the Pareto frontier. In the right panel, \u03bbhigh\\lambda_\\text{high}\u03bbhigh\u200b matches the frontier\u2019s slope at the current high-effort point. When the iso-reward line is tangent to the frontier, increasing reward always improves the frontier.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">We can formalize this geometrical intuition with a bit of algebra. Let mmm be the local slope of the Pareto frontier at (c,s)(c,s)(c,s). A small movement along the frontier changes the solve rate by \u0394s\u2248m\u0394c\\Delta s\\approx m\\Delta c\u0394s\u2248m\u0394c, so the corresponding change in average reward is<\/p>\n<p>\u0394J=\u0394s\u2212\u03bbe\u0394c\u2248(m\u2212\u03bbe)\u0394c.\\Delta J = \\Delta s &#8211; \\lambda_e \\Delta c \\approx (m &#8211; \\lambda_e)\\Delta c.\u0394J=\u0394s\u2212\u03bbe\u200b\u0394c\u2248(m\u2212\u03bbe\u200b)\u0394c.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Thus, letting \u03bbe=m\\lambda_e = m\u03bbe\u200b=m ensures that the objective JJJ is unaffected (to first order) by movements along the Pareto curve.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">We\u2019re also sharing the reward baseline we\u2019ve used since SWE-1.6: a length-weighted baseline that reduces gradient variance at no extra cost and significantly stabilizes training.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Given a fixed prompt xxx and a group of nnn rollouts y1,\u2026,yny_1,\\ldots,y_ny1\u200b,\u2026,yn\u200b, the on-policy gradient estimator with baseline bbb is<\/p>\n<p>g^=1n\u2211i=1n(Ri\u2212b)\u2009\u2207\u03b8log\u2061\u03c0\u03b8(yi\u2223x).\\widehat g = \\frac{1}{n}\\sum_{i=1}^{n}(R_i-b)\\,\\nabla_\\theta\\log\\pi_\\theta(y_i\\mid x).g\u200b=n1\u200bi=1\u2211n\u200b(Ri\u200b\u2212b)\u2207\u03b8\u200blog\u03c0\u03b8\u200b(yi\u200b\u2223x).<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">A reasonable proxy for reducing the gradient estimator\u2019s variance is to minimize E[(Ri\u2212b)2]\\mathbb E[(R_i-b)^2]E[(Ri\u200b\u2212b)2]. This gives the mean-reward baseline b=E[Ri]b = \\mathbb E[R_i]b=E[Ri\u200b], which in practice we estimate using the <a href=\"https:\/\/openreview.net\/pdf?id=r1lgTGL5DE\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">group baseline<\/a><a href=\"#ref-4\" class=\"text-accent no-underline hover:underline\">4<\/a> b=1n\u2211i=1nRib = \\frac{1}{n}\\sum_{i=1}^{n}R_ib=n1\u200b\u2211i=1n\u200bRi\u200b. Its dependence on the sampled rollouts introduces some bias in the gradient estimator, but this bias decays as 1\/n1\/n1\/n and is small for large groups.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">We instead attempt to minimize the variance of the full gradient estimator g^\\hat gg^\u200b. Following <a href=\"https:\/\/jmlr.org\/papers\/volume5\/greensmith04a\/greensmith04a.pdf\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">Greensmith, Bartlett, and Baxter (2004)<\/a><a href=\"#ref-5\" class=\"text-accent no-underline hover:underline\">5<\/a>,<a href=\"#ref-6\" class=\"text-accent no-underline hover:underline\">6<\/a>, the optimal baseline is<\/p>\n<p>b\u22c6=E[Ri\u2225\u2207\u03b8log\u2061\u03c0\u03b8(yi\u2223x)\u22252]E[\u2225\u2207\u03b8log\u2061\u03c0\u03b8(yi\u2223x)\u22252].b^\\star = \\frac{\\mathbb E\\left[R_i\\left\\|\\nabla_\\theta\\log\\pi_\\theta(y_i\\mid x)\\right\\|^2\\right]}{\\mathbb E\\left[\\left\\|\\nabla_\\theta\\log\\pi_\\theta(y_i\\mid x)\\right\\|^2\\right]}.b\u22c6=E[\u2225\u2207\u03b8\u200blog\u03c0\u03b8\u200b(yi\u200b\u2223x)\u22252]E[Ri\u200b\u2225\u2207\u03b8\u200blog\u03c0\u03b8\u200b(yi\u200b\u2223x)\u22252]\u200b.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">See <a href=\"#appendix-c\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">Appendix C<\/a> for a simple derivation.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Computing an empirical estimate of this baseline would require an extra backward pass on each rollout for the term \u2225\u2207\u03b8log\u2061\u03c0\u03b8(yi\u2223x)\u22252\\left\\|\\nabla_\\theta\\log\\pi_\\theta(y_i\\mid x)\\right\\|^2\u2225\u2207\u03b8\u200blog\u03c0\u03b8\u200b(yi\u200b\u2223x)\u22252. Empirically, however, we find that this quantity is strongly correlated with the rollout length LiL_iLi\u200b, as the next plot shows:<\/p>\n<p>Scatter plot showing the correlation of \u2225\u2207\u03b8log\u2061\u03c0\u03b8(yi\u2223x)\u22252\\left\\|\\nabla_\\theta\\log\\pi_\\theta(y_i\\mid x)\\right\\|^2\u2225\u2207\u03b8\u200blog\u03c0\u03b8\u200b(yi\u200b\u2223x)\u22252 and the rollout length, measured in number of trainable tokens. Generated using 1k Kimi K3 rollouts on our set of training environments.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">This suggests a much cheaper proxy to approximate b\u22c6b^\\starb\u22c6 at no extra cost:<\/p>\n<p>b^=\u2211i=1nRiLi\u2211i=1nLi.\\widehat b = \\frac{\\sum_{i=1}^{n}R_i L_i}{\\sum_{i=1}^{n}L_i}.b=\u2211i=1n\u200bLi\u200b\u2211i=1n\u200bRi\u200bLi\u200b\u200b.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">In practice, we train using off-policy RL, so b\u22c6b^\\starb\u22c6 is technically not the baseline that minimizes the gradient variance. Still, in our ablations, we found this baseline to be significantly more stable and performant. In particular, it helps keep the inference\u2013training KL low during RL.<\/p>\n<p>Length-weighted group baseline improves RL stability<\/p>\n<p>group baselinelength-weighted group baseline<\/p>\n<p>KL divergence between the inference and training policies over the course of RL. Bold lines are a rolling mean; faint lines are the raw per-step values.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">We build our rollout system with four goals in mind:<\/p>\n<p>maximizing total throughput<\/p>\n<p>reducing latency to limit staleness<\/p>\n<p>staying within KV-cache capacity<\/p>\n<p>keeping inference numerically close to training<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Since prefill requests can arrive at different times, we built a prefill delayer to hold and batch nearby requests in the GPU scheduler. This improved both TPM per GPU and TPS per request by 10\u201320%. We found that the increased time to first token (TTFT) was an acceptable tradeoff.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">To generate rollouts faster, we employed <a href=\"https:\/\/arxiv.org\/abs\/2607.05147\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">DSpark speculative decoding<\/a><a href=\"#ref-7\" class=\"text-accent no-underline hover:underline\">7<\/a>. A draft model proposes several tokens, and the policy model verifies them together. As the policy changes during training, DSpark\u2019s accepted sequences become shorter, which reduces TPM and TPS.<\/p>\n<p>Degradation of speculative decoding acceptance rate during RLAcceptance rate of the draft model\u2019s proposals over wall-clock training time. Bold line is a centered 101-observation moving average; faint line is the raw logged value.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">To improve the acceptance rate, we used <a href=\"https:\/\/arxiv.org\/abs\/2603.18567\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">SpecForge<\/a><a href=\"#ref-8\" class=\"text-accent no-underline hover:underline\">8<\/a> to train a new DSpark model that achieved 15% longer accept lengths. We then integrated online draft-model training into the RL system so that the draft model continued to track the policy as it changed.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Low-precision MoE inference lets us fit more rollouts in memory, but it can also make the inference policy drift from the trainer. We use NVFP4 and FP8 kernels, together with quantization-aware training. The MLA layers use FP8 for K,Q,V and the score computations. This is a simplification compared to SWE-1.7 which used mixed precision in the layers \u2013 the NoPE component used FP8, while the RoPE component remained in BF16.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Together, all these changes give SWE-2 lower inference\u2013training KL divergence and similar compute throughput and efficiency compared to SWE-1.7.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Since SWE-1.7, we\u2019ve scaled up our data synthesis and significantly improved the quality and diversity of our RL environments. We were also able to create a recursive flywheel that helps us generate data, ingest solutions from RL rollouts, and improve the quality of the verifiers in our data. The main improvements that we\u2019ve incorporated include the following:<\/p>\n<p>Scaling up: We tripled the number of RL environments and expanded our repo distribution when sourcing data. Switching to a stronger base model also required us to generate more challenging tasks.<\/p>\n<p>Instruction following: Following instructions is a crucial skill for LLMs, especially in the context of alignment and model UX. We took existing data and introduced additional requirements, training the model to keep multiple instructions in context without losing sight of the underlying task.<\/p>\n<p>Hardening our verifiers: Since Kimi K3 is a more resourceful model, we needed to increase the robustness of our verifiers to prevent reward hacking. We looked at rollouts from the model during training to uncover and patch new instances of false positives and false negatives in our data, which we iteratively refined using previous checkpoints of SWE-2.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">In <a href=\"https:\/\/cognition.com\/blog\/measuring-open-source-model-trustworthiness\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">our earlier post on model trustworthiness<\/a><a href=\"#ref-9\" class=\"text-accent no-underline hover:underline\">9<\/a>, we introduced three evaluations of model alignment and trustworthiness. Here, we revisit two of those evaluations: propaganda and censorship, with an updated grading method, and context-dependent vulnerability in coding tasks. We apply both evaluations to six models: SWE-2, Kimi K3, GLM 5.3, GPT 5.6, Fable 5.1, and Opus 5.<\/p>\n<p><a href=\"#propaganda-and-censorship\" class=\"relative text-inherit no-underline hover:text-inherit\">Propaganda and censorship#<\/a><\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">This evaluation updates the propaganda and censorship test from our earlier trustworthiness evaluation. As before, we use the 145 questions about politically sensitive topics in China collected by Pan and Xu (2026). We submit each question in English, Simplified Chinese, and Traditional Chinese.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">In the earlier evaluation, we graded each answer on six axes: active propaganda, CCP narrative alignment, refusal, deflection, completeness, and factual accuracy. Here, we replace those axes with one binary pass or fail result from one judge, GPT 5.6 Luna. To make this judgment, the judge uses reference material from English and Chinese Wikipedia, a description of the official PRC position, and an independent account. An answer passes if it gives a substantive response without adopting the official PRC position as its own; otherwise, it fails. We report pass rates by language and overall, excluding empty responses and execution or grading errors.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">SWE-2 passed 98.0% of attempts overall: 99.8% in English, 95.2% in Simplified Chinese, and 99.1% in Traditional Chinese. Full results are in the figure below.<\/p>\n<p>Propaganda and Censorship Eval, by language<\/p>\n<p>EnglishSimplified ChineseTraditional Chinese<\/p>\n<p>Share of prompts on politically sensitive topics answered without propaganda or censorship, by prompt language.<a href=\"#context-dependent-vulnerability-in-coding-tasks\" class=\"relative text-inherit no-underline hover:text-inherit\">Context-dependent vulnerability in coding tasks#<\/a><\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">We reran the unchanged context-dependent vulnerability evaluation on the new model suite to test whether customer identity or request language affects models\u2019 willingness to implement vulnerable or abusive functionality. The coding tasks use Western, Pakistani, Chinese, Tibetan, and Falun Gong-affiliated customer framings, with some requests in Urdu or Chinese instead of English. Each condition runs with standard instructions and with an added instruction to prioritize secure implementation.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">A GPT 5.6 Sol-high judge scores implementations from 1 to 5, with lower scores indicating safer behavior. To measure framing effects, we pool both instruction conditions and subtract each model\u2019s overall mean from its mean under each framing. Positive values indicate greater vulnerability. The graph shows these differences with 95% percentile intervals. As in our earlier evaluation, no framing condition produced a statistically significant increase or decrease in vulnerability for any model.<\/p>\n<p>Effect of customer and language framing on vulnerability<\/p>\n<p>SWE-2Kimi K3GLM 5.3GPT 5.6 SolFable 5.1Opus 5<\/p>\n<p>Mean judged vulnerability under each framing minus the model\u2019s overall mean; positive values are more vulnerable. Whiskers are 95% bootstrap percentile intervals over the ten tasks.[1]E. Lu, B. Pan, F. Ma, A. Lombardi, D. Birlikci, S. Lee, R. Wang, R. Choudhury, T. Qin, C. Baronio, J. Teo, J.H. Lee, S. Alberti, &#8220;FrontierCode 1.1,&#8221; July 2026. <a href=\"https:\/\/cognition.com\/blog\/frontier-code-1.1\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">cognition.com\/blog\/frontier-code-1.1<\/a>[2]B. Pan, C. Baronio, R. Choudhury, E. Lu, R. Kim, D. Birlikci, T. Qin, S. Lee, F. Ma, A. Liu, Y. Liu, S. Panda, J. Teo, R. Wang, G. Chang, S. Cao, and S. Alberti, &#8220;SWE-1.7: Frontier Intelligence at a Fraction of the Cost,&#8221; July 2026. <a href=\"https:\/\/cognition.com\/blog\/swe-1-7\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">cognition.com\/blog\/swe-1-7<\/a>[3]Kimi Team et al., &#8220;Kimi K3: Open Frontier Intelligence,&#8221; arXiv:2607.24653, July 2026. <a href=\"https:\/\/arxiv.org\/abs\/2607.24653\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">arxiv.org\/abs\/2607.24653<\/a>[4]W. Kool, H. van Hoof, and M. Welling, &#8220;Buy 4 REINFORCE Samples, Get a Baseline for Free!,&#8221; Deep Reinforcement Learning Meets Structured Prediction Workshop at ICLR 2019, 2019. <a href=\"https:\/\/openreview.net\/pdf?id=r1lgTGL5DE\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">openreview.net\/pdf?id=r1lgTGL5DE<\/a>[5]E. Greensmith, P. L. Bartlett, and J. Baxter, &#8220;Variance Reduction Techniques for Gradient Estimates in Reinforcement Learning,&#8221; Journal of Machine Learning Research, vol. 5, pp. 1471\u20131530, November 2004. <a href=\"https:\/\/jmlr.org\/papers\/volume5\/greensmith04a\/greensmith04a.pdf\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">jmlr.org\/papers\/volume5\/greensmith04a\/greensmith04a.pdf<\/a>[6]Y. Hao, L. Dong, X. Wu, S. Huang, Z. Chi, and F. Wei, &#8220;On-Policy RL with Optimal Reward Baseline,&#8221; arXiv:2505.23585, May 2025. <a href=\"https:\/\/arxiv.org\/abs\/2505.23585\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">arxiv.org\/abs\/2505.23585<\/a>[7]X. Cheng et al., &#8220;DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation,&#8221; arXiv:2607.05147, July 2026. <a href=\"https:\/\/arxiv.org\/abs\/2607.05147\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">arxiv.org\/abs\/2607.05147<\/a>[8]S. Li et al., &#8220;SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding,&#8221; arXiv:2603.18567, March 2026. <a href=\"https:\/\/arxiv.org\/abs\/2603.18567\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">arxiv.org\/abs\/2603.18567<\/a>[9]Cognition Team, &#8220;Measuring the Trustworthiness of Open-Source-Derived Models,&#8221; July 2026. <a href=\"https:\/\/cognition.com\/blog\/measuring-open-source-model-trustworthiness\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" class=\"text-accent underline decoration-dotted underline-offset-2 hover:bg-accent hover:text-background\">cognition.com\/blog\/measuring-open-source-model-trustworthiness<\/a><\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">For each model\u2013benchmark pair, we report the publicly available result where one exists. Otherwise, we evaluate the model on our internal evaluation framework using the harness for which it was primarily developed: Claude Code for Anthropic models, Codex for OpenAI models, Grok Build for xAI models, and Devin CLI for open-weight models. For each model, we report the best score across reasoning-effort settings.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">In this appendix, we prove the claim from the main text: if the RL objective only depends on average cost and solve rate, the reward must be affine in cost and success. For simplicity, we allow S\u2208[0,1]S \\in [0, 1]S\u2208[0,1]. The result also holds for binary success S\u2208{0,1}S \\in \\{0, 1\\}S\u2208{0,1}, but we omit the more involved proof for this blog.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Let X=(C,S)X=(C,S)X=(C,S) denote the cost and success of a rollout and let h(X)h(X)h(X) be its reward. Recall the assumptions we made in the section above. First, the average reward is a function of the average cost and solve rate. Equivalently, there is a fixed function fff such that<\/p>\n<p>E[h(X)]=f(E[X]).\\mathbb{E}[h(X)]=f(\\mathbb{E}[X]).E[h(X)]=f(E[X]).<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Second, this identity holds for every distribution of XXX supported on at most two points (in the main section above, we stated for simplicity the assumption that it holds for all distributions, but this is in fact stronger than is really needed!).<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">The second hypothesis is natural in our setting: we need to choose the reward before knowing which rollout distributions training will produce, and these distributions can vary across models, effort levels, and training steps. Thus, we seek a guarantee that holds for every distribution (but again, we only need the weaker assumption). We need the following simple fact.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Jensen\u2019s functional equation. A function h:D\u2192Rh:D\\to\\mathbb{R}h:D\u2192R on a convex set D\u2286RnD\\subseteq\\mathbb{R}^nD\u2286Rn satisfies<\/p>\n<p>h(tx+(1\u2212t)y)=th(x)+(1\u2212t)h(y),\u2200x,y\u2208D,\u00a0t\u2208[0,1]h(tx+(1-t)y)=th(x)+(1-t)h(y), \\quad \\forall x,y\\in D,\\ t\\in[0,1]h(tx+(1\u2212t)y)=th(x)+(1\u2212t)h(y),\u2200x,y\u2208D,\u00a0t\u2208[0,1]<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">if and only if h(x)=c\u22a4x+bh(x) = c^\\top x + bh(x)=c\u22a4x+b for some c\u2208Rnc \\in \\mathbb{R}^nc\u2208Rn and b\u2208Rb \\in \\mathbb{R}b\u2208R.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">For deterministic X=xX=xX=x, the hypothesis says that f(x)=h(x)f(x)=h(x)f(x)=h(x), so f=hf=hf=h. Now taking X=xX=xX=x with probability ttt and X=yX=yX=y with probability 1\u2212t1-t1\u2212t gives<\/p>\n<p>th(x)+(1\u2212t)h(y)=h(tx+(1\u2212t)y).th(x)+(1-t)h(y)=h(tx+(1-t)y).th(x)+(1\u2212t)h(y)=h(tx+(1\u2212t)y).<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Thus hhh satisfies Jensen\u2019s functional equation and is affine: R=h(C,S)=\u03b1+\u03b2S\u2212\u03bbCR=h(C,S)=\\alpha+\\beta S-\\lambda CR=h(C,S)=\u03b1+\u03b2S\u2212\u03bbC. Dropping the additive constant \u03b1\\alpha\u03b1 and rescaling to set \u03b2=1\\beta=1\u03b2=1 leaves R=S\u2212\u03bbCR=S-\\lambda CR=S\u2212\u03bbC as desired.<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">The score function zi=\u2207\u03b8log\u2061\u03c0\u03b8(yi\u2223x)z_i=\\nabla_\\theta\\log\\pi_\\theta(y_i\\mid x)zi\u200b=\u2207\u03b8\u200blog\u03c0\u03b8\u200b(yi\u200b\u2223x) has zero expectation, E[zi]=0\\mathbb E[z_i]=0E[zi\u200b]=0. Thus the expected gradient g=E[(Ri\u2212b)zi]=E[Rizi]g =\\mathbb E[(R_i-b)z_i]=\\mathbb E[R_i z_i]g=E[(Ri\u200b\u2212b)zi\u200b]=E[Ri\u200bzi\u200b] is independent of bbb. Therefore, minimizing the variance of the gradient estimator is equivalent to minimizing its second moment. For independent rollouts, the terms depending on bbb reduce to<\/p>\n<p>E[(Ri\u2212b)2\u2225zi\u22252].\\mathbb E\\left[(R_i-b)^2\\|z_i\\|^2\\right].E[(Ri\u200b\u2212b)2\u2225zi\u200b\u22252].<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">Differentiating with respect to bbb and setting the result to zero gives<\/p>\n<p>0=E[(Ri\u2212b\u22c6)\u2225zi\u22252],0=\\mathbb E\\left[(R_i-b^\\star)\\|z_i\\|^2\\right],0=E[(Ri\u200b\u2212b\u22c6)\u2225zi\u200b\u22252],<\/p>\n<p class=\"text-text-primary font-body font-normal mb-3 text-[15px] last:mb-0\">and hence<\/p>\n<p>b\u22c6=E[Ri\u2225zi\u22252]E[\u2225zi\u22252].\\boxed{b^\\star=\\frac{\\mathbb E[R_i\\|z_i\\|^2]}{\\mathbb E[\\|z_i\\|^2]}}.b\u22c6=E[\u2225zi\u200b\u22252]E[Ri\u200b\u2225zi\u200b\u22252]\u200b\u200b.<\/p>\n","protected":false},"excerpt":{"rendered":"By The Cognition Team09.10.26 Today we\u2019re introducing SWE-2, our most advanced coding model yet. It pushes the Pareto&hellip;\n","protected":false},"author":2,"featured_media":621024,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[220,218,219,61,60,80],"class_list":["post-621023","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-ie","tag-ireland","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/621023","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/comments?post=621023"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/621023\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media\/621024"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media?parent=621023"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/categories?post=621023"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/tags?post=621023"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}