Your Site Name
返回新闻

RT Jean P.D. Meijer ― 🇪🇺 eu/acc:方便发布终端基准2.1的得分,而省略终端基准4的得分,这其实也不算差……

RT Jean P.D. Meijer ― 🇪🇺 eu/acc: convenient to post the Terminal-Bench 2.1 score, while leaving Terminal-Bench 4 out that's not a bad score, bu...

Peter SteinbergerAI2026-09-10
RT Jean P.D. Meijer ― 🇪🇺 eu/accconvenient to post the Terminal-Bench 2.1 score, while leaving Terminal-Bench 4 out that's not a bad score, but just be honest Cognition: Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.

原文

RT Jean P.D. Meijer ― 🇪🇺 eu/accconvenient to post the Terminal-Bench 2.1 score, while leaving Terminal-Bench 4 outthat's not a bad score, but just be honestCognition: Introducing SWE-2, our closest model yet to the frontier.On leading evals, it scores on par with recent frontier models – at up to 70% lower cost.We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.