Cursor Composer 2.5 beats GPT-5.5 on code at $0.07 per task
Cursor launched Composer 2.5 on the open-weight Kimi K2.5 base, with 85% of compute in Cursor's own RL. Third on the Coding Agent Index at $0.07 per task — 10 to 60× cheaper than rivals. Proof that specialised fine-tuning reaches the coding frontier.
Summary
On May 18, Cursor launched Composer 2.5. The model is built on Kimi K2.5 (Moonshot AI's open-weight base), with 85% of training compute dedicated to Cursor's own reinforcement learning. Result: third place on Artificial Analysis's Coding Agent Index, 77.6% on SWE-Bench Verified Multilingual, and a cost of $0.07 per task. Main rivals range from $0.70 to $4.20 per task — 10 to 60 times more expensive. The benchmark includes a direct win over GPT-5.5 on SWE-Bench Multilingual.
In practice
Composer 2.5 is available inside the Cursor editor as the default option for mid-length agentic tasks. The $0.07 per task cost changes usage behaviour: running the agent in fast iteration loops no longer requires billing control. For teams already using Cursor as their main environment, there's no workflow change — the model swapped in the background.
Context
What makes this launch relevant isn't the benchmark — it's the technical path. It shows that fine-tuning an open-weight model with specialised RL can compete with the best closed models on coding tasks. It's not imitation of a larger model: 85% of the compute went into Cursor's own post-training on top of an open base. This pattern — quality open-weight model plus vertical RL — is becoming viable for startups without the scale to train from scratch. Kimi K2.6, Moonshot AI's next model, meanwhile ranked 4th on the Artificial Analysis Intelligence Index overall, validating the base Cursor used.
Why it matters
- $0.07 per task against $0.70–$4.20 from rivals: 10 to 60× cheaper with top benchmark results
- 85% of compute in Cursor's own RL on open-weight base — specialised fine-tuning reaches the frontier without training from scratch
- Third on the global Coding Agent Index, behind only closed frontier models
- Proof of concept for startups: massive training compute is not required to compete in agentic coding