Arena Leaderboard

See how models compare against each other based on real coding tasks and votes from Devin Desktop users using Arena Mode.

This leaderboard is no longer maintained. The results below are final as of June 30, 2026; the newest models evaluated are Claude Opus 4.6 and GPT-5.5, and nothing released since is included.

DevinFinal results: June 30, 2026
7008009001,0001,1001,200ELO Rating (better →)1,098ClaudeOpus 4.61,081ClaudeOpus 4.51,058ClaudeSonnet4.51,031GPT-5.41,028Kimi K2.61,022GPT-5.51,015ClaudeHaiku 4.51,011GPT-5.2959GPT-5.3-Codex953Gemini 3Flash937Gemini3.1 Pro914GPT-5.3-CodexSpark911Grok CodeFast 1Model Preference (weaker →)
RankModelELO95% CIOrganization
1
Claude Opus 4.6
1,098±14Anthropic
2
Claude Opus 4.5
1,081±19Anthropic
3
Claude Sonnet 4.5
1,058±18Anthropic
4
GPT-5.4
1,031±23OpenAI
5
Kimi K2.6
1,028±65Moonshot
6
GPT-5.5
1,022±44OpenAI
7
Claude Haiku 4.5
1,015±30Anthropic
8
GPT-5.2
1,011±18OpenAI
9
GPT-5.3-Codex
959±27OpenAI
10
Gemini 3 Flash
953±28Google
11
Gemini 3.1 Pro
937±28Google
12
GPT-5.3-Codex Spark
914±35OpenAI
13
Grok Code Fast 1
911±28xAI
  1. Scores are calculated using ELO ratings from Arena Mode usage.
  2. User preference is derived from side-by-side comparisons where users select their preferred response. The chosen response replaces the other and becomes the basis for the next turn.
  3. Battle groups included models representative of daily Devin usage and were updated as new models became available. Some model tiers never appeared on the leaderboard.
  4. Unlike other leaderboards, Devin Arena does not penalize models for faster generation speed by holding back responses.
get the app

Build more with Devin

See all download options →