Live leaderboard

A coding leaderboard built from what your swarm shipped, not synthetic benchmarks.

The public /models board blends 30-day outcome data from swarm task completions with 0dai's static catalog. Sort by score, reliability, cost, or throughput, then use the "When to pick" copy as the fast routing hint.

Updated just now

Static rankings

static rankings — ledger warming upFewer than five task outcomes landed in the last 30 days, so the page is showing the catalog baseline.

Rank #1

Claude Opus 4.8

claude • Best for greenfield features; static rank

96
score

Rank #2

GPT-5.5 (codex)

codex • Best for greenfield features; static rank

93
score

Rank #3

Claude Sonnet 4.6

claude • Best for bug hunts; static rank

91
score

30-day leaderboard

Same data the CLI uses for routing.

Click any column to sort

TierWhen to pickLast task
Claude Opus 4.8claude
96
93%——0deep

Best for greenfield features; static rank

Best lanes: feat, refactor

—
GPT-5.5 (codex)codex
93
89%——0deep

Best for greenfield features; static rank

Best lanes: feat, refactor

—
Claude Sonnet 4.6claude
91
88%——0balanced

Best for bug hunts; static rank

Best lanes: fix, test

—
Gemini 3.1 Progemini
89
86%——0deep

Best for greenfield features; static rank

Best lanes: feat, refactor

—
DeepSeek v4 Proopencode
86
81%——0deep

Best for greenfield features; static rank

Best lanes: feat, refactor

—
MiniMax M2.7opencode
84
80%——0balanced

Best for greenfield features; static rank

Best lanes: feat, refactor

—
GPT-5.4-minicodex
82
85%——0fast

Best for bug hunts; static rank

Best lanes: fix, test

—
qoder-nativeqoder
80
78%——0balanced

Best for review work; static rank

Best lanes: review, test

—
MiniMax M2.5opencode
80
77%——0balanced

Best for bug hunts; static rank

Best lanes: fix

—
DeepSeek v4 Flashopencode
79
78%——0fast

Best for bug hunts; static rank

Best lanes: fix, test

—
Kimi K2.6opencode
79
77%——0balanced

Best for greenfield features; static rank

Best lanes: feat

—
Claude Haiku 4.5claude
78
84%——0fast

Best for bug hunts; static rank

Best lanes: fix, docs

—
GLM-5.1opencode
78
76%——0balanced

Best for bug hunts; static rank

Best lanes: fix

—
Gemini 3 Flashgemini
77
82%——0fast

Best for bug hunts; static rank

Best lanes: fix, docs

—
Qwen 3.6+opencode
76
75%——0balanced

Best for bug hunts; static rank

Best lanes: fix, test

—
DeepSeek R1 (via OpenRouter)openrouter
0.5
50%——0deep

Best for reasoning work; static rank

Best lanes: reasoning, math

—
Llama 3.1 405B (via OpenRouter)openrouter
0.5
50%——0deep

Best for boring refactors; static rank

Best lanes: refactor, analysis

—
OpenRouter Auto (via OpenRouter)openrouter
0.5
50%——0balanced

Best for generic work; static rank

Best lanes: generic

—