Coding Agent Leaderboard

How coding agents perform on real software-engineering tasks, with API cost and runtime per task

Coding Agent Index v1.5Snapshot: 2026-09-14
Sort
#Agent / modelIndexCost / taskTime / task
1
Claude Code
62.2%Best$12.3934.8m
2
Devin Fusion CLI
61.7%$7.9035.8m
3
Codex
61.6%$7.4729.4m
4
Claude Code
59.7%$10.7941.9m
5
Devin Fusion CLI
58.9%$4.5424.7m
6
Codex
54.6%$6.5820.6m
7
Muse Code
Muse Spark 1.3 (max)
54.3%$3.9818.4m
8
Opencode
53.6%$4.2448.1m
9
Kimi Code CLI
51.9%$5.051h
10
Grok Build
47.0%$3.5719.5m
11
Claude Code
43.3%$3.491.1h
12
Codex
43.1%$0.2440.3m
13
Antigravity SDK v0.1.12
41.9%$2.4711.7m
14
Codex
38.7%$0.0818.3m

How the index is built

Each agent runs three attempts per task across three suites — DeepSWE v1.1 (113 software-engineering tasks), Terminal-Bench 4.0 (66 terminal tasks) and SWE-Atlas-QnA (124 technical Q&A tasks) — scored as pass@1. The index is their equal-weight average. Cost is measured token usage × provider token prices, not subscription plans.