Skip to content

/ LLM benchmarks / Agentic Coding

Agentic coding benchmarks

Multi-step code editing and tool use in agentic workflows, from LiveBench.

DataLiveBenchrelease 2026-06-25fetched

LiveBench Agentic Coding · CC BY-SA 4.0

View original source
LiveBench Agentic Coding leaderboard, 58 models
#ModelScoreCost/solved task
1
DeepSeek V4.1 Flash Max EffortDeepSeek · open weights
77.3%
$0.029
2
Claude 5.5 Opus Thinking Max EffortAnthropic
71.7%
$0.80
3
Claude Fable 5.1 Max EffortAnthropic
66.1%
$1.21
4
Claude 5 Opus Thinking Max EffortAnthropic
65.2%
$0.70
5
DeepSeek V4 Flash Vision ExpDeepSeek · open weights
65.1%
$0.051
6
Qwen 3.8 MaxAlibaba · open weights
64.6%
$0.28
7
Muse Spark 1.3 xHigh EffortMeta
64.1%
$0.22
8
Claude Fable 5 Max EffortAnthropic
62.2%
$1.44
9
Kimi K3Moonshot AI · open weights
62.2%
$0.35
10
Qwen 3.8 Flash NextAlibaba · open weights
61.6%
$0.042
11
Qwen3.8 27BAlibaba · open weights
61.4%
$0.094
12
GLM-5.3Z.AI · open weights
60.9%
$0.45
13
Claude Sonnet 5 xHigh EffortAnthropic
59.4%
$0.51
14
Muse Spark 1.1 xHigh EffortMeta
58.5%
$0.20
15
Gemini 3.7 Flash HighGoogle
58.3%
$0.16
16
Muse Spark 1.2 xHigh EffortMeta
57.6%
$0.38
17
GPT-6 Astra Max EffortOpenAI
57.3%
$0.74
18
Grok 4.6 xHighxAI
57.0%
$0.21
19
GLM-5.3 FlashZ.AI · open weights
56.8%
$0.031
20
Grok 4.5xAI
56.5%
$0.13
21
GPT-5.6 Sol Max EffortOpenAI
56.2%
$0.52
22
GPT-5.6 Terra Max EffortOpenAI
54.9%
$0.35
23
DeepSeek V4 Pro 0813DeepSeek · open weights
54.9%
$0.044
24
GPT-6.1 Sol Max EffortOpenAI
54.5%
$0.14
25
Gemini 3.8 Flash HighGoogle
54.2%
$0.31
26
GPT-5.5 Thinking xHigh EffortOpenAI
54.0%
$0.43
27
Grok 4.7 xHighxAI
54.0%
$0.72
28
GPT-5.4 Thinking xHigh EffortOpenAI
53.8%
$0.39
29
GPT-6 Sol Max EffortOpenAI
52.9%
$0.27
30
Ox Alpha MaxStealth
52.6%
–
31
GLM-5.2Z.AI · open weights
51.8%
$0.23
32
GPT-6 Luna Max EffortOpenAI
51.2%
$0.026
33
Claude 4.7 Opus Thinking xHigh EffortAnthropic
50.7%
$0.53
34
Claude 4.8 Opus Thinking Max EffortAnthropic
50.5%
$0.98
35
GPT-5.2 HighOpenAI
50.3%
$0.23
36
GPT-5.2 CodexOpenAI
49.4%
$0.19
37
Inkling xHigh EffortThinking Machines · open weights
49.4%
$0.31
38
Gemini 3.5 Flash HighGoogle
49.0%
$0.25
39
Claude 4.6 Opus Thinking High EffortAnthropic
49.0%
$0.40
40
GPT-5.6 Luna Max EffortOpenAI
48.4%
$0.17
41
Kimi K2.6 ThinkingMoonshot AI · open weights
46.9%
$0.17
42
DeepSeek V4 Flash 0731DeepSeek · open weights
46.8%
$0.060
43
GPT-5.4 Nano xHighOpenAI
46.8%
$0.091
44
Grok Build 0.1xAI
45.8%
$0.024
45
Kimi K2.7 CodeMoonshot AI · open weights
45.7%
$0.10
46
Gemini 3.5 Flash-Lite HighGoogle
45.3%
$0.069
47
Gemini 3.1 Pro Preview HighGoogle
44.1%
$0.29
48
Qwen 3.7 MaxAlibaba
43.6%
$0.18
49
Gemini 3.6 Flash HighGoogle
43.4%
$0.23
50
Claude 4.6 Sonnet Thinking Medium EffortAnthropic
42.6%
$0.31
51
GPT-5.4 Mini xHighOpenAI
41.7%
$0.33
52
Qwen 3.6 PlusAlibaba
41.4%
$0.23
53
Minimax M3MiniMax
40.7%
$0.060
54
Claude 4.5 Opus Thinking High EffortAnthropic
39.7%
$0.61
55
Claude Sonnet 5.5 xHigh EffortAnthropic
39.3%
$0.14
56
Qwen 3.6 27BAlibaba · open weights
39.3%
$0.20
57
Nemotron 3 Ultra 550B A55BNVIDIA · open weights
38.7%
$0.37
58
Grok 4.3xAI
18.5%
$0.061

Need a model picked for your use case?

Benchmarks rank models on other people's tasks. For a business workflow I test the shortlist on your own data, costs and privacy requirements, and recommend one.

AI consulting