/ LLM benchmarks
LLM benchmarks: the best model for each job
The current leader for each task on Arena and LiveBench, with price or cost per task where the source reports one. Every score is the source's own.
DataText Arena2 Oct 2026Code Arena1 Oct 2026LiveBenchrelease 2026-06-25fetched
/ Best model for the job
- CodingClaude Opus 5.5 MaxAnthropic1815Elo$4.00 / $20.00 per 1M tokensCode Arena
- Writing and chatGemini 4 Argon HighGoogle1525Elo$2.00 / $10.00 per 1M tokensText Arena
- ReasoningGPT-6 Astra Max EffortOpenAI92.7%$0.74 per solved task, LiveBench overallLiveBench Reasoning
- MathClaude 5.5 Opus Thinking Max EffortAnthropic97.1%$0.80 per solved task, LiveBench overallLiveBench Math
- Data analysisGPT-6 Astra Max EffortOpenAI83.0%$0.74 per solved task, LiveBench overallLiveBench Data Analysis
- Agentic codingDeepSeek V4.1 Flash Max EffortDeepSeek · open weights77.3%$0.029 per solved task, LiveBench overallLiveBench Agentic Coding
- Instruction followingGemini 3.8 Flash HighGoogle81.4%$0.31 per solved task, LiveBench overallLiveBench Instruction Following
- Best valueDeepSeek V4.1 Flash Max EffortDeepSeek · open weights81.1%overall$0.029 per solved task, LiveBench overallLiveBench Overall · computed
- Best open-weights modelDeepSeek V4.1 Flash Max EffortDeepSeek · open weights81.1%overall$0.029 per solved task, LiveBench overallLiveBench Overall · computed
- Cheapest strong modelDeepSeek V4.1 Flash Max EffortDeepSeek · open weights81.1%overall$0.029 per solved task, LiveBench overallLiveBench Overall · computed
- Best value: Highest LiveBench overall score per dollar, among models within 5 points of the leader.
- Best open-weights model: Highest LiveBench overall score among open-weights models.
- Cheapest strong model: Lowest cost per successful task with at least 90% of the leader's overall score.
Cost vs quality
Each dot is a model on LiveBench: overall score against cost per successful task. The line is the Pareto frontier, where no other model is both cheaper and better.
Swipe sideways to see the whole chart.
57 models. On the frontier, from cheapest to best: Grok Build 0.1 ($0.024, 67.8%), GPT-6 Luna Max Effort ($0.026, 72.0%), DeepSeek V4.1 Flash Max Effort ($0.029, 81.1%), GPT-6.1 Sol Max Effort ($0.14, 81.6%), GPT-6 Astra Max Effort ($0.74, 82.2%), Claude 5.5 Opus Thinking Max Effort ($0.80, 83.2%), Claude Fable 5.1 Max Effort ($1.21, 83.4%).
Show the chart data as a table
| Model | Organization | Overall | Cost/solved task | Open weights | Frontier |
|---|---|---|---|---|---|
| Claude Fable 5.1 Max Effort | Anthropic | 83.4 | $1.21 | yes | |
| Claude 5.5 Opus Thinking Max Effort | Anthropic | 83.2 | $0.80 | yes | |
| Claude Fable 5 Max Effort | Anthropic | 83.0 | $1.44 | ||
| GPT-6 Astra Max Effort | OpenAI | 82.2 | $0.74 | yes | |
| GPT-6.1 Sol Max Effort | OpenAI | 81.6 | $0.14 | yes | |
| Muse Spark 1.3 xHigh Effort | Meta | 81.6 | $0.22 | ||
| DeepSeek V4.1 Flash Max Effort | DeepSeek | 81.1 | $0.029 | yes | yes |
| GPT-5.6 Sol Max Effort | OpenAI | 81.0 | $0.52 | ||
| GPT-5.5 Thinking xHigh Effort | OpenAI | 80.2 | $0.43 | ||
| Claude 5 Opus Thinking Max Effort | Anthropic | 80.1 | $0.70 | ||
| GPT-6 Sol Max Effort | OpenAI | 79.3 | $0.27 | ||
| Kimi K3 | Moonshot AI | 79.2 | $0.35 | yes | |
| Gemini 3.7 Flash High | 78.8 | $0.16 | |||
| Qwen 3.8 Max | Alibaba | 78.5 | $0.28 | yes | |
| Grok 4.6 xHigh | xAI | 78.0 | $0.21 | ||
| GPT-5.4 Thinking xHigh Effort | OpenAI | 78.0 | $0.39 | ||
| Muse Spark 1.2 xHigh Effort | Meta | 78.0 | $0.38 | ||
| GPT-5.6 Terra Max Effort | OpenAI | 77.9 | $0.35 | ||
| Claude Sonnet 5.5 xHigh Effort | Anthropic | 77.8 | $0.14 | ||
| DeepSeek V4 Pro 0813 | DeepSeek | 77.4 | $0.044 | yes | |
| Grok 4.7 xHigh | xAI | 77.4 | $0.72 | ||
| Gemini 3.1 Pro Preview High | 77.0 | $0.29 | |||
| DeepSeek V4 Flash Vision Exp | DeepSeek | 76.8 | $0.051 | yes | |
| Claude 4.7 Opus Thinking xHigh Effort | Anthropic | 76.5 | $0.53 | ||
| Claude 4.8 Opus Thinking Max Effort | Anthropic | 76.2 | $0.98 | ||
| Qwen 3.8 Flash Next | Alibaba | 76.2 | $0.042 | yes | |
| GLM-5.3 | Z.AI | 76.1 | $0.45 | yes | |
| Claude Sonnet 5 xHigh Effort | Anthropic | 76.0 | $0.51 | ||
| Gemini 3.8 Flash High | 75.8 | $0.31 | |||
| Grok 4.5 | xAI | 75.8 | $0.13 | ||
| Muse Spark 1.1 xHigh Effort | Meta | 75.3 | $0.20 | ||
| Qwen3.8 27B | Alibaba | 75.3 | $0.094 | yes | |
| Gemini 3.5 Flash High | 74.6 | $0.25 | |||
| GPT-5.2 High | OpenAI | 74.6 | $0.23 | ||
| Claude 4.6 Opus Thinking High Effort | Anthropic | 74.5 | $0.40 | ||
| DeepSeek V4 Flash 0731 | DeepSeek | 74.2 | $0.060 | yes | |
| GPT-5.2 Codex | OpenAI | 74.0 | $0.19 | ||
| Gemini 3.6 Flash High | 73.6 | $0.23 | |||
| GPT-5.6 Luna Max Effort | OpenAI | 73.6 | $0.17 | ||
| GLM-5.2 | Z.AI | 73.2 | $0.23 | yes | |
| Qwen 3.7 Max | Alibaba | 73.1 | $0.18 | ||
| Claude 4.6 Sonnet Thinking Medium Effort | Anthropic | 73.0 | $0.31 | ||
| Claude 4.5 Opus Thinking High Effort | Anthropic | 72.6 | $0.61 | ||
| GPT-6 Luna Max Effort | OpenAI | 72.0 | $0.026 | yes | |
| Inkling xHigh Effort | Thinking Machines | 71.9 | $0.31 | yes | |
| GLM-5.3 Flash | Z.AI | 71.6 | $0.031 | yes | |
| Kimi K2.6 Thinking | Moonshot AI | 70.5 | $0.17 | yes | |
| GPT-5.4 Nano xHigh | OpenAI | 69.6 | $0.091 | ||
| Qwen 3.6 Plus | Alibaba | 68.9 | $0.23 | ||
| Kimi K2.7 Code | Moonshot AI | 68.4 | $0.10 | yes | |
| Grok Build 0.1 | xAI | 67.8 | $0.024 | yes | |
| Nemotron 3 Ultra 550B A55B | NVIDIA | 67.4 | $0.37 | yes | |
| Minimax M3 | MiniMax | 67.3 | $0.060 | ||
| GPT-5.4 Mini xHigh | OpenAI | 66.4 | $0.33 | ||
| Qwen 3.6 27B | Alibaba | 64.0 | $0.20 | yes | |
| Gemini 3.5 Flash-Lite High | 63.9 | $0.069 | |||
| Grok 4.3 | xAI | 62.3 | $0.061 |
All leaderboards by category
Coding benchmarks
Code generation and web development tasks from Code Arena (Elo) and LiveBench.
2 boards: Code Arena, LiveBench Coding
View leaderboardsAgentic coding benchmarks
Multi-step code editing and tool use in agentic workflows, from LiveBench.
1 board: LiveBench Agentic Coding
View leaderboardsReasoning benchmarks
Logic, deduction and inference tasks from LiveBench, plus the LiveBench overall score.
2 boards: LiveBench Reasoning, LiveBench Overall
View leaderboardsMath benchmarks
Numerical reasoning and mathematical problem solving from LiveBench.
1 board: LiveBench Math
View leaderboardsData analysis benchmarks
Structured data interpretation, querying and analysis from LiveBench.
1 board: LiveBench Data Analysis
View leaderboardsLanguage benchmarks
Chat preference rankings (Text Arena Elo) and language comprehension (LiveBench).
2 boards: Text Arena, LiveBench Language
View leaderboardsInstruction following benchmarks
Adherence to formatting constraints and complex instructions, from LiveBench.
1 board: LiveBench Instruction Following
View leaderboards
Sources and method
I fetch three public leaderboards twice a day, at 06:00 and 18:00 UTC, and publish their numbers unchanged. Scores, prices and costs are the sources' own, not mine. I compute only the rank order of the seven LiveBench category tables, the cost and quality frontier, and the three picks marked as computed.
JSON API
Every category is available as JSON with open CORS, for example /api/benchmarks/coding
- Measures
- Human preference: people compare two anonymous answers to their own prompt and vote. Scores are Elo-style ratings with a confidence interval.
- Dates
- source date 2 Oct 2026 · fetched 3 Oct 2026 10:46 UTC · 413 models
- Licence
- Data by Arena (arena.ai), also published as an open dataset under CC BY 4.0.
- Measures
- Human preference on web development tasks, including agentic coding workflows, rated the same way.
- Dates
- source date 1 Oct 2026 · fetched 3 Oct 2026 10:46 UTC · 138 models
- Licence
- Data by Arena (arena.ai), also published as an open dataset under CC BY 4.0.
- Measures
- 23 objective tasks across 7 categories with verifiable answers, plus the cost per successful task. The questions are refreshed every six months.
- Dates
- release 2026-06-25 · fetched 3 Oct 2026 10:46 UTC · 58 models · New models are added between releases.
- Licence
- Data by LiveBench (livebench.ai) under CC BY-SA 4.0. The LiveBench tables on this site are shared under the same licence.
Need a model picked for your use case?
Benchmarks rank models on other people's tasks. For a business workflow I test the shortlist on your own data, costs and privacy requirements, and recommend one.