Skip to content

/ LLM benchmarks

LLM benchmarks: the best model for each job

The current leader for each task on Arena and LiveBench, with price or cost per task where the source reports one. Every score is the source's own.

DataText Arena2 Oct 2026Code Arena1 Oct 2026LiveBenchrelease 2026-06-25fetched

/ Best model for the job

  • Best value: Highest LiveBench overall score per dollar, among models within 5 points of the leader.
  • Best open-weights model: Highest LiveBench overall score among open-weights models.
  • Cheapest strong model: Lowest cost per successful task with at least 90% of the leader's overall score.

Cost vs quality

Each dot is a model on LiveBench: overall score against cost per successful task. The line is the Pareto frontier, where no other model is both cheaper and better.

ProprietaryOpen weightsPareto frontier

Swipe sideways to see the whole chart.

LiveBench overall score against cost per successful task57 models. On the frontier, from cheapest to best: Grok Build 0.1 ($0.024, 67.8%), GPT-6 Luna Max Effort ($0.026, 72.0%), DeepSeek V4.1 Flash Max Effort ($0.029, 81.1%), GPT-6.1 Sol Max Effort ($0.14, 81.6%), GPT-6 Astra Max Effort ($0.74, 82.2%), Claude 5.5 Opus Thinking Max Effort ($0.80, 83.2%), Claude Fable 5.1 Max Effort ($1.21, 83.4%).606570758085$0.02$0.05$0.1$0.2$0.5$1$2Cost per successful task (USD, log scale)LiveBench overall (%)Claude Fable 5.1 Max Effort (Anthropic): 83.4%, $1.21Claude 5.5 Opus Thinking Max Effort (Anthropic): 83.2%, $0.80Claude Fable 5 Max Effort (Anthropic): 83.0%, $1.44GPT-6 Astra Max Effort (OpenAI): 82.2%, $0.74GPT-6.1 Sol Max Effort (OpenAI): 81.6%, $0.14Muse Spark 1.3 xHigh Effort (Meta): 81.6%, $0.22DeepSeek V4.1 Flash Max Effort (DeepSeek): 81.1%, $0.029GPT-5.6 Sol Max Effort (OpenAI): 81.0%, $0.52GPT-5.5 Thinking xHigh Effort (OpenAI): 80.2%, $0.43Claude 5 Opus Thinking Max Effort (Anthropic): 80.1%, $0.70GPT-6 Sol Max Effort (OpenAI): 79.3%, $0.27Kimi K3 (Moonshot AI): 79.2%, $0.35Gemini 3.7 Flash High (Google): 78.8%, $0.16Qwen 3.8 Max (Alibaba): 78.5%, $0.28Grok 4.6 xHigh (xAI): 78.0%, $0.21GPT-5.4 Thinking xHigh Effort (OpenAI): 78.0%, $0.39Muse Spark 1.2 xHigh Effort (Meta): 78.0%, $0.38GPT-5.6 Terra Max Effort (OpenAI): 77.9%, $0.35Claude Sonnet 5.5 xHigh Effort (Anthropic): 77.8%, $0.14DeepSeek V4 Pro 0813 (DeepSeek): 77.4%, $0.044Grok 4.7 xHigh (xAI): 77.4%, $0.72Gemini 3.1 Pro Preview High (Google): 77.0%, $0.29DeepSeek V4 Flash Vision Exp (DeepSeek): 76.8%, $0.051Claude 4.7 Opus Thinking xHigh Effort (Anthropic): 76.5%, $0.53Claude 4.8 Opus Thinking Max Effort (Anthropic): 76.2%, $0.98Qwen 3.8 Flash Next (Alibaba): 76.2%, $0.042GLM-5.3 (Z.AI): 76.1%, $0.45Claude Sonnet 5 xHigh Effort (Anthropic): 76.0%, $0.51Gemini 3.8 Flash High (Google): 75.8%, $0.31Grok 4.5 (xAI): 75.8%, $0.13Muse Spark 1.1 xHigh Effort (Meta): 75.3%, $0.20Qwen3.8 27B (Alibaba): 75.3%, $0.094Gemini 3.5 Flash High (Google): 74.6%, $0.25GPT-5.2 High (OpenAI): 74.6%, $0.23Claude 4.6 Opus Thinking High Effort (Anthropic): 74.5%, $0.40DeepSeek V4 Flash 0731 (DeepSeek): 74.2%, $0.060GPT-5.2 Codex (OpenAI): 74.0%, $0.19Gemini 3.6 Flash High (Google): 73.6%, $0.23GPT-5.6 Luna Max Effort (OpenAI): 73.6%, $0.17GLM-5.2 (Z.AI): 73.2%, $0.23Qwen 3.7 Max (Alibaba): 73.1%, $0.18Claude 4.6 Sonnet Thinking Medium Effort (Anthropic): 73.0%, $0.31Claude 4.5 Opus Thinking High Effort (Anthropic): 72.6%, $0.61GPT-6 Luna Max Effort (OpenAI): 72.0%, $0.026Inkling xHigh Effort (Thinking Machines): 71.9%, $0.31GLM-5.3 Flash (Z.AI): 71.6%, $0.031Kimi K2.6 Thinking (Moonshot AI): 70.5%, $0.17GPT-5.4 Nano xHigh (OpenAI): 69.6%, $0.091Qwen 3.6 Plus (Alibaba): 68.9%, $0.23Kimi K2.7 Code (Moonshot AI): 68.4%, $0.10Grok Build 0.1 (xAI): 67.8%, $0.024Nemotron 3 Ultra 550B A55B (NVIDIA): 67.4%, $0.37Minimax M3 (MiniMax): 67.3%, $0.060GPT-5.4 Mini xHigh (OpenAI): 66.4%, $0.33Qwen 3.6 27B (Alibaba): 64.0%, $0.20Gemini 3.5 Flash-Lite High (Google): 63.9%, $0.069Grok 4.3 (xAI): 62.3%, $0.061Claude Fable 5.1 Max EffortClaude 5.5 Opus Thinking Max EffortClaude Fable 5 Max EffortGPT-6 Astra Max EffortGPT-6.1 Sol Max EffortMuse Spark 1.3 xHigh EffortDeepSeek V4.1 Flash Max EffortGPT-5.6 Sol Max Effort

57 models. On the frontier, from cheapest to best: Grok Build 0.1 ($0.024, 67.8%), GPT-6 Luna Max Effort ($0.026, 72.0%), DeepSeek V4.1 Flash Max Effort ($0.029, 81.1%), GPT-6.1 Sol Max Effort ($0.14, 81.6%), GPT-6 Astra Max Effort ($0.74, 82.2%), Claude 5.5 Opus Thinking Max Effort ($0.80, 83.2%), Claude Fable 5.1 Max Effort ($1.21, 83.4%).

1 model without a cost figure is left out.
Show the chart data as a table
LiveBench overall score against cost per successful task
ModelOrganizationOverallCost/solved taskOpen weightsFrontier
Claude Fable 5.1 Max EffortAnthropic83.4$1.21yes
Claude 5.5 Opus Thinking Max EffortAnthropic83.2$0.80yes
Claude Fable 5 Max EffortAnthropic83.0$1.44
GPT-6 Astra Max EffortOpenAI82.2$0.74yes
GPT-6.1 Sol Max EffortOpenAI81.6$0.14yes
Muse Spark 1.3 xHigh EffortMeta81.6$0.22
DeepSeek V4.1 Flash Max EffortDeepSeek81.1$0.029yesyes
GPT-5.6 Sol Max EffortOpenAI81.0$0.52
GPT-5.5 Thinking xHigh EffortOpenAI80.2$0.43
Claude 5 Opus Thinking Max EffortAnthropic80.1$0.70
GPT-6 Sol Max EffortOpenAI79.3$0.27
Kimi K3Moonshot AI79.2$0.35yes
Gemini 3.7 Flash HighGoogle78.8$0.16
Qwen 3.8 MaxAlibaba78.5$0.28yes
Grok 4.6 xHighxAI78.0$0.21
GPT-5.4 Thinking xHigh EffortOpenAI78.0$0.39
Muse Spark 1.2 xHigh EffortMeta78.0$0.38
GPT-5.6 Terra Max EffortOpenAI77.9$0.35
Claude Sonnet 5.5 xHigh EffortAnthropic77.8$0.14
DeepSeek V4 Pro 0813DeepSeek77.4$0.044yes
Grok 4.7 xHighxAI77.4$0.72
Gemini 3.1 Pro Preview HighGoogle77.0$0.29
DeepSeek V4 Flash Vision ExpDeepSeek76.8$0.051yes
Claude 4.7 Opus Thinking xHigh EffortAnthropic76.5$0.53
Claude 4.8 Opus Thinking Max EffortAnthropic76.2$0.98
Qwen 3.8 Flash NextAlibaba76.2$0.042yes
GLM-5.3Z.AI76.1$0.45yes
Claude Sonnet 5 xHigh EffortAnthropic76.0$0.51
Gemini 3.8 Flash HighGoogle75.8$0.31
Grok 4.5xAI75.8$0.13
Muse Spark 1.1 xHigh EffortMeta75.3$0.20
Qwen3.8 27BAlibaba75.3$0.094yes
Gemini 3.5 Flash HighGoogle74.6$0.25
GPT-5.2 HighOpenAI74.6$0.23
Claude 4.6 Opus Thinking High EffortAnthropic74.5$0.40
DeepSeek V4 Flash 0731DeepSeek74.2$0.060yes
GPT-5.2 CodexOpenAI74.0$0.19
Gemini 3.6 Flash HighGoogle73.6$0.23
GPT-5.6 Luna Max EffortOpenAI73.6$0.17
GLM-5.2Z.AI73.2$0.23yes
Qwen 3.7 MaxAlibaba73.1$0.18
Claude 4.6 Sonnet Thinking Medium EffortAnthropic73.0$0.31
Claude 4.5 Opus Thinking High EffortAnthropic72.6$0.61
GPT-6 Luna Max EffortOpenAI72.0$0.026yes
Inkling xHigh EffortThinking Machines71.9$0.31yes
GLM-5.3 FlashZ.AI71.6$0.031yes
Kimi K2.6 ThinkingMoonshot AI70.5$0.17yes
GPT-5.4 Nano xHighOpenAI69.6$0.091
Qwen 3.6 PlusAlibaba68.9$0.23
Kimi K2.7 CodeMoonshot AI68.4$0.10yes
Grok Build 0.1xAI67.8$0.024yes
Nemotron 3 Ultra 550B A55BNVIDIA67.4$0.37yes
Minimax M3MiniMax67.3$0.060
GPT-5.4 Mini xHighOpenAI66.4$0.33
Qwen 3.6 27BAlibaba64.0$0.20yes
Gemini 3.5 Flash-Lite HighGoogle63.9$0.069
Grok 4.3xAI62.3$0.061

Sources and method

I fetch three public leaderboards twice a day, at 06:00 and 18:00 UTC, and publish their numbers unchanged. Scores, prices and costs are the sources' own, not mine. I compute only the rank order of the seven LiveBench category tables, the cost and quality frontier, and the three picks marked as computed.

JSON API

Every category is available as JSON with open CORS, for example /api/benchmarks/coding

  • Measures
    Human preference: people compare two anonymous answers to their own prompt and vote. Scores are Elo-style ratings with a confidence interval.
    Dates
    source date 2 Oct 2026 · fetched 3 Oct 2026 10:46 UTC · 413 models
    Licence
    Data by Arena (arena.ai), also published as an open dataset under CC BY 4.0.
  • Measures
    Human preference on web development tasks, including agentic coding workflows, rated the same way.
    Dates
    source date 1 Oct 2026 · fetched 3 Oct 2026 10:46 UTC · 138 models
    Licence
    Data by Arena (arena.ai), also published as an open dataset under CC BY 4.0.
  • Measures
    23 objective tasks across 7 categories with verifiable answers, plus the cost per successful task. The questions are refreshed every six months.
    Dates
    release 2026-06-25 · fetched 3 Oct 2026 10:46 UTC · 58 models · New models are added between releases.
    Licence
    Data by LiveBench (livebench.ai) under CC BY-SA 4.0. The LiveBench tables on this site are shared under the same licence.

Need a model picked for your use case?

Benchmarks rank models on other people's tasks. For a business workflow I test the shortlist on your own data, costs and privacy requirements, and recommend one.

AI consulting