Skip to content

/ LLM benchmarks / live

Live LLM benchmark: new models and ranking changes

What moved on Code Arena, Text Arena and LiveBench. Every model that entered a leaderboard, every new #1 and every overtake inside a top 10, read from the leaderboards twice a day.

Which model is best right now? The current leaders per job

Last crawl per source

The crawl is scheduled for 06:00 and 18:00 UTC and often finishes hours later. The table shows when each run fetched. The source date is the leaderboard's own update, which can be older than the crawl.

Last crawl per source
SourceFetchedSource dateModels
Text Arena2 Oct 2026413
Code Arena1 Oct 2026138
LiveBenchrelease 2026-06-2558

Twice a day, scheduled 06:00 and 18:00 UTC

New #1s

Every change at the top of a leaderboard in the last 90 days.

  1. Gemini 4 Argon High (Google) took #1 on Text Arena
  2. Claude Opus 5.5 High (Anthropic) took #1 on Text Arena
  3. Claude Opus 5.5 Max (Anthropic) took #1 on Code Arena
  4. Claude Fable 5 High (Anthropic) took #1 on Text Arena
  5. GPT-6 Astra Max (OpenAI) took #1 on Code Arena
  6. Claude Fable 5.1 Max (Anthropic) took #1 on Code Arena
  7. Claude Opus 5 Max (Anthropic) took #1 on Code Arena
  8. Kimi K3 (Moonshot AI) took #1 on Code Arena
  9. GPT-5.6 Sol xHigh (codex-harness) (OpenAI) took #1 on Code Arena (from #2)

Ranking changes, day by day

The last 30 days. A climb counts when a model overtakes another one inside a top 10 and its rank changes. Models that fell behind are counted, not listed.

  • GPT-6.1 Sol Max (OpenAI): Text Arena entered at #21
  • Claude Sonnet 5.5 xHigh (Anthropic): Text Arena entered at #45
  • Step 5 Preview (Stepfun): Text Arena entered at #74
  • Gemini 3.8 Flash High (Google): Text Arena climbed from #11 to #8
  • 3 models moved down

  • Claude Sonnet 5.5 xHigh (Anthropic): Code Arena entered at #3

  • Claude Opus 5.5 High (Anthropic): took #1 on Text Arena
  • MiMo V2.6 Pro (Xiaomi): Text Arena entered at #23
  • DeepSeek V4.1 Flash Max (DeepSeek): Text Arena entered at #29
  • GPT-6 Sol Max (OpenAI): Text Arena entered at #60
  • MiMo V2.6 Flash (Xiaomi): Text Arena entered at #71
  • GPT-6 Luna Max (OpenAI): Text Arena entered at #86
  • Grok 4.7 xHigh (SpaceXAI): Text Arena entered at #92
  • 2 models moved down

  • Step 5 Preview High (Stepfun): Code Arena entered at #29

  • Claude Opus 5.5 Max (Anthropic): took #1 on Code Arena
  • MiMo V2.6 Pro (Xiaomi): Code Arena entered at #19
  • GPT-6 Luna Max (OpenAI): Code Arena entered at #24
  • 1 model moved down

  • GPT-6 Sol Max (OpenAI): Code Arena entered at #4
  • 2 models moved down

  • Grok 4.7 xHigh (SpaceXAI): Code Arena entered at #10
  • Qwen 3.8 Max (Alibaba): Code Arena climbed from #6 to #4
  • Claude Opus 5 High (Anthropic): Code Arena climbed from #7 to #6
  • 3 models moved down

  • Muse Spark 1.2 (xHigh) (Meta): Text Arena climbed from #5 to #4
  • Claude Opus 4.7 (Anthropic): Text Arena climbed from #8 to #7
  • Gemini 3.8 Flash High (Google): Text Arena climbed from #11 to #9
  • 3 models moved down

  • Muse Spark 1.3 Max (Meta): Text Arena entered at #7

  • DeepSeek V4.1 Flash Max (DeepSeek): Code Arena entered at #16
  • GPT-6 Astra Max (OpenAI): Text Arena entered at #24
  • Claude Opus 4.7 High (Anthropic): Text Arena climbed from #4 to #3
  • Muse Spark 1.1 (Meta): Text Arena climbed from #10 to #8
  • 2 models moved down

  • Muse Spark 1.3 Max (Meta): Code Arena entered at #8
  • 2 models moved down

  • GPT-6 Astra Max (OpenAI): took #1 on Code Arena
  • Muse Spark 1.3 (xHigh) (Meta): Code Arena entered at #11
  • Claude Fable 5 (Anthropic): Code Arena climbed from #8 to #7
  • Qwen 3.8 Flash Next (Alibaba): Code Arena climbed from #10 to #9
  • 4 models moved down

Oldest recorded change: 9 Apr 2026.

Need a model picked for your use case?

Benchmarks rank models on other people's tasks. For a business workflow I test the shortlist on your own data, costs and privacy requirements, and recommend one.

AI consulting