Zum Inhalt springen

/ LLM-Benchmarks / Mathematik

KI Mathe: welches Modell am meisten löst

Numerisches Denken und mathematische Problemlösung aus LiveBench.

DatenLiveBenchRelease 2026-06-25abgerufen

Leaderboard LiveBench Math, 58 Modelle
#ModellWertungKosten pro gelöster Aufgabe
1
Claude 5.5 Opus Thinking Max EffortAnthropic
97.1%
$0.80
2
Claude Fable 5.1 Max EffortAnthropic
97.0%
$1.21
3
GPT-6 Astra Max EffortOpenAI
96.8%
$0.74
4
GPT-6.1 Sol Max EffortOpenAI
96.8%
$0.14
5
Claude Sonnet 5.5 xHigh EffortAnthropic
96.7%
$0.14
6
GPT-6 Sol Max EffortOpenAI
96.4%
$0.27
7
GPT-5.6 Sol Max EffortOpenAI
96.2%
$0.52
8
Claude Fable 5 Max EffortAnthropic
96.0%
$1.44
9
Muse Spark 1.3 xHigh EffortMeta
95.9%
$0.22
10
GPT-5.5 Thinking xHigh EffortOpenAI
95.9%
$0.43
11
Claude 5 Opus Thinking Max EffortAnthropic
95.7%
$0.70
12
Grok 4.7 xHighxAI
95.7%
$0.72
13
DeepSeek V4 Pro 0813DeepSeek · Open Weights
95.1%
$0.044
14
GPT-5.6 Terra Max EffortOpenAI
94.9%
$0.35
15
Claude 4.8 Opus Thinking Max EffortAnthropic
94.3%
$0.98
16
GPT-5.4 Thinking xHigh EffortOpenAI
94.1%
$0.39
17
Gemini 3.7 Flash HighGoogle
93.5%
$0.16
18
DeepSeek V4.1 Flash Max EffortDeepSeek · Open Weights
93.3%
$0.029
19
GPT-5.2 HighOpenAI
93.2%
$0.23
20
Claude 4.7 Opus Thinking xHigh EffortAnthropic
92.9%
$0.53
21
Claude Sonnet 5 xHigh EffortAnthropic
92.9%
$0.51
22
Grok 4.6 xHighxAI
92.6%
$0.21
23
Gemini 3.8 Flash HighGoogle
91.6%
$0.31
24
Qwen 3.8 MaxAlibaba · Open Weights
91.3%
$0.28
25
Muse Spark 1.2 xHigh EffortMeta
91.2%
$0.38
26
Gemini 3.1 Pro Preview HighGoogle
91.0%
$0.29
27
GPT-5.4 Nano xHighOpenAI
91.0%
$0.091
28
Grok 4.5xAI
90.8%
$0.13
29
Claude 4.5 Opus Thinking High EffortAnthropic
90.4%
$0.61
30
GLM-5.2Z.AI · Open Weights
89.8%
$0.23
31
Claude 4.6 Opus Thinking High EffortAnthropic
89.3%
$0.40
32
GPT-6 Luna Max EffortOpenAI
89.1%
$0.026
33
GPT-5.2 CodexOpenAI
88.8%
$0.19
34
Nemotron 3 Ultra 550B A55BNVIDIA · Open Weights
88.7%
$0.37
35
Inkling xHigh EffortThinking Machines · Open Weights
88.4%
$0.31
36
Gemini 3.5 Flash HighGoogle
88.2%
$0.25
37
GLM-5.3Z.AI · Open Weights
87.9%
$0.45
38
DeepSeek V4 Flash Vision ExpDeepSeek · Open Weights
87.8%
$0.051
39
GPT-5.6 Luna Max EffortOpenAI
87.2%
$0.17
40
Muse Spark 1.1 xHigh EffortMeta
87.1%
$0.20
41
Claude 4.6 Sonnet Thinking Medium EffortAnthropic
87.0%
$0.31
42
DeepSeek V4 Flash 0731DeepSeek · Open Weights
86.8%
$0.060
43
Gemini 3.6 Flash HighGoogle
86.4%
$0.23
44
Qwen3.8 27BAlibaba · Open Weights
86.2%
$0.094
45
Qwen 3.8 Flash NextAlibaba · Open Weights
85.8%
$0.042
46
Qwen 3.7 MaxAlibaba
85.2%
$0.18
47
Kimi K3Moonshot AI · Open Weights
84.4%
$0.35
48
Kimi K2.6 ThinkingMoonshot AI · Open Weights
84.3%
$0.17
49
Grok 4.3xAI
84.3%
$0.061
50
Qwen 3.6 PlusAlibaba
83.7%
$0.23
51
GLM-5.3 FlashZ.AI · Open Weights
81.2%
$0.031
52
Qwen 3.6 27BAlibaba · Open Weights
79.9%
$0.20
53
Kimi K2.7 CodeMoonshot AI · Open Weights
79.6%
$0.10
54
GPT-5.4 Mini xHighOpenAI
78.5%
$0.33
55
Grok Build 0.1xAI
78.4%
$0.024
56
Ox Alpha MaxStealth
77.5%
–
57
Minimax M3MiniMax
76.9%
$0.060
58
Gemini 3.5 Flash-Lite HighGoogle
73.7%
$0.069

Sie brauchen ein Modell für Ihren konkreten Fall?

Benchmarks bewerten Modelle anhand fremder Aufgaben. Für einen Geschäftsprozess teste ich die engere Wahl mit Ihren eigenen Daten, Kosten und Datenschutzvorgaben und empfehle eines.

KI-Beratung