LiveBench
LiveBench · 2026-06-25 · A contamination-resistant language-model benchmark with scores and token accounting published by the benchmark owner.
Model Ranking
Source | |||||||
|---|---|---|---|---|---|---|---|
| #1 | Claude Fable 5Anthropic | 86.55 | 20,255 | 1,270 | claude-fable-5-max-effortaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #2 | GPT-5.6 SolOpenAI | 85.26 | 11,729 | 1,270 | gpt-5.6-sol-maxaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #3 | Claude Opus 5Anthropic | 83.44 | 14,826 | 1,271 | claude-opus-5-max-effortaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #4 | Gemini 3.7 FlashGoogle DeepMind | 83.15 | 22,610 | 1,270 | gemini-3.7-flash-highaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #5 | GPT-5.6 TerraOpenAI | 82.31 | 22,145 | 1,270 | gpt-5.6-terra-maxaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #6 | Grok 4.6xAI | 82.17 | 16,154 | 1,270 | grok-4.6average per case | Recomputedwinner eligible | Third-party benchmark |
| #7 | Qwen3.8-MaxQwen | 81.88 | 21,637 | 1,270 | qwen3.8-maxaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #8 | Gemini 3.1 ProGoogle DeepMind | 81.66 | 13,380 | 1,270 | gemini-3.1-pro-preview-highaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #9 | Claude Opus 4.8AnthropicSelected model | 81.18 | 24,171 | 1,270 | claude-opus-4-8-max-effortaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #10 | Kimi-K3Moonshot AI | 81.02 | 13,647 | 1,270 | kimi-k3average per case | Recomputedwinner eligible | Third-party benchmark |
| #11 | DeepSeek-V4-Flash-Vision-ExpDeepSeek | 79.67 | 52,644 | 1,270 | deepseek-v4-flash-vision-expaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #12 | GLM-5.3Z.ai | 79.15 | 62,090 | 1,270 | glm-5.3average per case | Recomputedwinner eligible | Third-party benchmark |
| #13 | Gemini 3.5 FlashGoogle DeepMind | 78.84 | 15,522 | 1,270 | gemini-3.5-flash-highaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #14 | Gemini 3.6 FlashGoogle DeepMind | 78.05 | 14,746 | 1,270 | gemini-3.6-flash-highaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #15 | Qwen3.8-27BQwen | 78.02 | 28,740 | 1,270 | qwen3.8-27baverage per case | Recomputedwinner eligible | Third-party benchmark |
| #16 | GPT-5.6 LunaOpenAI | 77.05 | 21,799 | 1,270 | gpt-5.6-luna-maxaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #17 | GLM-5.2Z.ai | 76.96 | 23,463 | 1,270 | glm-5.2average per case | Recomputedwinner eligible | Third-party benchmark |
| #18 | DeepSeek-V4-ProDeepSeek | 76.79 | 35,014 | 1,270 | deepseek-v4-proaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #19 | Kimi-K2.6Moonshot AI | 74.18 | 27,001 | 1,270 | kimi-k2.6-thinkingaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #20 | GLM-5.3-FlashZ.ai | 73.27 | 34,707 | 1,270 | glm-5.3-flashaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #21 | MiniMax-M3MiniMax | 70.26 | 15,530 | 1,270 | minimax-m3average per case | Recomputedwinner eligible | Third-party benchmark |
| #22 | DeepSeek-V4-FlashDeepSeek | 69.67 | 34,434 | 1,270 | deepseek-v4-flashaverage per case | Recomputedwinner eligible | Third-party benchmark |
| #23 | Gemini 3.5 Flash-LiteGoogle DeepMind | 66.35 | 11,526 | 1,270 | gemini-3.5-flash-lite-highaverage per case | Recomputedwinner eligible | Third-party benchmark |
23 models ranked by the best winner-eligible current-version score when available, otherwise the best publisher-artifact score; highest first. Observed Sep 2, 2026. The model you came from is highlighted.
Methodology and Coverage
Primary Evidence
Publisher Artifacts
Questions
LiveBench FAQs
What does LiveBench measure?+
A contamination-resistant language-model benchmark with scores and token accounting published by the benchmark owner. Model Markets classifies it as a general language benchmark and preserves the publisher's 2026-06-25 release as a distinct comparison cohort.
How are models ranked on LiveBench?+
Models are ordered by overall in percent, with higher scores ranked first. Each model appears once; a verified winner-eligible current-version result takes precedence over a publisher-artifact-only result, and missing scores are not estimated.
Which model currently leads LiveBench?+
Claude Fable 5 leads the current verified table with 86.55 percent on 2026-06-25. This is a benchmark-specific result, not a universal model-quality claim.
How many models have a published LiveBench score?+
23 catalog models are published from 49 source rows. 26 source identities remain quarantined rather than guessed.
Can LiveBench scores be compared with other benchmarks?+
Raw scores should be compared only within the same benchmark version, metric, and protocol. Model Markets normalizes eligible scores only for aggregate rankings, and groups models by an identical benchmark set before ranking them. View aggregate rankings →
Why might a model be missing from LiveBench?+
A model remains absent when the publisher has no current result, the source model identity is unresolved, the evaluation protocol is incompatible, or the evidence cannot be verified. Model Markets does not substitute a provider claim or infer a score from a related model.
How current is the LiveBench leaderboard?+
The current Model Markets snapshot was observed Sep 2, 2026 from LiveBench artifacts. The exact publisher source and each retained result artifact are linked on this page.