LiveBench

LiveBench · 2026-06-25 · A contamination-resistant language-model benchmark with scores and token accounting published by the benchmark owner.

Model Ranking

Source
#1Claude Fable 5Anthropic86.55percent20,2551,270claude-fable-5-max-effortaverage per caseRecomputedwinner eligibleThird-party benchmark
#2GPT-5.6 SolOpenAI85.26percent11,7291,270gpt-5.6-sol-maxaverage per caseRecomputedwinner eligibleThird-party benchmark
#3Claude Opus 5Anthropic83.44percent14,8261,271claude-opus-5-max-effortaverage per caseRecomputedwinner eligibleThird-party benchmark
#4Gemini 3.7 FlashGoogle DeepMind83.15percent22,6101,270gemini-3.7-flash-highaverage per caseRecomputedwinner eligibleThird-party benchmark
#5GPT-5.6 TerraOpenAI82.31percent22,1451,270gpt-5.6-terra-maxaverage per caseRecomputedwinner eligibleThird-party benchmark
#6Grok 4.6xAI82.17percent16,1541,270grok-4.6average per caseRecomputedwinner eligibleThird-party benchmark
#7Qwen3.8-MaxQwen81.88percent21,6371,270qwen3.8-maxaverage per caseRecomputedwinner eligibleThird-party benchmark
#8Gemini 3.1 ProGoogle DeepMind81.66percent13,3801,270gemini-3.1-pro-preview-highaverage per caseRecomputedwinner eligibleThird-party benchmark
#9Claude Opus 4.8Anthropic81.18percent24,1711,270claude-opus-4-8-max-effortaverage per caseRecomputedwinner eligibleThird-party benchmark
#10Kimi-K3Moonshot AI81.02percent13,6471,270kimi-k3average per caseRecomputedwinner eligibleThird-party benchmark
#11DeepSeek-V4-Flash-Vision-ExpDeepSeek79.67percent52,6441,270deepseek-v4-flash-vision-expaverage per caseRecomputedwinner eligibleThird-party benchmark
#12GLM-5.3Z.ai79.15percent62,0901,270glm-5.3average per caseRecomputedwinner eligibleThird-party benchmark
#13Gemini 3.5 FlashGoogle DeepMind78.84percent15,5221,270gemini-3.5-flash-highaverage per caseRecomputedwinner eligibleThird-party benchmark
#14Gemini 3.6 FlashGoogle DeepMind78.05percent14,7461,270gemini-3.6-flash-highaverage per caseRecomputedwinner eligibleThird-party benchmark
#15Qwen3.8-27BQwen78.02percent28,7401,270qwen3.8-27baverage per caseRecomputedwinner eligibleThird-party benchmark
#16GPT-5.6 LunaOpenAI77.05percent21,7991,270gpt-5.6-luna-maxaverage per caseRecomputedwinner eligibleThird-party benchmark
#17GLM-5.2Z.ai76.96percent23,4631,270glm-5.2average per caseRecomputedwinner eligibleThird-party benchmark
#18DeepSeek-V4-ProDeepSeek76.79percent35,0141,270deepseek-v4-proaverage per caseRecomputedwinner eligibleThird-party benchmark
#19Kimi-K2.6Moonshot AI74.18percent27,0011,270kimi-k2.6-thinkingaverage per caseRecomputedwinner eligibleThird-party benchmark
#20GLM-5.3-FlashZ.aiSelected model73.27percent34,7071,270glm-5.3-flashaverage per caseRecomputedwinner eligibleThird-party benchmark
#21MiniMax-M3MiniMax70.26percent15,5301,270minimax-m3average per caseRecomputedwinner eligibleThird-party benchmark
#22DeepSeek-V4-FlashDeepSeek69.67percent34,4341,270deepseek-v4-flashaverage per caseRecomputedwinner eligibleThird-party benchmark
#23Gemini 3.5 Flash-LiteGoogle DeepMind66.35percent11,5261,270gemini-3.5-flash-lite-highaverage per caseRecomputedwinner eligibleThird-party benchmark

23 models ranked by the best winner-eligible current-version score when available, otherwise the best publisher-artifact score; highest first. Observed Sep 2, 2026. The model you came from is highlighted.

Methodology and Coverage

Ranking protocoloverall · percent. Each model appears once; verified winner-eligible runs take precedence, and missing models are not imputed.
Acquisition coverage23 published models from 49 source rows; 26 source identities remain quarantined.

Primary Evidence

Publisher Artifacts

Questions

LiveBench FAQs

What does LiveBench measure?+

A contamination-resistant language-model benchmark with scores and token accounting published by the benchmark owner. Model Markets classifies it as a general language benchmark and preserves the publisher's 2026-06-25 release as a distinct comparison cohort.

How are models ranked on LiveBench?+

Models are ordered by overall in percent, with higher scores ranked first. Each model appears once; a verified winner-eligible current-version result takes precedence over a publisher-artifact-only result, and missing scores are not estimated.

Which model currently leads LiveBench?+

Claude Fable 5 leads the current verified table with 86.55 percent on 2026-06-25. This is a benchmark-specific result, not a universal model-quality claim.

How many models have a published LiveBench score?+

23 catalog models are published from 49 source rows. 26 source identities remain quarantined rather than guessed.

Can LiveBench scores be compared with other benchmarks?+

Raw scores should be compared only within the same benchmark version, metric, and protocol. Model Markets normalizes eligible scores only for aggregate rankings, and groups models by an identical benchmark set before ranking them. View aggregate rankings

Why might a model be missing from LiveBench?+

A model remains absent when the publisher has no current result, the source model identity is unresolved, the evaluation protocol is incompatible, or the evidence cannot be verified. Model Markets does not substitute a provider claim or infer a score from a related model.

How current is the LiveBench leaderboard?+

The current Model Markets snapshot was observed Sep 2, 2026 from LiveBench artifacts. The exact publisher source and each retained result artifact are linked on this page.

Send Feedback