LMArena Document Arena
LMArena · document-2026-07-30-011508720696 · Blind human preferences for model-generated documents in LMArena.
Model Ranking
Source | |||||||
|---|---|---|---|---|---|---|---|
| #1 | Claude Opus 4.6Anthropic | 1,509.94 | — | 37,271 | claude-opus-4-6; 95% CI [1503.54038690, 1516.34687585]; votes 37271; rank 2unknown | Original sourcewinner eligible | Third-party benchmark |
| #2 | Claude Fable 5Anthropic | 1,504.25 | — | 5,792 | claude-fable-5; 95% CI [1495.07964874, 1513.41865883]; votes 5792; rank 4unknown | Original sourcewinner eligible | Third-party benchmark |
| #3 | Claude Opus 4.7Anthropic | 1,498.15 | — | 18,990 | claude-opus-4-7; 95% CI [1491.06893431, 1505.22120086]; votes 18990; rank 5unknown | Original sourcewinner eligible | Third-party benchmark |
| #4 | GPT-5.5OpenAI | 1,484.70 | — | 16,816 | gpt-5.5-high; 95% CI [1477.69349247, 1491.71461103]; votes 16816; rank 7unknown | Original sourcewinner eligible | Third-party benchmark |
| #5 | Claude Sonnet 4.6Anthropic | 1,482.80 | — | 54,427 | claude-sonnet-4-6; 95% CI [1476.49799837, 1489.09368537]; votes 54427; rank 8unknown | Original sourcewinner eligible | Third-party benchmark |
| #6 | GPT-5.6 TerraOpenAI | 1,479.21 | — | 1,969 | gpt-5.6-terra-xhigh; 95% CI [1466.19331424, 1492.21945897]; votes 1969; rank 10unknown | Original sourcewinner eligible | Third-party benchmark |
| #7 | GPT-5.6 SolOpenAI | 1,478.68 | — | 1,152 | gpt-5.6-sol-xhigh; 95% CI [1461.94144922, 1495.42273628]; votes 1152; rank 11unknown | Original sourcewinner eligible | Third-party benchmark |
| #8 | Claude Opus 4.8Anthropic | 1,474.59 | — | 8,388 | claude-opus-4-8-thinking; 95% CI [1466.45231317, 1482.72568328]; votes 8388; rank 12unknown | Original sourcewinner eligible | Third-party benchmark |
| #9 | Muse Spark 1.1Meta | 1,471.66 | — | 2,054 | muse-spark-1.1; 95% CI [1458.80838072, 1484.51367096]; votes 2054; rank 13unknown | Original sourcewinner eligible | Third-party benchmark |
| #10 | Claude Sonnet 5Anthropic | 1,470.22 | — | 3,601 | claude-sonnet-5-high; 95% CI [1459.84859101, 1480.59820001]; votes 3601; rank 14unknown | Original sourcewinner eligible | Third-party benchmark |
| #11 | GPT-5.4OpenAI | 1,470.14 | — | 29,809 | gpt-5.4; 95% CI [1463.54357191, 1476.73400825]; votes 29809; rank 15unknown | Original sourcewinner eligible | Third-party benchmark |
| #12 | Gemini 3.5 FlashGoogle DeepMind | 1,464.65 | — | 3,689 | gemini-3.5-flash-medium; 95% CI [1454.40953837, 1474.88391002]; votes 3689; rank 17unknown | Original sourcewinner eligible | Third-party benchmark |
| #13 | GPT-5.6 LunaOpenAI | 1,462.30 | — | 1,888 | gpt-5.6-luna-xhigh; 95% CI [1449.02133383, 1475.58666708]; votes 1888; rank 18unknown | Original sourcewinner eligible | Third-party benchmark |
| #14 | Claude Opus 4.5AnthropicSelected model | 1,461.65 | — | 7,965 | claude-opus-4-5-20251101; 95% CI [1451.22442739, 1472.08183979]; votes 7965; rank 19unknown | Original sourcewinner eligible | Third-party benchmark |
| #15 | Grok 4.5xAI | 1,453.73 | — | 2,435 | grok-4.5; 95% CI [1441.93744334, 1465.52785378]; votes 2435; rank 20unknown | Original sourcewinner eligible | Third-party benchmark |
| #16 | Kimi-K2.6Moonshot AI | 1,450.90 | — | 11,181 | kimi-k2.6; 95% CI [1443.21872745, 1458.58325377]; votes 11181; rank 21unknown | Original sourcewinner eligible | Third-party benchmark |
| #17 | Claude Sonnet 4.5Anthropic | 1,446.19 | — | 28,634 | claude-sonnet-4-5-20250929; 95% CI [1439.60004073, 1452.78970477]; votes 28634; rank 22unknown | Original sourcewinner eligible | Third-party benchmark |
| #18 | Gemini 3.1 ProGoogle DeepMind | 1,445.32 | — | 45,162 | gemini-3.1-pro-preview; 95% CI [1439.35538449, 1451.29339146]; votes 45162; rank 23unknown | Original sourcewinner eligible | Third-party benchmark |
| #19 | Qwen3.7-PlusQwen | 1,439.57 | — | 3,088 | qwen3.7-plus; 95% CI [1428.95936566, 1450.17849072]; votes 3088; rank 25unknown | Original sourcewinner eligible | Third-party benchmark |
| #20 | MiniMax-M3MiniMax | 1,434.82 | — | 6,263 | minimax-m3; 95% CI [1426.45297239, 1443.17975284]; votes 6263; rank 26unknown | Original sourcewinner eligible | Third-party benchmark |
| #21 | Kimi-K2.5Moonshot AI | 1,429.49 | — | 19,891 | kimi-k2.5-thinking; 95% CI [1422.63887395, 1436.33791058]; votes 19891; rank 28unknown | Original sourcewinner eligible | Third-party benchmark |
| #22 | Gemma 4 31BGoogle DeepMind | 1,423.70 | — | 10,577 | gemma-4-31b; 95% CI [1415.51297276, 1431.88143488]; votes 10577; rank 29unknown | Original sourcewinner eligible | Third-party benchmark |
| #23 | Claude Haiku 4.5Anthropic | 1,423.03 | — | 30,937 | claude-haiku-4-5-20251001; 95% CI [1416.59570790, 1429.45877492]; votes 30937; rank 30unknown | Original sourcewinner eligible | Third-party benchmark |
| #24 | Gemini 2.5 ProGoogle DeepMind | 1,420.87 | — | 24,963 | gemini-2.5-pro; 95% CI [1414.58328718, 1427.16017237]; votes 24963; rank 31unknown | Original sourcewinner eligible | Third-party benchmark |
| #25 | Gemini 3 FlashGoogle DeepMind | 1,413.32 | — | 7,173 | gemini-3-flash; 95% CI [1403.94387714, 1422.69054127]; votes 7173; rank 34unknown | Original sourcewinner eligible | Third-party benchmark |
| #26 | GPT-5.2OpenAI | 1,405.02 | — | 7,073 | gpt-5.2-high; 95% CI [1395.48938512, 1414.55044523]; votes 7073; rank 35unknown | Original sourcewinner eligible | Third-party benchmark |
| #27 | GPT-5.1OpenAI | 1,400.84 | — | 8,220 | gpt-5.1; 95% CI [1391.40144375, 1410.27736947]; votes 8220; rank 37unknown | Original sourcewinner eligible | Third-party benchmark |
27 models ranked by the best winner-eligible current-version score when available, otherwise the best publisher-artifact score; highest first. Observed Sep 3, 2026. The model you came from is highlighted.
Methodology and Coverage
Primary Evidence
Publisher Artifacts
Questions
LMArena Document Arena FAQs
What does LMArena Document Arena measure?+
Blind human preferences for model-generated documents in LMArena. Model Markets classifies it as a knowledge work benchmark and preserves the publisher's document-2026-07-30-011508720696 release as a distinct comparison cohort.
How are models ranked on LMArena Document Arena?+
Models are ordered by arena rating in rating, with higher scores ranked first. Each model appears once; a verified winner-eligible current-version result takes precedence over a publisher-artifact-only result, and missing scores are not estimated.
Which model currently leads LMArena Document Arena?+
Claude Opus 4.6 leads the current verified table with 1,509.94 rating on document-2026-07-30-011508720696. This is a benchmark-specific result, not a universal model-quality claim.
How many models have a published LMArena Document Arena score?+
27 catalog models are published from 38 source rows. 6 source identities remain quarantined rather than guessed.
Can LMArena Document Arena scores be compared with other benchmarks?+
Raw scores should be compared only within the same benchmark version, metric, and protocol. Model Markets normalizes eligible scores only for aggregate rankings, and groups models by an identical benchmark set before ranking them. View aggregate rankings →
Why might a model be missing from LMArena Document Arena?+
A model remains absent when the publisher has no current result, the source model identity is unresolved, the evaluation protocol is incompatible, or the evidence cannot be verified. Model Markets does not substitute a provider claim or infer a score from a related model.
How current is the LMArena Document Arena leaderboard?+
The current Model Markets snapshot was observed Sep 3, 2026 from LMArena artifacts. The exact publisher source and each retained result artifact are linked on this page.