ToneBench

Towards AI · 2026-08-28-10-task-cd9819ab6e4d · Blind, panel-scored writing benchmark across a fixed set of real editorial scripts.

Model Ranking

Source
#1Claude Fable 5Anthropic89.30points2,84750Claude Fable 5 (max effort + 4.8 fallback)average per caseOriginal sourcewinner eligibleThird-party benchmark
#2Claude Opus 5Anthropic88.76points2,62850Claude Opus 5 (max effort)average per caseOriginal sourcewinner eligibleThird-party benchmark
#3Kimi-K3Moonshot AI88.19points15,02350Kimi K3average per caseOriginal sourcewinner eligibleThird-party benchmark
#4GPT-5.6 SolOpenAI87.97points2,89450GPT-5.6 Sol (ultra)average per caseOriginal sourcewinner eligibleThird-party benchmark
#5Claude Opus 4.8Anthropic87.91points2,74150Claude Opus 4.8 (max effort)average per caseOriginal sourcewinner eligibleThird-party benchmark
#6Grok 4.6xAI86.61points15,57550Grok 4.6average per caseOriginal sourcewinner eligibleThird-party benchmark
#7GPT-5.6 TerraOpenAI86.39points2,64250GPT-5.6 Terra (ultra)average per caseOriginal sourcewinner eligibleThird-party benchmark
#8GPT-5.4OpenAI85.60points2,99850GPT-5.4 (xhigh)average per caseOriginal sourcewinner eligibleThird-party benchmark
#9Claude Opus 4.7Anthropic85.49points2,57850Claude Opus 4.7average per caseOriginal sourcewinner eligibleThird-party benchmark
#10GPT-5.6 LunaOpenAI84.47points2,94850GPT-5.6 Luna (ultra)average per caseOriginal sourcewinner eligibleThird-party benchmark
#11Claude Opus 4.6Anthropic84.21points2,60750Claude Opus 4.6average per caseOriginal sourcewinner eligibleThird-party benchmark
#12DeepSeek-V4-Pro-0813DeepSeek84.04points18,92250DeepSeek V4 Pro 0813 (max)average per caseOriginal sourcewinner eligibleThird-party benchmark
#13Claude Sonnet 4.6Anthropic83.63points2,52550Claude Sonnet 4.6average per caseOriginal sourcewinner eligibleThird-party benchmark
#14GPT-5.5OpenAISelected model82.80points3,51750GPT-5.5 (high)average per caseOriginal sourcewinner eligibleThird-party benchmark
#15Grok 4.5xAI82.23points2,68850Grok 4.5average per caseOriginal sourcewinner eligibleThird-party benchmark
#16GPT-5.4 MiniOpenAI82.16points2,49250GPT-5.4 mini (high)average per caseOriginal sourcewinner eligibleThird-party benchmark
#17Qwen3.7 MaxQwen81.26points8,63650Qwen3.7 Max (default)average per caseOriginal sourcewinner eligibleThird-party benchmark
#18Qwen3.8-MaxQwen81.06points23,33450Qwen3.8 Maxaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#19Muse Spark 1.1Meta80.76points6,03750Muse Spark 1.1 (thinking)average per caseOriginal sourcewinner eligibleThird-party benchmark
#20Gemini 3.1 ProGoogle DeepMind80.45points9,80150Gemini 3.1 Pro (default)average per caseOriginal sourcewinner eligibleThird-party benchmark
#21InklingThinking Machines Lab79.41points5,78550Inklingaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#22Gemini 3.7 FlashGoogle DeepMind79.18points5,19650Gemini 3.7 Flash (default)average per caseOriginal sourcewinner eligibleThird-party benchmark
#23Gemini 2.5 ProGoogle DeepMind78.27points4,52150Gemini 2.5 Proaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#24Gemini 3.6 FlashGoogle DeepMind78.25points10,92950Gemini 3.6 Flash (high thinking)average per caseOriginal sourcewinner eligibleThird-party benchmark
#25GPT-5.2OpenAI78.08points4,26650GPT-5.2average per caseOriginal sourcewinner eligibleThird-party benchmark
#26GPT-5OpenAI77.82points5,89650GPT-5average per caseOriginal sourcewinner eligibleThird-party benchmark
#27Claude Sonnet 4.5Anthropic77.25points2,44150Claude Sonnet 4.5average per caseOriginal sourcewinner eligibleThird-party benchmark
#28Gemini 3.5 FlashGoogle DeepMind76.21points14,25250Gemini 3.5 Flash (default)average per caseOriginal sourcewinner eligibleThird-party benchmark
#29GPT-5 MiniOpenAI76.07points3,92550GPT-5 miniaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#30GPT-5.1OpenAI75.98points3,97050GPT-5.1average per caseOriginal sourcewinner eligibleThird-party benchmark
#31Grok 4.3xAI74.91points5,15450Grok 4.3 (high)average per caseOriginal sourcewinner eligibleThird-party benchmark
#32Gemini 3.5 Flash-LiteGoogle DeepMind73.42points1,96650Gemini 3.5 Flash-Liteaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#33Gemini 2.5 FlashGoogle DeepMind72.17points4,39950Gemini 2.5 Flashaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#34Gemini 3.1 Flash-LiteGoogle DeepMind71.30points1,62850Gemini 3.1 Flash-Liteaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#35Mistral Large 3Mistral AI70.39points2,05650Mistral Large 3average per caseOriginal sourcewinner eligibleThird-party benchmark
#36GPT-5.4 NanoOpenAI69.12points3,83150GPT-5.4 nanoaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#37Gemini 2.5 Flash-LiteGoogle DeepMind68.76points2,36050Gemini 2.5 Flash-Liteaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#38Command A ReasoningCohere62.90points2,00250Cohere Command A Reasoningaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#39GPT-5 NanoOpenAI61.62points7,46550GPT-5 nanoaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#40Command ACohere60.90points1,17750Cohere Command A (03-2025)average per caseOriginal sourcewinner eligibleThird-party benchmark
#41GPT-4oOpenAI58.66points1,17450GPT-4oaverage per caseOriginal sourcewinner eligibleThird-party benchmark
#42GPT-4o MiniOpenAI58.13points1,34750GPT-4o miniaverage per caseOriginal sourcewinner eligibleThird-party benchmark

42 models ranked by the best winner-eligible current-version score when available, otherwise the best publisher-artifact score; highest first. Observed Sep 3, 2026. The model you came from is highlighted.

Methodology and Coverage

Ranking protocoloverall score · points. Each model appears once; verified winner-eligible runs take precedence, and missing models are not imputed.
Acquisition coverage42 published models from 138 source rows; 53 source identities remain quarantined.

Primary Evidence

Publisher Artifacts

Questions

ToneBench FAQs

What does ToneBench measure?+

Blind, panel-scored writing benchmark across a fixed set of real editorial scripts. Model Markets classifies it as a writing benchmark and preserves the publisher's 2026-08-28-10-task-cd9819ab6e4d release as a distinct comparison cohort.

How are models ranked on ToneBench?+

Models are ordered by overall score in points, with higher scores ranked first. Each model appears once; a verified winner-eligible current-version result takes precedence over a publisher-artifact-only result, and missing scores are not estimated.

Which model currently leads ToneBench?+

Claude Fable 5 leads the current verified table with 89.30 points on 2026-08-28-10-task-cd9819ab6e4d. This is a benchmark-specific result, not a universal model-quality claim.

How many models have a published ToneBench score?+

42 catalog models are published from 138 source rows. 53 source identities remain quarantined rather than guessed.

Can ToneBench scores be compared with other benchmarks?+

Raw scores should be compared only within the same benchmark version, metric, and protocol. Model Markets normalizes eligible scores only for aggregate rankings, and groups models by an identical benchmark set before ranking them. View aggregate rankings

Why might a model be missing from ToneBench?+

A model remains absent when the publisher has no current result, the source model identity is unresolved, the evaluation protocol is incompatible, or the evidence cannot be verified. Model Markets does not substitute a provider claim or infer a score from a related model.

How current is the ToneBench leaderboard?+

The current Model Markets snapshot was observed Sep 3, 2026 from Towards AI artifacts. The exact publisher source and each retained result artifact are linked on this page.

Send Feedback