Claude Sonnet 4.6 vs GPT-5.1

Benchmark Performance

Available Benchmarks

BenchmarkClaude Sonnet 4.6GPT-5.1
ARC-AGI-1verified-v1-15fb467fd4fc · verified_score · leader86.50100% of row best · percent · Claude Sonnet 4.6 (High)72.8384% of row best · percent · GPT-5.1 (Thinking, High)
ARC-AGI-2verified-v2-9a3db289984e · verified_score · leader60.42100% of row best · percent · Claude Sonnet 4.6 (High)17.6429% of row best · percent · GPT-5.1 (Thinking, High)
LMArena Document Arenadocument-2026-07-30-011508720696 · arena_rating · leader1,482.80100% of row best · rating · claude-sonnet-4-6; 95% CI [1476.49799837, 1489.09368537]; votes 54427; rank 81,400.8494% of row best · rating · gpt-5.1; 95% CI [1391.40144375, 1410.27736947]; votes 8220; rank 37
LMArena Search Arenasearch-2026-08-24-011508720696 · arena_rating · leader1,221.23100% of row best · rating · claude-sonnet-4-6-search; 95% CI [1216.39472071, 1226.06821163]; votes 134905; rank 71,199.4498% of row best · rating · gpt-5.1-search; 95% CI [1194.20592738, 1204.68165473]; votes 59909; rank 14
LMArena Text Arenatext-2026-09-01-011508720696 · arena_rating · leader1,458.01100% of row best · rating · claude-sonnet-4-6; 95% CI [1454.41608210, 1461.60383515]; votes 66343; rank 321,422.4798% of row best · rating · gpt-5.1; 95% CI [1418.85253526, 1426.08681113]; votes 43002; rank 96
LMArena Vision Arenavision-2026-08-27-011508720696 · arena_rating · leader1,282.25100% of row best · rating · claude-sonnet-4-6; 95% CI [1275.84615489, 1288.65223369]; votes 25610; rank 251,235.1996% of row best · rating · gpt-5.1; 95% CI [1227.44529223, 1242.93837299]; votes 10223; rank 60
ToneBench2026-08-28-10-task-cd9819ab6e4d · overall_score · leader83.63100% of row best · points · Claude Sonnet 4.6 · 2,525 output tokens / case75.9891% of row best · points · GPT-5.1 · 3,970 output tokens / case
Overall ResultCounted from the protocol-matched rows above7 benchmark winsOverall lead0 benchmark wins

Third-party benchmark Only like-for-like primary-publisher results are shown; raw scores, relative scores, configuration, and token spend remain visible.

FieldAt a Glance
Anthropic · activeClaude Sonnet 4.6Verified Sep 3, 2026
Openai · activeGPT-5.1Verified Sep 3, 2026
Not a Valid ComparisonThese records do not share a sourced modality and workload.

Technical Differences

Side-by-Side Facts

Interactive only
FieldClaude Sonnet 4.6GPT-5.1
DeveloperAnthropicOpenai
FamilyClaude 4Gpt 5 1
ModelClaude Sonnet 4.6GPT-5.1
VersionClaude Sonnet 4.6GPT-5.1
Lifecycleactiveactive
Released2026-02-172025-11-13
Knowledge cutoff2025-08-012024-09-30
Input modalitiesUnknownUnknown
Output modalitiesUnknownUnknown
Context window1,000,000400,000
Total parametersUnknownUnknown
Active parametersUnknownUnknown
LicenseUnknownUnknown
Open weightsNoNo
API availableYesYes
Self-hostableNoNo
Provider accessAnthropic (Standard), Deepinfra (Standard)Openai (Standard), Openrouter (Standard)
Capabilitieschat, generation, reasoning, structured_outputs, toolschat, generation, reasoning, structured_outputs, tools

13 comparable fields · 8 material differences · Interactive comparison only; indexing gate not met

Claude Sonnet 4.6 Capabilities

chatgenerationreasoningstructured outputstools
Input price$3.00
Output price$15.00
Serving providers2
Canonical IDanthropic/claude-sonnet-4-6

GPT-5.1 Capabilities

chatgenerationreasoningstructured outputstools
Input price$1.25
Output price$10.00
Serving providers2
Canonical IDopenai/gpt-5.1

Internal Comparison Graph

Related Comparisons

APairBContext
vsfamily variantsimage, text
vsfamily variantsimage, text
vsfamily variantsimage, text
vsfamily variantsimage, text
vsfamily variantsimage, text
vsfamily variantsimage, text
vsfamily variantsimage, text
vsfamily variantsimage, text
vscross-developer peersimage, text
vscross-developer peersimage, text
vscross-developer peersimage, text
vscross-developer peersimage, text

Primary Evidence

Sources and Freshness

Questions

Claude Sonnet 4.6 vs GPT-5.1 FAQs

Is Claude Sonnet 4.6 or GPT-5.1 better for coding?+

This comparison does not currently contain a protocol-matched coding benchmark for both Claude Sonnet 4.6 and GPT-5.1, so Model Markets cannot name a coding leader from pricing, context size, or capability labels alone.

Which is cheaper, Claude Sonnet 4.6 or GPT-5.1?+

Claude Sonnet 4.6 is $3.00 and GPT-5.1 is $1.25 per million tokens, so GPT-5.1 is cheaper on this metric. Claude Sonnet 4.6 is $15.00 and GPT-5.1 is $10.00 per million tokens, so GPT-5.1 is cheaper on this metric.

Which has a larger context window, Claude Sonnet 4.6 or GPT-5.1?+

Claude Sonnet 4.6 has the larger sourced context window. Claude Sonnet 4.6 supports 1,000,000 and GPT-5.1 supports 400,000.

Which performs better in benchmarks, Claude Sonnet 4.6 or GPT-5.1?+

Claude Sonnet 4.6 leads the current overall benchmark count. The result uses 7 protocol-matched benchmarks from 3 publishers; it is not a universal quality score.

Can Claude Sonnet 4.6 or GPT-5.1 be self-hosted?+

Both models have the same recorded self-hosting status: unsupported. Claude Sonnet 4.6 is not marked open weight; GPT-5.1 is not marked open weight.

Can Claude Sonnet 4.6 and GPT-5.1 understand images?+

Claude Sonnet 4.6 is not documented with image input; GPT-5.1 is not documented with image input. This reflects supported input modalities, not vision quality.

Which can generate longer answers, Claude Sonnet 4.6 or GPT-5.1?+

Neither has a larger sourced maximum output. Claude Sonnet 4.6 is 128,000 and GPT-5.1 is 128,000.

Do Claude Sonnet 4.6 and GPT-5.1 support reasoning and tool use?+

Claude Sonnet 4.6: reasoning and tool calling. GPT-5.1: reasoning and tool calling. Feature support does not establish relative quality.

Which is available from more inference providers, Claude Sonnet 4.6 or GPT-5.1?+

Claude Sonnet 4.6 has 2 sourced provider routes; GPT-5.1 has 2, a tie.

Which offers better value, Claude Sonnet 4.6 or GPT-5.1?+

There is no universal value winner. Compare the input and output prices above with the matched benchmark result for your workload: cheaper tokens can be offset by different quality, token usage, latency, or provider availability.

Send Feedback