ARC-AGI-2

ARC Prize Foundation · verified-v2-b5cb5ec6e8c7 · Verified ARC Prize results on the harder ARC-AGI-2 abstract reasoning benchmark.

Model Ranking

Source
#1GPT-6 AstraOpenAI95.00percentGPT-6 Astra (Max)unknownOriginal sourcewinner eligibleThird-party benchmark
#2GPT-5.6 SolOpenAI92.50percentGPT-5.6 Sol (Max)unknownOriginal sourcewinner eligibleThird-party benchmark
#3Claude Opus 5Anthropic90.42percentClaude Opus 5 (Max)unknownOriginal sourcewinner eligibleThird-party benchmark
#4Claude Fable 5.1Anthropic90.00percentClaude Fable 5.1 (XHigh)unknownOriginal sourcewinner eligibleThird-party benchmark
#5Claude Fable 5Anthropic89.17percentClaude Fable 5 (Max)unknownOriginal sourcewinner eligibleThird-party benchmark
#6GPT-5.5OpenAI85.00percentGPT-5.5 (XHigh)unknownOriginal sourcewinner eligibleThird-party benchmark
#7Gemini 3.7 FlashGoogle DeepMind84.58percentGemini 3.7 Flash (High)unknownOriginal sourcewinner eligibleThird-party benchmark
#8GPT-5.5 ProOpenAI84.58percentGPT-5.5 Pro (High)unknownOriginal sourcewinner eligibleThird-party benchmark
#9GPT-5.6 TerraOpenAI83.90percentGPT-5.6 Terra (Max)unknownOriginal sourcewinner eligibleThird-party benchmark
#10GPT-5.4 ProOpenAISelected model83.33percentGPT-5.4 Pro (XHigh)unknownOriginal sourcewinner eligibleThird-party benchmark
#11Claude Opus 4.7Anthropic75.83percentClaude 4.7 (Max)unknownOriginal sourcewinner eligibleThird-party benchmark
#12GPT-5.4OpenAI73.95percentGPT-5.4 (XHigh)unknownOriginal sourcewinner eligibleThird-party benchmark
#13Claude Opus 4.8Anthropic72.08percentClaude Opus 4.8 (High)unknownOriginal sourcewinner eligibleThird-party benchmark
#13Gemini 3.5 FlashGoogle DeepMind72.08percentGemini 3.5 Flash (High)unknownOriginal sourcewinner eligibleThird-party benchmark
#15Claude Opus 4.6Anthropic69.17percentClaude Opus 4.6 (120K, High)unknownOriginal sourcewinner eligibleThird-party benchmark
#16Grok 4.6xAI67.08percentGrok 4.6 (XHigh)unknownOriginal sourcewinner eligibleThird-party benchmark
#17Grok 4.20 Multi AgentxAI65.14percentGrok 4.20 (Reasoning)unknownOriginal sourcewinner eligibleThird-party benchmark
#18DeepSeek V4 FlashDeepSeek61.39percentDeepSeek V4 Flash 0731 (Max)unknownOriginal sourcewinner eligibleThird-party benchmark
#19DeepSeek V4 Pro 0813DeepSeek61.25percentDeepSeek V4 Pro 0813 (Max)unknownOriginal sourcewinner eligibleThird-party benchmark
#20Claude Sonnet 4.6Anthropic60.42percentClaude Sonnet 4.6 (High)unknownOriginal sourcewinner eligibleThird-party benchmark
#21Gemini 3.6 FlashGoogle DeepMind60.42percentGemini 3.6 Flash (High)unknownOriginal sourcewinner eligibleThird-party benchmark
#21Kimi K3Moonshot AI60.42percentKimi K3 (Max)unknownOriginal sourcewinner eligibleThird-party benchmark
#23GPT-5.6 LunaOpenAI59.58percentGPT-5.6 Luna 2026-07-30 (Max)unknownOriginal sourcewinner eligibleThird-party benchmark
#24GPT-5.2 ProOpenAI54.16percentGPT-5.2 Pro (High)unknownOriginal sourcewinner eligibleThird-party benchmark
#25GPT-5.2OpenAI52.91percentGPT-5.2 (XHigh)unknownOriginal sourcewinner eligibleThird-party benchmark
#26Grok 4.5xAI52.64percentGrok 4.5 (Medium)unknownOriginal sourcewinner eligibleThird-party benchmark
#27InklingThinking Machines Lab36.53percentInklingunknownOriginal sourcewinner eligibleThird-party benchmark
#28GPT-5.4 MiniOpenAI18.90percentGPT-5.4 Mini (XHigh)unknownOriginal sourcewinner eligibleThird-party benchmark
#29GPT-5 ProOpenAI18.33percentGPT-5 ProunknownOriginal sourcewinner eligibleThird-party benchmark
#30GPT-5.1OpenAI17.64percentGPT-5.1 (Thinking, High)unknownOriginal sourcewinner eligibleThird-party benchmark
#31Claude Sonnet 4.5Anthropic13.61percentClaude Sonnet 4.5 (Thinking 32K)unknownOriginal sourcewinner eligibleThird-party benchmark
#32Gemini 3.5 Flash LiteGoogle DeepMind10.28percentGemini 3.5 Flash-Lite (High)unknownOriginal sourcewinner eligibleThird-party benchmark
#33GPT-5OpenAI9.86percentGPT-5 (High)unknownOriginal sourcewinner eligibleThird-party benchmark
#34GPT-5.4 NanoOpenAI5.69percentGPT-5.4 Nano (XHigh)unknownOriginal sourcewinner eligibleThird-party benchmark
#35GPT-5 MiniOpenAI4.44percentGPT-5 Mini (High)unknownOriginal sourcewinner eligibleThird-party benchmark
#36Claude Haiku 4.5Anthropic4.03percentClaude Haiku 4.5 (Thinking 32K)unknownOriginal sourcewinner eligibleThird-party benchmark
#37GPT-5 NanoOpenAI2.61percentGPT-5 Nano (High)unknownOriginal sourcewinner eligibleThird-party benchmark
#38GPT-4.1OpenAI0.42percentGPT-4.1unknownOriginal sourcewinner eligibleThird-party benchmark

38 models ranked by the best winner-eligible current-version score when available, otherwise the best publisher-artifact score; highest first. Observed Sep 4, 2026. The model you came from is highlighted.

Methodology and Coverage

Ranking protocolverified score · percent. Each model appears once; verified winner-eligible runs take precedence, and missing models are not imputed.
Acquisition coverage38 published models from 217 source rows; 43 source identities remain quarantined.

Primary Evidence

Publisher Artifacts

Questions

ARC-AGI-2 FAQs

What does ARC-AGI-2 measure?+

Verified ARC Prize results on the harder ARC-AGI-2 abstract reasoning benchmark. Model Markets classifies it as an abstract reasoning benchmark and preserves the publisher's verified-v2-b5cb5ec6e8c7 release as a distinct comparison cohort.

How are models ranked on ARC-AGI-2?+

Models are ordered by verified score in percent, with higher scores ranked first. Each model appears once; a verified winner-eligible current-version result takes precedence over a publisher-artifact-only result, and missing scores are not estimated.

Which model currently leads ARC-AGI-2?+

GPT-6 Astra leads the current verified table with 95.00 percent on verified-v2-b5cb5ec6e8c7. This is a benchmark-specific result, not a universal model-quality claim.

How many models have a published ARC-AGI-2 score?+

38 catalog models are published from 217 source rows. 43 source identities remain quarantined rather than guessed.

Can ARC-AGI-2 scores be compared with other benchmarks?+

Raw scores should be compared only within the same benchmark version, metric, and protocol. Model Markets normalizes eligible scores only for aggregate rankings, and groups models by an identical benchmark set before ranking them. View aggregate rankings

Why might a model be missing from ARC-AGI-2?+

A model remains absent when the publisher has no current result, the source model identity is unresolved, the evaluation protocol is incompatible, or the evidence cannot be verified. Model Markets does not substitute a provider claim or infer a score from a related model.

How current is the ARC-AGI-2 leaderboard?+

The current Model Markets snapshot was observed Sep 4, 2026 from ARC Prize Foundation artifacts. The exact publisher source and each retained result artifact are linked on this page.

Send Feedback