ARC-AGI-2
ARC Prize Foundation · verified-v2-326661568d5f · Verified ARC Prize results on the harder ARC-AGI-2 abstract reasoning benchmark.
Model Ranking
Source | |||||||
|---|---|---|---|---|---|---|---|
| #1 | GPT-5.6 SolOpenAI | 92.50 | — | — | GPT-5.6 Sol (Max)unknown | Original sourcewinner eligible | Third-party benchmark |
| #2 | Claude Opus 5Anthropic | 90.42 | — | — | Claude Opus 5 (Max)unknown | Original sourcewinner eligible | Third-party benchmark |
| #3 | Claude Fable 5.1Anthropic | 90.00 | — | — | Claude Fable 5.1 (XHigh)unknown | Original sourcewinner eligible | Third-party benchmark |
| #4 | Claude Fable 5Anthropic | 89.17 | — | — | Claude Fable 5 (Max)unknown | Original sourcewinner eligible | Third-party benchmark |
| #5 | GPT-5.5OpenAISelected model | 85.00 | — | — | GPT-5.5 (XHigh)unknown | Original sourcewinner eligible | Third-party benchmark |
| #6 | Gemini 3.7 FlashGoogle DeepMind | 84.58 | — | — | Gemini 3.7 Flash (High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #7 | GPT-5.5 ProOpenAI | 84.58 | — | — | GPT-5.5 Pro (High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #8 | GPT-5.6 TerraOpenAI | 83.90 | — | — | GPT-5.6 Terra (Max)unknown | Original sourcewinner eligible | Third-party benchmark |
| #9 | GPT-5.4 ProOpenAI | 83.33 | — | — | GPT-5.4 Pro (XHigh)unknown | Original sourcewinner eligible | Third-party benchmark |
| #10 | Claude Opus 4.7Anthropic | 75.83 | — | — | Claude 4.7 (Max)unknown | Original sourcewinner eligible | Third-party benchmark |
| #11 | GPT-5.4OpenAI | 73.95 | — | — | GPT-5.4 (XHigh)unknown | Original sourcewinner eligible | Third-party benchmark |
| #12 | Claude Opus 4.8Anthropic | 72.08 | — | — | Claude Opus 4.8 (High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #12 | Gemini 3.5 FlashGoogle DeepMind | 72.08 | — | — | Gemini 3.5 Flash (High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #14 | Claude Opus 4.6Anthropic | 69.17 | — | — | Claude Opus 4.6 (120K, High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #15 | Grok 4.6xAI | 67.08 | — | — | Grok 4.6 (XHigh)unknown | Original sourcewinner eligible | Third-party benchmark |
| #16 | Grok 4.20 Multi-AgentxAI | 65.14 | — | — | Grok 4.20 (Reasoning)unknown | Original sourcewinner eligible | Third-party benchmark |
| #17 | DeepSeek-V4-FlashDeepSeek | 61.39 | — | — | DeepSeek V4 Flash 0731 (Max)unknown | Original sourcewinner eligible | Third-party benchmark |
| #18 | DeepSeek-V4-Pro-0813DeepSeek | 61.25 | — | — | DeepSeek V4 Pro 0813 (Max)unknown | Original sourcewinner eligible | Third-party benchmark |
| #19 | Claude Sonnet 4.6Anthropic | 60.42 | — | — | Claude Sonnet 4.6 (High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #20 | Gemini 3.6 FlashGoogle DeepMind | 60.42 | — | — | Gemini 3.6 Flash (High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #20 | Kimi-K3Moonshot AI | 60.42 | — | — | Kimi K3 (Max)unknown | Original sourcewinner eligible | Third-party benchmark |
| #22 | GPT-5.6 LunaOpenAI | 59.58 | — | — | GPT-5.6 Luna 2026-07-30 (Max)unknown | Original sourcewinner eligible | Third-party benchmark |
| #23 | GPT-5.2 ProOpenAI | 54.16 | — | — | GPT-5.2 Pro (High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #24 | GPT-5.2OpenAI | 52.91 | — | — | GPT-5.2 (XHigh)unknown | Original sourcewinner eligible | Third-party benchmark |
| #25 | Grok 4.5xAI | 52.64 | — | — | Grok 4.5 (Medium)unknown | Original sourcewinner eligible | Third-party benchmark |
| #26 | InklingThinking Machines Lab | 36.53 | — | — | Inklingunknown | Original sourcewinner eligible | Third-party benchmark |
| #27 | GPT-5.4 MiniOpenAI | 18.90 | — | — | GPT-5.4 Mini (XHigh)unknown | Original sourcewinner eligible | Third-party benchmark |
| #28 | GPT-5 ProOpenAI | 18.33 | — | — | GPT-5 Prounknown | Original sourcewinner eligible | Third-party benchmark |
| #29 | GPT-5.1OpenAI | 17.64 | — | — | GPT-5.1 (Thinking, High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #30 | Claude Sonnet 4.5Anthropic | 13.61 | — | — | Claude Sonnet 4.5 (Thinking 32K)unknown | Original sourcewinner eligible | Third-party benchmark |
| #31 | Gemini 3.5 Flash-LiteGoogle DeepMind | 10.28 | — | — | Gemini 3.5 Flash-Lite (High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #32 | GPT-5OpenAI | 9.86 | — | — | GPT-5 (High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #33 | GPT-5.4 NanoOpenAI | 5.69 | — | — | GPT-5.4 Nano (XHigh)unknown | Original sourcewinner eligible | Third-party benchmark |
| #34 | GPT-5 MiniOpenAI | 4.44 | — | — | GPT-5 Mini (High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #35 | Claude Haiku 4.5Anthropic | 4.03 | — | — | Claude Haiku 4.5 (Thinking 32K)unknown | Original sourcewinner eligible | Third-party benchmark |
| #36 | GPT-5 NanoOpenAI | 2.61 | — | — | GPT-5 Nano (High)unknown | Original sourcewinner eligible | Third-party benchmark |
| #37 | GPT-4.1OpenAI | 0.42 | — | — | GPT-4.1unknown | Original sourcewinner eligible | Third-party benchmark |
37 models ranked by the best winner-eligible current-version score when available, otherwise the best publisher-artifact score; highest first. Observed Sep 3, 2026. The model you came from is highlighted.
Methodology and Coverage
Primary Evidence
Publisher Artifacts
Questions
ARC-AGI-2 FAQs
What does ARC-AGI-2 measure?+
Verified ARC Prize results on the harder ARC-AGI-2 abstract reasoning benchmark. Model Markets classifies it as an abstract reasoning benchmark and preserves the publisher's verified-v2-326661568d5f release as a distinct comparison cohort.
How are models ranked on ARC-AGI-2?+
Models are ordered by verified score in percent, with higher scores ranked first. Each model appears once; a verified winner-eligible current-version result takes precedence over a publisher-artifact-only result, and missing scores are not estimated.
Which model currently leads ARC-AGI-2?+
GPT-5.6 Sol leads the current verified table with 92.50 percent on verified-v2-326661568d5f. This is a benchmark-specific result, not a universal model-quality claim.
How many models have a published ARC-AGI-2 score?+
37 catalog models are published from 211 source rows. 43 source identities remain quarantined rather than guessed.
Can ARC-AGI-2 scores be compared with other benchmarks?+
Raw scores should be compared only within the same benchmark version, metric, and protocol. Model Markets normalizes eligible scores only for aggregate rankings, and groups models by an identical benchmark set before ranking them. View aggregate rankings →
Why might a model be missing from ARC-AGI-2?+
A model remains absent when the publisher has no current result, the source model identity is unresolved, the evaluation protocol is incompatible, or the evidence cannot be verified. Model Markets does not substitute a provider claim or infer a score from a related model.
How current is the ARC-AGI-2 leaderboard?+
The current Model Markets snapshot was observed Sep 3, 2026 from ARC Prize Foundation artifacts. The exact publisher source and each retained result artifact are linked on this page.