Methodology
The capability gate requires documented reasoning support. Ranking can use both maximum-token performance and token-efficiency results while keeping missing benchmark coverage explicit.
Model Selection
Compare models with documented reasoning support by benchmark evidence, token efficiency, context, and price.
Interactive Tool
| Result | Value | How to read it |
|---|---|---|
| Eligible models | 69 | Models that pass the selected evidence and constraint gates. |
| Ranked shortlist | 20 | Top candidates shown below. |
| Priced candidates | 54 | Models with both input and output price minima. |
| Rank | Model | Fit score | Context | Input / 1M | Workload cost | Published benchmark rank | Providers |
|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 LunaOpenAI | 85.66 | 1.05M tokens | $0.20 / 1M | $44.00 | 4 | 2 |
| 2 | GPT-5.6 SolOpenAI | 84.37 | 1.05M tokens | $2.00 / 1M | $400.00 | 2 | 2 |
| 3 | GPT-5.6 TerraOpenAI | 81.61 | 1.05M tokens | $2.00 / 1M | $440.00 | 3 | 2 |
| 4 | Gemini 3.1 ProGoogle DeepMind | 76.95 | 1.05M tokens | $2.00 / 1M | $440.00 | 5 | 1 |
| 5 | Qwen3.8-MaxQwen | 75.82 | 1M tokens | $1.65 / 1M | $264.02 | 9 | 3 |
| 6 | Gemini 3.5 FlashGoogle DeepMind | 74.26 | 1.05M tokens | $1.50 / 1M | $330.00 | 8 | 1 |
| 7 | Gemini 3 FlashGoogle DeepMind | 72.94 | 1.05M tokens | $0.50 / 1M | $110.00 | 11 | 1 |
| 8 | Gemini 3.6 FlashGoogle DeepMind | 71.54 | 1.05M tokens | $0.75 / 1M | $150.00 | 12 | 1 |
| 9 | Gemini 2.5 ProGoogle DeepMind | 71.16 | 1.05M tokens | $1.25 / 1M | $325.00 | 10 | 1 |
| 10 | Claude Opus 5Anthropic | 70.68 | 1M tokens | $5.00 / 1M | $1,000.00 | 6 | 3 |
| 11 | Gemini 3.5 Flash-LiteGoogle DeepMind | 70.62 | 1.05M tokens | $0.30 / 1M | $80.00 | 14 | 1 |
| 12 | Claude Opus 4.8Anthropic | 68.28 | 1M tokens | $5.00 / 1M | $1,000.00 | 7 | 2 |
| 13 | Claude Sonnet 5Anthropic | 67.35 | 1M tokens | $2.00 / 1M | $400.00 | 15 | 3 |
| 14 | Kimi-K3Moonshot AI | 66.89 | 1.05M tokens | $2.85 / 1M | $570.00 | 13 | 2 |
| 15 | Claude Fable 5Anthropic | 66.77 | 1M tokens | $10.00 / 1M | $2,000.00 | 1 | 3 |
| 16 | Kimi-K2.6Moonshot AI | 66.34 | 262.14K tokens | $0.75 / 1M | $145.00 | 18 | 2 |
| 17 | MiniMax-M3MiniMax | 65.73 | 1.05M tokens | $0.28 / 1M | $50.00 | 20 | 2 |
| 18 | DeepSeek-V4-ProDeepSeek | 64.84 | 1.05M tokens | $0.66 / 1M | $96.88 | 22 | 3 |
| 19 | Grok 4.6xAI | 64.45 | 500K tokens | $2.00 / 1M | $320.00 | 16 | 1 |
| 20 | DeepSeek-V4-FlashDeepSeek | 64.1 | 1.05M tokens | $0.09 / 1M | $12.15 | 24 | 3 |
Fit scores are transparent screening aids, not universal quality claims. Unknown capability, context, or price fields are never filled by inference.
Inputs
| Input | How it is used |
|---|---|
| Context requirement | Minimum request envelope for the reasoning task. |
| Output budget | Expected monthly generated or reasoning tokens. |
| Priority | Favor performance, token efficiency, or price. |
The capability gate requires documented reasoning support. Ranking can use both maximum-token performance and token-efficiency results while keeping missing benchmark coverage explicit.
Reasoning controls and token budgets differ by model and provider. Compare the exact configuration you plan to call, not only the family name.
Boundaries
Catalog values retain their source and freshness on the linked model, provider, benchmark, or comparison page. Editable scenario assumptions are not Model Markets measurements.
Continue the Analysis
| Tool | Next question |
|---|---|
| Benchmark Comparison | Inspect like-for-like published results without blending incompatible metrics or versions. |
| Price vs Performance | Screen for models that combine useful published performance with acceptable token economics. |
| Model Selector | Turn workload constraints into a short, inspectable model shortlist instead of a universal best-model claim. |
Questions
Shortlist reasoning models while keeping output-token economics and evidence coverage visible. It returns reasoning-capable candidates with performance rank, efficiency rank, context, and workload cost.
Where the calculation needs model facts, it uses the current Model Markets catalog snapshot updated 2026-09-02. User-entered assumptions remain clearly editable, and unsupported values stay unknown rather than being inferred.
Reasoning tokens may be billed or reported differently by provider. Higher output budgets can improve quality and increase latency or cost. A benchmark configuration may not match the API default. Open the linked canonical records and primary sources before making a production or purchasing decision.