Model Selection

Reasoning Model Selector

Compare models with documented reasoning support by benchmark evidence, token efficiency, context, and price.

Interactive Tool

Build a Scenario

Runs in your browser
Ranking priority
Changes weights, not the eligibility gate.
ResultValueHow to read it
Eligible models69Models that pass the selected evidence and constraint gates.
Ranked shortlist20Top candidates shown below.
Priced candidates54Models with both input and output price minima.
RankModelFit scoreContextInput / 1MWorkload costPublished benchmark rankProviders
1GPT-5.6 LunaOpenAI85.661.05M tokens$0.20 / 1M$44.0042
2GPT-5.6 SolOpenAI84.371.05M tokens$2.00 / 1M$400.0022
3GPT-5.6 TerraOpenAI81.611.05M tokens$2.00 / 1M$440.0032
4Gemini 3.1 ProGoogle DeepMind76.951.05M tokens$2.00 / 1M$440.0051
5Qwen3.8-MaxQwen75.821M tokens$1.65 / 1M$264.0293
6Gemini 3.5 FlashGoogle DeepMind74.261.05M tokens$1.50 / 1M$330.0081
7Gemini 3 FlashGoogle DeepMind72.941.05M tokens$0.50 / 1M$110.00111
8Gemini 3.6 FlashGoogle DeepMind71.541.05M tokens$0.75 / 1M$150.00121
9Gemini 2.5 ProGoogle DeepMind71.161.05M tokens$1.25 / 1M$325.00101
10Claude Opus 5Anthropic70.681M tokens$5.00 / 1M$1,000.0063
11Gemini 3.5 Flash-LiteGoogle DeepMind70.621.05M tokens$0.30 / 1M$80.00141
12Claude Opus 4.8Anthropic68.281M tokens$5.00 / 1M$1,000.0072
13Claude Sonnet 5Anthropic67.351M tokens$2.00 / 1M$400.00153
14Kimi-K3Moonshot AI66.891.05M tokens$2.85 / 1M$570.00132
15Claude Fable 5Anthropic66.771M tokens$10.00 / 1M$2,000.0013
16Kimi-K2.6Moonshot AI66.34262.14K tokens$0.75 / 1M$145.00182
17MiniMax-M3MiniMax65.731.05M tokens$0.28 / 1M$50.00202
18DeepSeek-V4-ProDeepSeek64.841.05M tokens$0.66 / 1M$96.88223
19Grok 4.6xAI64.45500K tokens$2.00 / 1M$320.00161
20DeepSeek-V4-FlashDeepSeek64.11.05M tokens$0.09 / 1M$12.15243

Fit scores are transparent screening aids, not universal quality claims. Unknown capability, context, or price fields are never filled by inference.

Inputs

What the Calculation Needs

3 input groups
InputHow it is used
Context requirementMinimum request envelope for the reasoning task.
Output budgetExpected monthly generated or reasoning tokens.
PriorityFavor performance, token efficiency, or price.

Methodology

The capability gate requires documented reasoning support. Ranking can use both maximum-token performance and token-efficiency results while keeping missing benchmark coverage explicit.

How to Interpret the Result

Reasoning controls and token budgets differ by model and provider. Compare the exact configuration you plan to call, not only the family name.

Boundaries

What the Result Does Not Prove

  1. Reasoning tokens may be billed or reported differently by provider.
  2. Higher output budgets can improve quality and increase latency or cost.
  3. A benchmark configuration may not match the API default.

Catalog values retain their source and freshness on the linked model, provider, benchmark, or comparison page. Editable scenario assumptions are not Model Markets measurements.

Continue the Analysis

Related Tools

All tools →
ToolNext question
Benchmark ComparisonInspect like-for-like published results without blending incompatible metrics or versions.
Price vs PerformanceScreen for models that combine useful published performance with acceptable token economics.
Model SelectorTurn workload constraints into a short, inspectable model shortlist instead of a universal best-model claim.

Questions

Reasoning Models FAQs

What does the Reasoning Model Selector calculate?+

Shortlist reasoning models while keeping output-token economics and evidence coverage visible. It returns reasoning-capable candidates with performance rank, efficiency rank, context, and workload cost.

Does the Reasoning Model Selector use current model data?+

Where the calculation needs model facts, it uses the current Model Markets catalog snapshot updated 2026-09-02. User-entered assumptions remain clearly editable, and unsupported values stay unknown rather than being inferred.

What should I verify before using the Reasoning Model Selector result?+

Reasoning tokens may be billed or reported differently by provider. Higher output budgets can improve quality and increase latency or cost. A benchmark configuration may not match the API default. Open the linked canonical records and primary sources before making a production or purchasing decision.

Send Feedback