Methodology
Candidates need documented image input or vision capability. Ranking uses only sourced catalog fields and published benchmark coverage; it does not infer visual quality from a model name.
Model Selection
Rank models with documented image input using context, price, providers, and published benchmark evidence.
Interactive Tool
| Result | Value | How to read it |
|---|---|---|
| Eligible models | 70 | Models that pass the selected evidence and constraint gates. |
| Ranked shortlist | 20 | Top candidates shown below. |
| Priced candidates | 46 | Models with both input and output price minima. |
| Rank | Model | Fit score | Context | Input / 1M | Workload cost | Published benchmark rank | Providers |
|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 LunaOpenAI | 85.66 | 1.05M tokens | $0.20 / 1M | $44.00 | 4 | 2 |
| 2 | GPT-5.6 SolOpenAI | 84.37 | 1.05M tokens | $2.00 / 1M | $400.00 | 2 | 2 |
| 3 | GPT-5.6 TerraOpenAI | 81.61 | 1.05M tokens | $2.00 / 1M | $440.00 | 3 | 2 |
| 4 | Gemini 3.1 ProGoogle DeepMind | 76.95 | 1.05M tokens | $2.00 / 1M | $440.00 | 5 | 1 |
| 5 | Qwen3.8-MaxQwen | 75.82 | 1M tokens | $1.65 / 1M | $264.02 | 9 | 3 |
| 6 | Gemini 3.5 FlashGoogle DeepMind | 74.26 | 1.05M tokens | $1.50 / 1M | $330.00 | 8 | 1 |
| 7 | Gemini 3 FlashGoogle DeepMind | 72.94 | 1.05M tokens | $0.50 / 1M | $110.00 | 11 | 1 |
| 8 | Gemini 3.6 FlashGoogle DeepMind | 71.54 | 1.05M tokens | $0.75 / 1M | $150.00 | 12 | 1 |
| 9 | Gemini 2.5 ProGoogle DeepMind | 71.16 | 1.05M tokens | $1.25 / 1M | $325.00 | 10 | 1 |
| 10 | Claude Opus 5Anthropic | 70.68 | 1M tokens | $5.00 / 1M | $1,000.00 | 6 | 3 |
| 11 | Gemini 3.5 Flash-LiteGoogle DeepMind | 70.62 | 1.05M tokens | $0.30 / 1M | $80.00 | 14 | 1 |
| 12 | Claude Opus 4.8Anthropic | 68.28 | 1M tokens | $5.00 / 1M | $1,000.00 | 7 | 2 |
| 13 | Claude Sonnet 5Anthropic | 67.35 | 1M tokens | $2.00 / 1M | $400.00 | 15 | 3 |
| 14 | Kimi-K3Moonshot AI | 66.89 | 1.05M tokens | $2.85 / 1M | $570.00 | 13 | 2 |
| 15 | Claude Fable 5Anthropic | 66.77 | 1M tokens | $10.00 / 1M | $2,000.00 | 1 | 3 |
| 16 | Kimi-K2.6Moonshot AI | 66.34 | 262.14K tokens | $0.75 / 1M | $145.00 | 18 | 2 |
| 17 | MiniMax-M3MiniMax | 65.73 | 1.05M tokens | $0.28 / 1M | $50.00 | 20 | 2 |
| 18 | Grok 4.6xAI | 64.45 | 500K tokens | $2.00 / 1M | $320.00 | 16 | 1 |
| 19 | Gemini 3.7 FlashGoogle DeepMind | 63.88 | 1.05M tokens | $0.75 / 1M | $150.00 | 19 | 1 |
| 20 | Qwen3.7-PlusQwen | 63.05 | 1M tokens | $2.00 / 1M | $360.00 | 17 | 1 |
Fit scores are transparent screening aids, not universal quality claims. Unknown capability, context, or price fields are never filled by inference.
Inputs
| Input | How it is used |
|---|---|
| Minimum context | Text and image-token envelope required by the workload. |
| Monthly tokens | Text-token volume used for a comparable cost screen. |
| Priority | Weight performance, cost, or context capacity. |
Candidates need documented image input or vision capability. Ranking uses only sourced catalog fields and published benchmark coverage; it does not infer visual quality from a model name.
Check image size, count, file-type, OCR, and video limits on the exact endpoint. Text-token pricing may not capture image-unit charges.
Boundaries
Catalog values retain their source and freshness on the linked model, provider, benchmark, or comparison page. Editable scenario assumptions are not Model Markets measurements.
Continue the Analysis
| Tool | Next question |
|---|---|
| Model Selector | Turn workload constraints into a short, inspectable model shortlist instead of a universal best-model claim. |
| Benchmark Comparison | Inspect like-for-like published results without blending incompatible metrics or versions. |
| Image Generation Cost | Convert successful-image demand into actual generated attempts and monthly unit cost. |
Questions
Create a sourced shortlist for image understanding, document analysis, and multimodal prompts. It returns vision-capable candidates with context, price, provider coverage, and benchmark rank.
Where the calculation needs model facts, it uses the current Model Markets catalog snapshot updated 2026-09-02. User-entered assumptions remain clearly editable, and unsupported values stay unknown rather than being inferred.
Image billing units vary and may not map to text tokens. Vision support does not imply OCR, localization, or video support. Multimodal benchmark coverage remains uneven. Open the linked canonical records and primary sources before making a production or purchasing decision.