Methodology
The table joins source-linked token prices to Model Markets' verified benchmark ranking. Missing prices or benchmark rows stay unknown and do not receive fabricated ratios.
Model Selection
Compare published benchmark performance with estimated workload cost and visible evidence coverage.
Interactive Tool
| Model | Performance rank | Performance score | Coverage | Workload cost | Cost / score point |
|---|---|---|---|---|---|
| DeepSeek-V4-FlashDeepSeek | 24 | 30.22 | 36.36% | $12.15 | $0.40 |
| GPT-4.1 NanoOpenAI | 36 | 20.23 | 36.36% | $9.00 | $0.44 |
| GLM-5.3-FlashZ.ai | 30 | 25.45 | 27.27% | $12.50 | $0.49 |
| GPT-5.6 LunaOpenAI | 4 | 64.84 | 72.73% | $44.00 | $0.68 |
| Llama-4-Scout-17B-16E-InstructMeta | 37 | 19.70 | 27.27% | $16.00 | $0.81 |
| GPT-4o MiniOpenAI | 35 | 21.13 | 36.36% | $27.00 | $1.28 |
| GPT-4.1 MiniOpenAI | 27 | 25.74 | 36.36% | $36.00 | $1.40 |
| MiniMax-M3MiniMax | 20 | 33.31 | 36.36% | $50.00 | $1.50 |
| Llama-4-Maverick-17B-128E-InstructMeta | 39 | 17.97 | 27.27% | $33.92 | $1.89 |
| Gemini 3.5 Flash-LiteGoogle DeepMind | 14 | 39.06 | 45.45% | $80.00 | $2.05 |
| Gemini 2.5 FlashGoogle DeepMind | 21 | 32.27 | 36.36% | $80.00 | $2.48 |
| Gemini 3 FlashGoogle DeepMind | 11 | 41.80 | 45.45% | $110.00 | $2.63 |
| DeepSeek-V4-ProDeepSeek | 22 | 31.16 | 36.36% | $96.88 | $3.11 |
| Mistral Large 3Mistral AI | 33 | 24.12 | 27.27% | $80.00 | $3.32 |
| Kimi-K2.5Moonshot AI | 26 | 26.03 | 27.27% | $90.00 | $3.46 |
| Qwen3.8-27BQwen | 29 | 25.63 | 27.27% | $91.00 | $3.55 |
| Gemini 3.6 FlashGoogle DeepMind | 12 | 41.76 | 45.45% | $150.00 | $3.59 |
| Kimi-K2.6Moonshot AI | 18 | 34.11 | 36.36% | $145.00 | $4.25 |
| Gemini 3.7 FlashGoogle DeepMind | 19 | 33.82 | 36.36% | $150.00 | $4.44 |
| GPT-5.6 SolOpenAI | 2 | 80.28 | 81.82% | $400.00 | $4.98 |
| GPT-4.1OpenAI | 23 | 30.76 | 36.36% | $180.00 | $5.85 |
| Qwen3.8-MaxQwen | 9 | 43.53 | 45.45% | $264.02 | $6.06 |
| GPT-5.6 TerraOpenAI | 3 | 69.25 | 72.73% | $440.00 | $6.35 |
| Gemini 3.5 FlashGoogle DeepMind | 8 | 50.40 | 54.55% | $330.00 | $6.55 |
| Grok 4.20 Multi-AgentxAI | 25 | 26.08 | 27.27% | $175.00 | $6.71 |
| Gemini 3.1 ProGoogle DeepMind | 5 | 59.80 | 63.64% | $440.00 | $7.36 |
| Gemini 2.5 ProGoogle DeepMind | 10 | 42.19 | 45.45% | $325.00 | $7.70 |
| GLM-5.3Z.ai | 28 | 25.68 | 27.27% | $200.00 | $7.79 |
| Grok 4.6xAI | 16 | 34.83 | 36.36% | $320.00 | $9.19 |
| Qwen3.7-PlusQwen | 17 | 34.17 | 36.36% | $360.00 | $10.54 |
| Claude Sonnet 5Anthropic | 15 | 35.09 | 36.36% | $400.00 | $11.40 |
| Kimi-K3Moonshot AI | 13 | 41.06 | 45.45% | $570.00 | $13.88 |
| GPT-4oOpenAI | 32 | 24.33 | 36.36% | $450.00 | $18.50 |
| Claude Opus 5Anthropic | 6 | 53.92 | 54.55% | $1,000.00 | $18.55 |
| Claude Opus 4.8Anthropic | 7 | 52.56 | 54.55% | $1,000.00 | $19.03 |
| Command ACohere | 34 | 22.64 | 36.36% | $450.00 | $19.88 |
| Claude Fable 5Anthropic | 1 | 81.15 | 81.82% | $2,000.00 | $24.65 |
| Mistral-Medium-3.5-128BMistral AI | 31 | 24.59 | 27.27% | Unknown | Unknown |
Compare rows with similar benchmark coverage. A low ratio is a screening signal, not a workload-specific quality guarantee.
Inputs
| Input | How it is used |
|---|---|
| Monthly workload | Input and output tokens used to estimate each model's cost. |
| Minimum coverage | Required share of the published ranking benchmark set. |
The table joins source-linked token prices to Model Markets' verified benchmark ranking. Missing prices or benchmark rows stay unknown and do not receive fabricated ratios.
Compare rows within similar evidence coverage. A cheap model with sparse benchmark participation can look efficient simply because the hard tests are missing.
Boundaries
Catalog values retain their source and freshness on the linked model, provider, benchmark, or comparison page. Editable scenario assumptions are not Model Markets measurements.
Continue the Analysis
| Tool | Next question |
|---|---|
| Benchmark Comparison | Inspect like-for-like published results without blending incompatible metrics or versions. |
| Model Selector | Turn workload constraints into a short, inspectable model shortlist instead of a universal best-model claim. |
| LLM Cost Calculator | Price one explicit token workload against an exact published provider endpoint. |
Questions
Screen for models that combine useful published performance with acceptable token economics. It returns a sortable frontier-style table with benchmark score, coverage, workload cost, and cost per performance point.
Where the calculation needs model facts, it uses the current Model Markets catalog snapshot updated 2026-09-02. User-entered assumptions remain clearly editable, and unsupported values stay unknown rather than being inferred.
Benchmark aggregation is not a workload-specific quality guarantee. Lowest token prices can come from different providers. Latency and reliability are not part of the ratio. Open the linked canonical records and primary sources before making a production or purchasing decision.