Methodology
Provisioned cost uses all fleet hours; useful tokens use only utilized seconds at the measured throughput. Their ratio is the effective direct GPU cost per million tokens.
Infrastructure and Migration
Convert GPU hourly price, throughput, utilization, and fleet size into token capacity and unit economics.
Interactive Tool
| Result | Value | How to read it |
|---|---|---|
| Hourly fleet cost | $24.00 | Accelerators only. |
| Monthly fleet cost | $17,280.00 | Provisioned hours, including idle time. |
| Useful tokens / month | 1.24B | Sustained generation at useful utilization. |
| Direct cost / 1M tokens | $13.89 | GPU cost divided by useful generated tokens. |
| Idle-cost share | $6,912.00 | Provisioned cost outside useful inference time. |
Use sustained measured throughput for the exact model, quantization, batch policy, sequence length, and latency target.
Catalog model: Anthropic Claude Fable 5 →
Inputs
| Input | How it is used |
|---|---|
| GPU count and price | Provisioned accelerators and hourly cost per accelerator. |
| Throughput | Measured generated tokens per second per GPU for the target model and batch. |
| Utilization | Useful inference time as a share of provisioned time. |
| Provisioned hours | Monthly time the fleet remains available. |
Provisioned cost uses all fleet hours; useful tokens use only utilized seconds at the measured throughput. Their ratio is the effective direct GPU cost per million tokens.
Benchmark the exact model, quantization, sequence length, batch policy, and latency SLO. A peak-throughput benchmark is not a sustainable production rate.
Boundaries
Catalog values retain their source and freshness on the linked model, provider, benchmark, or comparison page. Editable scenario assumptions are not Model Markets measurements.
Continue the Analysis
| Tool | Next question |
|---|---|
| Self-Host vs API | Find the utilization and volume conditions under which a self-hosted inference cluster can beat an API on direct compute cost. |
| Price vs Performance | Screen for models that combine useful published performance with acceptable token economics. |
| Latency Comparison | Normalize latency measurements into user-visible wait time for the same output workload. |
Questions
Normalize a GPU deployment into cost per million generated tokens. It returns hourly and monthly spend, useful token capacity, effective token cost, and idle-cost share.
Where the calculation needs model facts, it uses the current Model Markets catalog snapshot updated 2026-09-02. User-entered assumptions remain clearly editable, and unsupported values stay unknown rather than being inferred.
CPU, memory, storage, networking, and platform margins are excluded. Redundancy and autoscaling headroom reduce utilization. Prompt ingestion may require a separate prefill measurement. Open the linked canonical records and primary sources before making a production or purchasing decision.