Methodology
The baseline prices every input token at the standard rate. The scenario prices hit-eligible reusable tokens at the cache rate and all remaining tokens normally.
Token and API Costs
Estimate prompt-cache savings from reusable input share, cache-hit rate, calls, and editable token prices.
Interactive Tool
| Result | Value | How to read it |
|---|---|---|
| Uncached baseline | $2,000.00 | Every input token at the normal rate. |
| Cached scenario | $920.00 | Hit-eligible prefix tokens at the cache-read rate. |
| Monthly savings | $1,080.00 | 54% below the baseline. |
| Cache-hit tokens | 480M | Reusable tokens expected to receive the cache rate. |
| Effective input rate / 1M | $1.15 | Blended rate across all input tokens. |
This scenario excludes cache writes, retention charges, minimum prefix lengths, and time-to-live rules.
Catalog model: Anthropic Claude Fable 5 →
Inputs
| Input | How it is used |
|---|---|
| Input and reusable tokens | Total input per call and the stable prefix eligible for caching. |
| Cache-hit rate | Share of calls expected to reuse the cached prefix. |
| Calls | Monthly request count. |
| Input and cache rates | Editable per-million-token prices for the endpoint. |
The baseline prices every input token at the standard rate. The scenario prices hit-eligible reusable tokens at the cache rate and all remaining tokens normally.
Use a measured hit rate after accounting for provider cache keys, TTL, minimum prefix sizes, and prompt churn.
Boundaries
Catalog values retain their source and freshness on the linked model, provider, benchmark, or comparison page. Editable scenario assumptions are not Model Markets measurements.
Continue the Analysis
| Tool | Next question |
|---|---|
| Prompt Cost | Make the cost of a system prompt, tool schema, or repeated instruction visible before scale. |
| Batch API Savings | Quantify the value of trading immediate responses for a batch discount. |
| LLM Cost Calculator | Price one explicit token workload against an exact published provider endpoint. |
Questions
Separate reusable prefix economics from total input volume before adopting prompt caching. It returns uncached baseline, cached scenario, monthly savings, and effective input rate.
Where the calculation needs model facts, it uses the current Model Markets catalog snapshot updated 2026-09-02. User-entered assumptions remain clearly editable, and unsupported values stay unknown rather than being inferred.
Cache writes, retention charges, and minimum token rules vary by provider. A declared reusable prefix does not guarantee a cache hit. Output cost is unchanged unless the workload behavior also changes. Open the linked canonical records and primary sources before making a production or purchasing decision.