Token and API Costs

Prompt Caching Savings Calculator

Estimate prompt-cache savings from reusable input share, cache-hit rate, calls, and editable token prices.

Interactive Tool

Build a Scenario

Runs in your browser
ResultValueHow to read it
Uncached baseline$2,000.00Every input token at the normal rate.
Cached scenario$920.00Hit-eligible prefix tokens at the cache-read rate.
Monthly savings$1,080.0054% below the baseline.
Cache-hit tokens480MReusable tokens expected to receive the cache rate.
Effective input rate / 1M$1.15Blended rate across all input tokens.

This scenario excludes cache writes, retention charges, minimum prefix lengths, and time-to-live rules.

Inputs

What the Calculation Needs

4 input groups
InputHow it is used
Input and reusable tokensTotal input per call and the stable prefix eligible for caching.
Cache-hit rateShare of calls expected to reuse the cached prefix.
CallsMonthly request count.
Input and cache ratesEditable per-million-token prices for the endpoint.

Methodology

The baseline prices every input token at the standard rate. The scenario prices hit-eligible reusable tokens at the cache rate and all remaining tokens normally.

How to Interpret the Result

Use a measured hit rate after accounting for provider cache keys, TTL, minimum prefix sizes, and prompt churn.

Boundaries

What the Result Does Not Prove

  1. Cache writes, retention charges, and minimum token rules vary by provider.
  2. A declared reusable prefix does not guarantee a cache hit.
  3. Output cost is unchanged unless the workload behavior also changes.

Catalog values retain their source and freshness on the linked model, provider, benchmark, or comparison page. Editable scenario assumptions are not Model Markets measurements.

Continue the Analysis

Related Tools

All tools →
ToolNext question
Prompt CostMake the cost of a system prompt, tool schema, or repeated instruction visible before scale.
Batch API SavingsQuantify the value of trading immediate responses for a batch discount.
LLM Cost CalculatorPrice one explicit token workload against an exact published provider endpoint.

Questions

Caching Savings FAQs

What does the Prompt Caching Savings Calculator calculate?+

Separate reusable prefix economics from total input volume before adopting prompt caching. It returns uncached baseline, cached scenario, monthly savings, and effective input rate.

Does the Prompt Caching Savings Calculator use current model data?+

Where the calculation needs model facts, it uses the current Model Markets catalog snapshot updated 2026-09-02. User-entered assumptions remain clearly editable, and unsupported values stay unknown rather than being inferred.

What should I verify before using the Prompt Caching Savings Calculator result?+

Cache writes, retention charges, and minimum token rules vary by provider. A declared reusable prefix does not guarantee a cache hit. Output cost is unchanged unless the workload behavior also changes. Open the linked canonical records and primary sources before making a production or purchasing decision.

Send Feedback