Prompt cache break-even calculator
Caching a prompt prefix is not automatically cheaper. Anthropic charges a premium to write the cache, so a prefix that is written often and read rarely costs more than never caching it. This finds the hit rate where caching starts to pay, and prices your own workload at the hit rate you actually see.
Cost per call, no caching $0.00
Cost per call, with caching $0.00
Saved per call $0.00
Saved per month, 30 days $0
Break-even hit rate 0%
Why a cache can cost more
Three costs decide it. Write the cache on a miss and you pay write multiplier x prefix price. Read it on a hit and you pay read multiplier x prefix price. Not caching at all costs 1 x prefix price on every call. Setting the expected cost of caching equal to the cost of not caching, the hit rate that breaks even is (write multiplier - 1) / (write multiplier - read multiplier).
At Anthropic's 1.25x write and 0.1x read that is 21.7 percent. Below that hit rate, caching the prefix is strictly worse than sending it plainly, and the more calls you make, the more you lose. OpenAI's write premium is effectively zero, so the formula gives a break-even of zero: caching there is never worse, only unhelpfully small when the prefix is short.
The arithmetic hides two caveats. Caches expire, so a hit rate measured over a burst can be far above the rate you get overnight, when the cache has gone cold and every call pays the write. And a cached prefix only counts as a hit when the bytes match exactly from the start, so a timestamp or a user id near the top of the prompt drops your real hit rate to zero without any error appearing.
The rest of the cost model, including batch and the long-context cliff, is in how to estimate an AI API bill. To price a real trace rather than a guess, use the Python counter.
Where the numbers come from
The multipliers are from Anthropic's prompt caching documentation and the OpenAI and Google pricing pages, read on 4 October 2026. Anthropic also states a minimum cacheable prefix of 1,024 tokens for most models, so a shorter prefix will not cache at all. Re-read those pages before you reuse a default here.