Long-context surcharge calculator
OpenAI and Google both reprice the entire request once the input crosses a token threshold, rather than only the tokens above it. That turns the cost curve into a cliff, and it makes the obvious calculation wrong. This shows the cliff and the size of the error.
Cost per request at your token count $0.00
Cost just below the threshold $0.00
Cost one token above it $0.00
The cliff, for one token $0.00
The whole request is repriced
Every other rate in an API bill applies to the tokens you used. This one applies to the whole input. Cross the threshold and the multiplier hits your total, including the first 272,000 tokens that were being charged cheaply a moment ago. The cost jumps by a fixed amount at the boundary instead of growing smoothly.
That is why the intuitive calculation, charging the higher rate only on the overflow, is quietly wrong. The table below shows the difference for your numbers.
| Input tokens | Repriced, whole request | Naive, overflow only | Difference |
|---|
What to do about it
Because the jump is a step rather than a slope, shaving a few tokens does not help. Stay under the threshold, or stop pretending the request is small. Above the line, the sensible moves are the ones that cut input outright: retrieval that returns fewer chunks, a cached prefix for the stable part, or a cheaper tier that has no threshold this low.
Where each vendor puts its threshold, and the multipliers that apply above it, are on the OpenAI and Google pricing pages. The full cost model is in how to estimate an AI API bill. Caching is the lever that changes this number most, and it has its own break-even calculator.