Tools

Long-context surcharge calculator

OpenAI and Google both reprice the entire request once the input crosses a token threshold, rather than only the tokens above it. That turns the cost curve into a cliff, and it makes the obvious calculation wrong. This shows the cliff and the size of the error.

Rates are editable and stay in the browser. Defaults follow the OpenAI and Google pricing pages as read on 4 October 2026. No prefix caching is applied, since a cached prefix is charged at the cache rate, not the surcharge rate.

Cost per request at your token count $0.00

Cost just below the threshold $0.00

Cost one token above it $0.00

The cliff, for one token $0.00

The whole request is repriced

Every other rate in an API bill applies to the tokens you used. This one applies to the whole input. Cross the threshold and the multiplier hits your total, including the first 272,000 tokens that were being charged cheaply a moment ago. The cost jumps by a fixed amount at the boundary instead of growing smoothly.

That is why the intuitive calculation, charging the higher rate only on the overflow, is quietly wrong. The table below shows the difference for your numbers.

Input tokensRepriced, whole requestNaive, overflow onlyDifference

What to do about it

Because the jump is a step rather than a slope, shaving a few tokens does not help. Stay under the threshold, or stop pretending the request is small. Above the line, the sensible moves are the ones that cut input outright: retrieval that returns fewer chunks, a cached prefix for the stable part, or a cheaper tier that has no threshold this low.

Where each vendor puts its threshold, and the multipliers that apply above it, are on the OpenAI and Google pricing pages. The full cost model is in how to estimate an AI API bill. Caching is the lever that changes this number most, and it has its own break-even calculator.