Engineering

How to estimate an AI API bill before you ship

Two million output tokens a day costs $20 on GPT-6.1 Sol and $100 on Astra. Here is the arithmetic, from list price to a monthly bill.

A coding agent that writes 2 million output tokens a day costs $20 a day on GPT-6.1 Sol and $100 a day on GPT-6 Astra. Same loop, five times the bill, and that gap is the whole decision. Most teams find it on the invoice.

The list prices below were read off the vendor pricing pages on 4 October 2026. OpenAI's come from its API pricing page, Anthropic's from the Claude pricing page, Google's from the Gemini pricing page. They will move. The arithmetic will not.

The only formula that matters

Monthly bill, in dollars:

input_tokens  / 1_000_000 * input_price
+ output_tokens / 1_000_000 * output_price

A token is not a word. OpenAI's rule of thumb is about four characters of English per token, so 750 words is roughly 1,000 tokens. Code tokenises worse than prose, and Claude 4.7 and later uses a newer tokenizer that Anthropic says produces about 30 percent more tokens for the same text. If you skip that and budget from a word count, you will be low. The fix is to count tokens on a real trace, which is what the Python walk-through does.

Output is where the money goes. On every flagship pair below, output costs five times input. An agent that reasons in public, writing its plan into the response, is an output-heavy workload even when the prompt looks expensive.

What the three labs charge right now

Short context, standard API, USD per million tokens. OpenAI's long-context threshold is 272,000 input tokens; above it, input doubles and output is 1.5 times, on the whole request. Google's Gemini 3.8 Flash rates are promotional through 31 December 2026, then double.

Model Input Output Batch output
GPT-6 Luna 0.10 0.50 0.25
Gemini 3.8 Flash 0.75 3.75 about half
Claude Haiku 4.5 1.00 5.00 2.50
Claude Sonnet 5.5 2.00 10.00 5.00
GPT-6.1 Sol 2.00 10.00 5.00
Claude Opus 5.5 4.00 20.00 10.00
GPT-6 Astra 10.00 50.00 25.00
Claude Fable 5.1 10.00 50.00 25.00
Bar chart of output prices: Luna $0.50 to Astra $50 per million tokens.
Standard output price per million tokens, on a log scale, 4 October 2026. Sources: OpenAI, Anthropic and Google pricing pages.

Two things in that chart are easy to miss. Sol and Sonnet 5.5 land on the same $2 and $10. Astra and Fable 5.1 land on the same $10 and $50. At the list price, the mid tier and the flagship are no longer an OpenAI-versus-Anthropic argument. They are a tier argument, and you can swap vendors inside a tier without the bill moving.

Luna is the outlier, at one twentieth of Sol. It is also the model OpenAI pointed at the Decisions API, the one that classifies and routes. That is the right shape: a cheap model decides, an expensive model does the work that needs it. Routing that gets this wrong is the subject of the fixed-tier routing result.

Worked example: one agent, one day

Take a loop a team might actually run. A coding agent, 200 turns a day. Each turn sends 8,000 tokens of context and gets 1,500 tokens back. No caching yet.

Tokens a day Sol Astra
Input 1.6 million $3.20 $16.00
Output 0.3 million $3.00 $15.00
Day $6.20 $31.00
Month, 22 weekdays $136 $682

That is the polite version. It assumes the agent answers and stops. The version that hurts is the one that thinks out loud for 10,000 tokens a turn, 200 times: 2 million output tokens, $20 a day on Sol, $100 a day on Astra, and $440 versus $2,200 across a working month. The DevDay recap is where Sol's price came from. The invoice is where you will meet it.

Caching changes the input column, not the output column. OpenAI charges cached input on Sol at $0.10 per million, against $2.00 uncached. If 6,000 of those 8,000 input tokens are a stable system prompt and tool list, the input line falls from $3.20 to $0.52. You still owe the output. This is why "we cache everything" is not a cost strategy for an agent that writes long traces. It is a cost strategy for a chatbot that rereads the same manual.

Batch is the other lever, and it is a real one: half price at every vendor above, in exchange for waiting. The comparison, including when the wait costs more than the discount, is in batch against standard pricing.

Where the estimate goes wrong

Four mistakes account for most of the surprise bills we hear about.

  1. Counting words. A 30 percent tokenizer gap on Claude 4.7 and later is a 30 percent miss on the whole bill. Count tokens.
  2. Ignoring the long-context cliff. One request over 272,000 input tokens on OpenAI reprices the entire request, not the overflow. A retrieval pipeline that sometimes stuffs the whole corpus in is not "usually cheap."
  3. Forgetting tool output. Anthropic adds several hundred to several thousand input tokens per request once tools are declared, and a screenshot from computer use is billed as image input. The tool trace is part of the prompt.
  4. Budgeting the demo. A demo is 20 turns with a short answer. Production is the retry loop, the nightly batch, and the one user who pastes a 90-page PDF. Multiply the demo by ten before you quote anyone a number.

Prices dated 4 October 2026. If you are reading this later, re-read the three pricing pages before you reuse the table. The calculator runs the same formula against whatever rates you type in, so a stale default does not have to become a stale estimate.

If you would rather model several workloads at once, the rows above are downloadable as a CSV. The price cells are plain numbers and the monthly columns are live formulas, so overwriting one rate reprices every line that uses it.

Sources