The bill is not the list price. The bill is the list price times a token count you have not measured yet. Here is the whole calculation, and then a script that does it on a trace you already have.
Eight thousand input tokens and 1,500 output tokens cost $0.031 on GPT-6.1 Sol and $0.155 on GPT-6 Astra, at the standard short-context rates on OpenAI's pricing page as of 4 October 2026. One turn is nothing. Two hundred turns a day is the monthly gap in the cost guide.
Price a single call
Rates are USD per million tokens. Cached input is a separate rate, not a discount you apply yourself: on Sol it is $0.10 against $2.00, which is 5 percent of the input price, not 50.
RATES = {
# USD per million tokens: input, cached input, output. Read 4 Oct 2026.
"gpt-6-luna": (0.10, 0.01, 0.50),
"gpt-6.1-sol": (2.00, 0.10, 10.00),
"gpt-6-astra": (10.00, 1.00, 50.00),
"sonnet-5.5": (2.00, 0.20, 10.00),
"opus-5.5": (4.00, 0.20, 20.00),
}
def call_cost(model, input_tokens, output_tokens, cached_tokens=0):
price_in, price_cached, price_out = RATES[model]
fresh = max(input_tokens - cached_tokens, 0)
return (
fresh / 1_000_000 * price_in
+ cached_tokens / 1_000_000 * price_cached
+ output_tokens / 1_000_000 * price_out
)
print(round(call_cost("gpt-6.1-sol", 8_000, 1_500), 4))
# 0.031, of which 0.015 is the answer and 0.016 is the prompt
The printed result is $0.031, not $0.019, because the snippet prices the prompt as uncached. Pass cached_tokens=6_000 and the same call falls to $0.0194. That is the entire caching saving on this request: 1.2 cents. Worth having. Not worth designing the architecture around, unless the prompt is enormous and repeated.
Count tokens instead of guessing
The character rule (four per token, for English prose) is a ceiling on your confidence, not an estimate you should ship. tiktoken is the counter OpenAI publishes, and it is the right one for GPT-family models. Claude's newer tokenizer runs about 30 percent more tokens on the same text, according to Anthropic's pricing notes, so a number from tiktoken is a floor for a Claude bill.
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
def tokens(text):
return len(enc.encode(text))
prompt = open("trace-prompt.txt").read()
reply = open("trace-reply.txt").read()
print(tokens(prompt), tokens(reply))
print(round(call_cost("gpt-6.1-sol", tokens(prompt), tokens(reply)), 4))
Run it on a real trace, not on a paragraph you wrote to test the script. A system prompt plus three tool definitions plus a retrieved document is a different order of magnitude from the user message you remember typing.
Price a log, not a turn
One turn lies. A day of turns is the number finance will ask you for, and you already have it if you log usage. The OpenAI and Anthropic responses both return a usage object. Write those two integers down and the script gets simpler, because the vendor has counted for you.
import json
day = 0.0
with open("usage.jsonl") as fh:
for line in fh:
row = json.loads(line)
day += call_cost(
row["model"],
row["input_tokens"],
row["output_tokens"],
row.get("cached_tokens", 0),
)
print(f"${day:,.2f} today, ${day * 22:,.0f} across 22 weekdays")
A JSONL row looks like this. One model field, three counts:
{"model": "gpt-6.1-sol", "input_tokens": 8200, "output_tokens": 1460, "cached_tokens": 6100}
If the log has no cached_tokens, assume zero and treat the result as a ceiling. Vendors round, and a ceiling you can defend beats an estimate you cannot.
Two limits, so this does not get used for something it cannot do. The table has no long-context surcharge: an OpenAI request over 272,000 input tokens is repriced in full, and this function does not know that. It also has no batch rate. Halve the output of call_cost for a batch job, or read the batch comparison before you do, because the discount is only worth it when the result can wait.
Rates dated 4 October 2026. When a vendor changes a number, change the tuple, not the function.