Engineering

Count the tokens, then price the call

A short Python script prices an API call from a real trace: 8,000 in and 1,500 out is $0.031 on Sol and $0.155 on Astra.

The bill is not the list price. The bill is the list price times a token count you have not measured yet. Here is the whole calculation, and then a script that does it on a trace you already have.

Eight thousand input tokens and 1,500 output tokens cost $0.031 on GPT-6.1 Sol and $0.155 on GPT-6 Astra, at the standard short-context rates on OpenAI's pricing page as of 4 October 2026. One turn is nothing. Two hundred turns a day is the monthly gap in the cost guide.

Price a single call

Rates are USD per million tokens. Cached input is a separate rate, not a discount you apply yourself: on Sol it is $0.10 against $2.00, which is 5 percent of the input price, not 50.

RATES = {
    # USD per million tokens: input, cached input, output. Read 4 Oct 2026.
    "gpt-6-luna":   (0.10, 0.01, 0.50),
    "gpt-6.1-sol":  (2.00, 0.10, 10.00),
    "gpt-6-astra":  (10.00, 1.00, 50.00),
    "sonnet-5.5":   (2.00, 0.20, 10.00),
    "opus-5.5":     (4.00, 0.20, 20.00),
}

def call_cost(model, input_tokens, output_tokens, cached_tokens=0):
    price_in, price_cached, price_out = RATES[model]
    fresh = max(input_tokens - cached_tokens, 0)
    return (
        fresh / 1_000_000 * price_in
        + cached_tokens / 1_000_000 * price_cached
        + output_tokens / 1_000_000 * price_out
    )

print(round(call_cost("gpt-6.1-sol", 8_000, 1_500), 4))
# 0.031, of which 0.015 is the answer and 0.016 is the prompt

The printed result is $0.031, not $0.019, because the snippet prices the prompt as uncached. Pass cached_tokens=6_000 and the same call falls to $0.0194. That is the entire caching saving on this request: 1.2 cents. Worth having. Not worth designing the architecture around, unless the prompt is enormous and repeated.

Count tokens instead of guessing

The character rule (four per token, for English prose) is a ceiling on your confidence, not an estimate you should ship. tiktoken is the counter OpenAI publishes, and it is the right one for GPT-family models. Claude's newer tokenizer runs about 30 percent more tokens on the same text, according to Anthropic's pricing notes, so a number from tiktoken is a floor for a Claude bill.

import tiktoken

enc = tiktoken.get_encoding("o200k_base")

def tokens(text):
    return len(enc.encode(text))

prompt = open("trace-prompt.txt").read()
reply = open("trace-reply.txt").read()
print(tokens(prompt), tokens(reply))
print(round(call_cost("gpt-6.1-sol", tokens(prompt), tokens(reply)), 4))

Run it on a real trace, not on a paragraph you wrote to test the script. A system prompt plus three tool definitions plus a retrieved document is a different order of magnitude from the user message you remember typing.

Price a log, not a turn

One turn lies. A day of turns is the number finance will ask you for, and you already have it if you log usage. The OpenAI and Anthropic responses both return a usage object. Write those two integers down and the script gets simpler, because the vendor has counted for you.

import json

day = 0.0
with open("usage.jsonl") as fh:
    for line in fh:
        row = json.loads(line)
        day += call_cost(
            row["model"],
            row["input_tokens"],
            row["output_tokens"],
            row.get("cached_tokens", 0),
        )
print(f"${day:,.2f} today, ${day * 22:,.0f} across 22 weekdays")

A JSONL row looks like this. One model field, three counts:

{"model": "gpt-6.1-sol", "input_tokens": 8200, "output_tokens": 1460, "cached_tokens": 6100}

If the log has no cached_tokens, assume zero and treat the result as a ceiling. Vendors round, and a ceiling you can defend beats an estimate you cannot.

Two limits, so this does not get used for something it cannot do. The table has no long-context surcharge: an OpenAI request over 272,000 input tokens is repriced in full, and this function does not know that. It also has no batch rate. Halve the output of call_cost for a batch job, or read the batch comparison before you do, because the discount is only worth it when the result can wait.

Rates dated 4 October 2026. When a vendor changes a number, change the tuple, not the function.

Sources