Engineering

Batch is half price. The wait is the cost

Batch cuts GPT-6.1 Sol from $10 to $5 per million output tokens. On a 2 million token nightly run that is $220 a month, if the job can wait.

Both OpenAI and Anthropic sell a batch tier at half the standard token price. On GPT-6.1 Sol that takes output from $10 to $5 per million tokens. On a nightly job of 2 million output tokens, batch saves $10 a day, $220 across 22 weekdays. The discount is identical at every tier. The question is whether your job can sit in a queue.

The rates were read on 4 October 2026 from OpenAI's pricing page and Anthropic's. OpenAI also lists a Flex tier at the same half-price figures as batch. Anthropic's batch discount is stated as 50 percent on input and output.

Same discount, every tier

Half of a large number is a large number. Half of a small one is a reason to stop optimising and go and build the feature.

Model Standard output Batch output Saved per million
GPT-6 Luna 0.50 0.25 0.25
Claude Haiku 4.5 5.00 2.50 2.50
GPT-6.1 Sol 10.00 5.00 5.00
Claude Sonnet 5.5 10.00 5.00 5.00
Claude Opus 5.5 20.00 10.00 10.00
GPT-6 Astra 50.00 25.00 25.00
Claude Fable 5.1 50.00 25.00 25.00

Input is halved too. For an agent loop it barely matters, because output dominates. For a classification job that sends long documents and gets a label back, the input half is the whole saving, and Luna at $0.05 batch input is already so cheap that batch is not why you would choose it.

When the queue costs more than it saves

Batch is for work with no user waiting on the response. Evaluation runs, nightly re-indexing, backfills, the weekly report nobody reads before Monday. It is the wrong tool for a chat reply, an autocomplete, or an agent turn that blocks the next tool call. OpenAI documents batch as asynchronous, with a completion window measured in hours. If your product promise is seconds, the discount is not available to you, whatever the table says.

Three cases where the saving loses to the wait:

  • A retry you could have avoided. A batch job that fails overnight and gets rerun interactively in the morning was billed twice, once at each rate. The second run eats the first run's discount.
  • A stale result. A support-ticket classifier that batches overnight is fine. The same classifier holding a customer on a web form for an hour is a product bug that happens to be cheap.
  • A small job. Saving $0.25 per million tokens on Luna, at 100,000 tokens a day, is under a cent. The engineering time to split the pipeline into a batch path costs more than a year of the discount.

The useful split is the one in the cost guide: anything a person is watching stays on standard, anything a cron is watching moves to batch. Price both paths with the token script before you build the second path. If the gap is under $50 a month, do not build it.

The other half-price traps

OpenAI's long-context surcharge applies to batch as well. A request over 272,000 input tokens is repriced on the whole request, and then the batch halving applies to that higher number. Batch does not get you under the cliff. It gets you half of whatever the cliff charged.

Anthropic's fast mode is not available with batch, and it prices the other way: Opus 5.5 fast mode is $8 and $40 per million, double the standard rate. Speed and batch are opposite purchases. Buying both, on different parts of one pipeline, is reasonable. Buying both on the same call is not a configuration the API will allow, and averaging them in a spreadsheet produces a price nobody is charging you.

One more, because it shows up in procurement conversations. A 50 percent discount on tokens is not a 50 percent discount on the invoice once you add tool calls, storage, and the engineer who babysits the batch window. Token price is the part of the bill these pages document. It is rarely all of the bill.

Rates dated 4 October 2026.

Sources