Tools

RAG index cost calculator

Embedding a corpus is a one-time bill. Answering questions against it is a monthly one, and the retrieved chunks are paid for again on every single query. Put your corpus size and traffic in and see both, plus what chunk overlap adds to the embedding bill.

Embedding price defaults to OpenAI text-embedding-3-small at $0.02 per million tokens. Set the price and the answer model rates to whatever you actually pay: they move, and the totals are only as good as these numbers.

Chunks in the index 0

Tokens embedded, with overlap 0

Overlap you pay for, against no overlap 0

One-time index build 0

Per query: embedding plus generation 0

Monthly at your traffic 0

Monthly re-embedding at your churn 0

First month, all in 0

What each query actually pays for

ComponentTokensCost

Overlap is the invisible multiplier

Overlap is there so a sentence split across a chunk boundary is still retrievable, and it works. What it also does is make each chunk advance by less than its own length, so a 500-token chunk with 50 tokens of overlap only moves the window forward 450 tokens. The corpus needs more chunks than a naive divide suggests, and every one of them is embedded in full. On a 50-million-token corpus that turns 100,000 chunks into about 111,000, so the embedding bill comes out 11 percent above what the corpus size suggests.

The rule is simple enough to hold in your head: chunks are ceil((corpus - overlap) / (chunk - overlap)), and embedding tokens are chunks times chunk size. Raise the overlap for better recall and the index build gets more expensive; there is no setting where overlap is free.

Retrieved chunks are paid for on every query

Most teams budget for the build and forget the query. Every question re-sends the retrieved chunks to the answer model, so at five chunks of 500 tokens your input is about 2,500 tokens before the question is even considered, multiplied by every query. At 20,000 queries a month that is 50 million input tokens, which is the same order as embedding the whole corpus once. Whether the monthly bill or the build bill dominates comes down to traffic, and the two totals above show you which one you are in.

If the retrieval half is what you are tuning, the prompt cache calculator covers the case where a fixed preamble or a stable document set is sent on every call, and the cost guide walks through the rest of an API bill.

Three things this page does not model

Vector storage is billed by the host, not by tokens, so it is not here. Reranking is an extra model call per query and can cost more than the retrieval it improves, so if you run one, add its rate as another input line by hand. Caching hits are not modelled either: if your queries repeat, the real bill is lower than this, possibly a lot lower.