RAG TCO Estimator: Infra + Inference

Published by Bles Software, a custom software and AI company based in Yehud-Monoson, Israel, building web apps, AI agents and API integrations for clients in Israel, the US, the UK and the EU.

The five lines a RAG bill is actually made of

A retrieval system's running cost is not one number, it is five, and they move independently. Embedding is paid once per document and again on every change. Storage is paid every month whether anyone searches or not. Retrieval is paid per query. Reranking is paid per query again, on a different meter. Generation is paid per token in and per token out. Put those five lines on one page with list prices against each and you have an estimate you can defend instead of a guess.

Everything priced below was read on 9 October 2026 from the vendor's own pricing page, with the unit the vendor bills in, because the unit is where estimates go wrong far more often than the rate.

Line one, embedding the corpus

Embedding is the cheapest line and the one people over-engineer. A million tokens is roughly 700,000 to 750,000 English words, so a corpus of five million words costs single dollars to embed at the rates above [1]. The cost that matters is not the first pass, it is re-embedding: every time you change chunking, change model, or the documents change, you pay it again. Budget for four or five full passes in the first year, not one.

Cohere publishes no per-token embedding rate on its current pricing page. What it does publish is dedicated-instance pricing: Embed 5 Fast and Pro at $3.00 to $5.00 per hour, or $2,000 to $3,250 per month [2]. That is a different cost shape, and it only pays back at constant high throughput.

Line two, keeping the vectors somewhere

Read the units, not the headline. Pinecone charges you per read unit and per write unit [3], so your bill tracks query volume. Weaviate charges per million vector dimensions stored [4], so your bill tracks how wide your embedding model is multiplied by how many chunks you keep: a 3,072-dimension model costs three times a 1,024-dimension one for the same chunk count. Neon charges for compute by the hour and storage by the gigabyte [5], so a search index that is idle overnight costs almost nothing and one under constant load costs like a database, because it is one.

This is the line where "we will just use Postgres" is often right and often wrong for the same reason: it is priced like infrastructure rather than per query, which is cheap at low query volume and expensive at high volume.

Line three and four, retrieving and reranking

Retrieval is a read against line two, so it is already priced above. Reranking is a second model call on every query and it is the line most estimates forget.

Cohere's own definition is the number to hold on to: a single search unit is one query with up to 100 documents to be ranked [2]. Reranking 100 candidates per query is not a rounding error on a token meter, because each candidate's text goes through the model. If your retrieval returns 100 chunks of 500 tokens, every query sends 50,000 tokens to the reranker, which at $0.05 per million [1] is about a quarter of a cent per query before the generator has seen anything.

A managed vector store can also carry its own tool-call meter. OpenAI's file search prices storage at $0.10 per GB per day with 1 GB free, and tool calls at $2.50 per 1,000 [6]. Per day and per thousand calls are both easy to under-count by an order of magnitude.

Line five, generation, priced per answer rather than per million

The generator is the only line you can compute directly in cost per answer, so do it that way. One answer costs (context tokens + question tokens) x input rate + (answer tokens) x output rate. A retrieval system sends far more in than it gets out, typically by a factor of twenty or more, so the input rate dominates and the output rate barely matters.

Put real numbers through it. Eight retrieved chunks of 500 tokens, a 400-token instruction block and a 600-token answer is 4,400 tokens in and 600 out, which on Claude Sonnet 5.5 at $2 per million in and $10 per million out [7] costs $0.0088 plus $0.0060, so $0.0148. Swap the generator for Haiku 5.5 at $0.10 and $0.50 per million for prompts under 100,000 tokens [7] and the same answer costs $0.00044 plus $0.00030, so $0.00074, which is twenty times less. On gpt-6.1-sol at $2.00 in and $10.00 out the arithmetic matches the Sonnet line, and its cached input rate of $0.10 per million is where the saving sits [8].

That twenty-fold gap is the single biggest lever on this page, and it is a quality question rather than a cost question: in a retrieval system the model is summarising text you already found, which is the kind of work the cheap tier does well.

The two levers that cut the dominant line

The instruction block and the retrieved prefix repeat on every single query, which makes caching worth more in retrieval than in any other workload.

Anthropic's published multipliers are 1.25 times the base input rate for a five minute cache write, 2 times for a one hour write, and 0.1 times for a cache read, falling to 0.05 times on Opus 5.5 and Sonnet 5.5 [7]. Apply the Sonnet read multiplier to the 400-token instruction block in the example above and that portion of every request costs a twentieth of what it did. OpenAI publishes the same idea as a rate rather than a multiplier, $0.10 against $2.00 per million on gpt-6.1-sol [8].

The second lever is batching, and it applies to the half of a retrieval system nobody notices: index building, evaluation runs over a question set, and bulk summarisation. Anthropic's Batch API is 50 per cent off both input and output [7]. Evaluation is the big one here, because a serious eval suite re-answers hundreds of questions on every change and that traffic has no latency requirement at all.

One adjustment to any spreadsheet older than this year. Anthropic states that Claude 4.7 and later use a tokenizer producing approximately 30 per cent more tokens for the same text [7], so a chunk size chosen under an older tokenizer now sends about a third more than the spreadsheet thinks, on the line that already dominates the bill.

If you self-host instead

Self-hosting replaces four variable lines with one fixed one, and the published hourly rate is the honest comparison point. In AWS US East (N. Virginia), on-demand Linux, g4dn.xlarge is $0.526 per hour, g6.xlarge is $0.8048 per hour, g5.xlarge is $1.006 per hour and g6e.xlarge is $1.861 per hour [9]. A g5.xlarge carries one GPU with 24 GB of GPU memory, 4 vCPU and 16 GiB of system memory [10].

Run that continuously and $1.006 per hour is about $734 a month [9], before storage, networking and the engineer who keeps the server alive. Public wage data prices that last part: the US Bureau of Labor Statistics puts software developers at a mean $71.20 per hour in its May 2025 survey [11]. One day a month of maintenance is in the same order as the instance itself.

How to turn this into your own number

Five inputs, measured rather than assumed, and the model above becomes arithmetic: corpus size in tokens, embedding dimensions multiplied by chunk count, queries per month, candidates per query sent to the reranker, and average input and output tokens per answer. Measure the fourth and fifth on a week of real traffic. Those two are where estimates miss by a factor of ten, because people guess the retrieved context and the reranker's input and both are much larger than they feel.

We build the measurement first on every retrieval system, for the same reason: the cheapest version of this bill is almost always a caching change and a smaller reranker candidate set, not a different vendor.

Sources

More Costs and Timelines from Bles Software