RAG TCO Estimator: Infra + Inference
Published by Bles Software, a custom software and AI company based in Yehud-Monoson, Israel, building web apps, AI agents and API integrations for clients in Israel, the US, the UK and the EU.
The five lines a RAG bill is actually made of
A retrieval system's running cost is not one number, it is five, and they move independently. Embedding is paid once per document and again on every change. Storage is paid every month whether anyone searches or not. Retrieval is paid per query. Reranking is paid per query again, on a different meter. Generation is paid per token in and per token out. Put those five lines on one page with list prices against each and you have an estimate you can defend instead of a guess.
Everything priced below was read on 9 October 2026 from the vendor's own pricing page, with the unit the vendor bills in, because the unit is where estimates go wrong far more often than the rate.
Line one, embedding the corpus
- voyage-4-large. $0.12 per 1M tokens [1]
- voyage-4. $0.06 per 1M tokens [1]
- voyage-4-lite. $0.02 per 1M tokens [1]
Embedding is the cheapest line and the one people over-engineer. A million tokens is roughly 700,000 to 750,000 English words, so a corpus of five million words costs single dollars to embed at the rates above [1]. The cost that matters is not the first pass, it is re-embedding: every time you change chunking, change model, or the documents change, you pay it again. Budget for four or five full passes in the first year, not one.
Cohere publishes no per-token embedding rate on its current pricing page. What it does publish is dedicated-instance pricing: Embed 5 Fast and Pro at $3.00 to $5.00 per hour, or $2,000 to $3,250 per month [2]. That is a different cost shape, and it only pays back at constant high throughput.
Line two, keeping the vectors somewhere
- Pinecone Standard. $50 per month minimum usage, storage $0.33 per GB per month [3]
- Pinecone read and write. Read units $16 to $18 per million, write units $4 to $4.50 per million, varying by cloud and region [3]
- Weaviate Cloud Flex. From $45 per month minimum, vector dimensions $0.00465 per 1M, storage from $0.12 per GiB [4]
- Postgres with pgvector on Neon. Compute $0.106 per CU-hour on Launch, $0.222 on Scale, storage $0.35 per GB-month [5]
Read the units, not the headline. Pinecone charges you per read unit and per write unit [3], so your bill tracks query volume. Weaviate charges per million vector dimensions stored [4], so your bill tracks how wide your embedding model is multiplied by how many chunks you keep: a 3,072-dimension model costs three times a 1,024-dimension one for the same chunk count. Neon charges for compute by the hour and storage by the gigabyte [5], so a search index that is idle overnight costs almost nothing and one under constant load costs like a database, because it is one.
This is the line where "we will just use Postgres" is often right and often wrong for the same reason: it is priced like infrastructure rather than per query, which is cheap at low query volume and expensive at high volume.
Line three and four, retrieving and reranking
Retrieval is a read against line two, so it is already priced above. Reranking is a second model call on every query and it is the line most estimates forget.
- rerank-3. $0.05 per 1M tokens [1]
- rerank-3-lite. $0.02 per 1M tokens [1]
- Cohere Rerank 4 on dedicated instances. $5.00 to $10.00 per hour, or $3,250 to $6,500 per month [2]
Cohere's own definition is the number to hold on to: a single search unit is one query with up to 100 documents to be ranked [2]. Reranking 100 candidates per query is not a rounding error on a token meter, because each candidate's text goes through the model. If your retrieval returns 100 chunks of 500 tokens, every query sends 50,000 tokens to the reranker, which at $0.05 per million [1] is about a quarter of a cent per query before the generator has seen anything.
A managed vector store can also carry its own tool-call meter. OpenAI's file search prices storage at $0.10 per GB per day with 1 GB free, and tool calls at $2.50 per 1,000 [6]. Per day and per thousand calls are both easy to under-count by an order of magnitude.
Line five, generation, priced per answer rather than per million
The generator is the only line you can compute directly in cost per answer, so do it that way. One answer costs (context tokens + question tokens) x input rate + (answer tokens) x output rate. A retrieval system sends far more in than it gets out, typically by a factor of twenty or more, so the input rate dominates and the output rate barely matters.
Put real numbers through it. Eight retrieved chunks of 500 tokens, a 400-token instruction block and a 600-token answer is 4,400 tokens in and 600 out, which on Claude Sonnet 5.5 at $2 per million in and $10 per million out [7] costs $0.0088 plus $0.0060, so $0.0148. Swap the generator for Haiku 5.5 at $0.10 and $0.50 per million for prompts under 100,000 tokens [7] and the same answer costs $0.00044 plus $0.00030, so $0.00074, which is twenty times less. On gpt-6.1-sol at $2.00 in and $10.00 out the arithmetic matches the Sonnet line, and its cached input rate of $0.10 per million is where the saving sits [8].
That twenty-fold gap is the single biggest lever on this page, and it is a quality question rather than a cost question: in a retrieval system the model is summarising text you already found, which is the kind of work the cheap tier does well.
The two levers that cut the dominant line
The instruction block and the retrieved prefix repeat on every single query, which makes caching worth more in retrieval than in any other workload.
Anthropic's published multipliers are 1.25 times the base input rate for a five minute cache write, 2 times for a one hour write, and 0.1 times for a cache read, falling to 0.05 times on Opus 5.5 and Sonnet 5.5 [7]. Apply the Sonnet read multiplier to the 400-token instruction block in the example above and that portion of every request costs a twentieth of what it did. OpenAI publishes the same idea as a rate rather than a multiplier, $0.10 against $2.00 per million on gpt-6.1-sol [8].
The second lever is batching, and it applies to the half of a retrieval system nobody notices: index building, evaluation runs over a question set, and bulk summarisation. Anthropic's Batch API is 50 per cent off both input and output [7]. Evaluation is the big one here, because a serious eval suite re-answers hundreds of questions on every change and that traffic has no latency requirement at all.
One adjustment to any spreadsheet older than this year. Anthropic states that Claude 4.7 and later use a tokenizer producing approximately 30 per cent more tokens for the same text [7], so a chunk size chosen under an older tokenizer now sends about a third more than the spreadsheet thinks, on the line that already dominates the bill.
If you self-host instead
Self-hosting replaces four variable lines with one fixed one, and the published hourly rate is the honest comparison point. In AWS US East (N. Virginia), on-demand Linux, g4dn.xlarge is $0.526 per hour, g6.xlarge is $0.8048 per hour, g5.xlarge is $1.006 per hour and g6e.xlarge is $1.861 per hour [9]. A g5.xlarge carries one GPU with 24 GB of GPU memory, 4 vCPU and 16 GiB of system memory [10].
Run that continuously and $1.006 per hour is about $734 a month [9], before storage, networking and the engineer who keeps the server alive. Public wage data prices that last part: the US Bureau of Labor Statistics puts software developers at a mean $71.20 per hour in its May 2025 survey [11]. One day a month of maintenance is in the same order as the instance itself.
How to turn this into your own number
Five inputs, measured rather than assumed, and the model above becomes arithmetic: corpus size in tokens, embedding dimensions multiplied by chunk count, queries per month, candidates per query sent to the reranker, and average input and output tokens per answer. Measure the fourth and fifth on a week of real traffic. Those two are where estimates miss by a factor of ten, because people guess the retrieved context and the reranker's input and both are much larger than they feel.
We build the measurement first on every retrieval system, for the same reason: the cheapest version of this bill is almost always a caching change and a smaller reranker candidate set, not a different vendor.
Sources
- [1] Voyage AI, pricing documentation: voyage-4-large $0.12, voyage-4 $0.06 and voyage-4-lite $0.02 per 1M tokens, rerank-3 $0.05 and rerank-3-lite $0.02 per 1M tokens, read 9 October 2026, https://docs.voyageai.com/docs/pricing
- [2] Cohere, pricing page: Embed 5 Fast and Pro dedicated instances $3.00 to $5.00 per hour ($2,000 to $3,250 per month), Rerank 4 Fast and Pro $5.00 to $10.00 per hour ($3,250 to $6,500 per month), and the definition of a search unit, read 9 October 2026, https://cohere.com/pricing
- [3] Pinecone, pricing page: Standard $50 per month minimum usage, storage $0.33 per GB per month, write units $4 to $4.50 per million and read units $16 to $18 per million varying by cloud and region, read 9 October 2026, https://www.pinecone.io/pricing/
- [4] Weaviate, pricing page: Flex from $45 per month minimum, $0.00465 per 1M vector dimensions, storage from $0.12 per GiB, read 9 October 2026, https://weaviate.io/pricing
- [5] Neon, pricing page: Launch compute $0.106 per CU-hour, Scale compute $0.222 per CU-hour, storage $0.35 per GB-month, read 9 October 2026, https://neon.com/pricing
- [6] OpenAI, API pricing, file search and vector store section: storage $0.10 per GB per day with 1 GB free, tool calls $2.50 per 1,000, read 9 October 2026, https://developers.openai.com/api/docs/pricing
- [7] Anthropic, Claude platform pricing: per-million-token prices, cache write and read multipliers, the 50 per cent Batch API discount, and the tokenizer note for Claude 4.7 and later, read 9 October 2026, https://platform.claude.com/docs/en/about-claude/pricing
- [8] OpenAI, API pricing: standard and cached input prices and output prices per million tokens, read 9 October 2026, https://developers.openai.com/api/docs/pricing
- [9] Amazon Web Services, EC2 on-demand pricing data for Linux in US East (N. Virginia), read 9 October 2026, https://aws.amazon.com/ec2/pricing/on-demand/
- [10] Amazon Web Services, G5 instance specifications: g5.xlarge with 1 GPU, 24 GB GPU memory, 4 vCPU and 16 GiB memory, read 9 October 2026, https://aws.amazon.com/ec2/instance-types/g5/
- [11] US Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025, Software Developers (15-1252): mean hourly wage $71.20, read 9 October 2026 via the BLS public API, https://www.bls.gov/oes/current/oes151252.htm
More Costs and Timelines from Bles Software
- Reverse ETL Implementation Cost and Timeline (2025)
- Salesforce Data Cloud Implementation Cost and Timeline (2025)
- SLA Impact on Cost: Uptime → Spend
- Warehouse‑Native CDP Implementation Cost and Timeline (2025)
- Web App Development Cost in 2025: Budgets, Timelines, and the Architecture Behind Them
- Web App Development Cost: Scope Matrix
- AI Assistant Cost: Build vs Buy
- API Integration Cost in 2025: Pricing the Work Behind Reliable Connections
- Daily AI Roundup: AI agent, model and enterprise AI news