A retrieval-augmented generation (RAG) app has several cost lines that are usually priced in different places: embedding your documents, storing the vectors, embedding each question, and generating each answer with the retrieved text pasted in. This calculator adds them into one monthly figure and a cost per question, so you can see which line actually matters.
Figures checked by Muhammad Ahmad against Claude Platform Docs: Pricing, OpenAI API pricing and Gemini API pricing. Prices and rules change, so confirm with the official source before relying on them. How we check the math
How it works
Your documents are split into chunks and each chunk is embedded once: chunks = document tokens ÷ chunk size, and the one-off embedding cost = document tokens ÷ 1,000,000 × the embedding price. Re-embedding new and changed documents each month costs that share of the one-off figure.
Each vector takes dimensions × 4 bytes as 32-bit floats. The calculator adds 20% for index structures and metadata, then multiplies the gigabytes by your vector database's monthly price.
Each question sends the retrieved chunks (chunks per question × chunk size) plus instructions and the question to the answer model, which writes the answer. That generation cost uses the verified model prices from the AI API Cost Calculator. Embedding the question itself is counted at 50 tokens.
Embedding and storage prices vary by provider and change, so they are inputs: enter the figures from your own providers' pricing pages. The defaults are examples.
A worked example
50 million tokens of documents in 500-token chunks is 100,000 vectors. At 1,536 dimensions that is about 0.74 GB with overhead, so storage at $0.33 per GB-month is about $0.24. Embedding everything once costs $1.00 at $0.02 per million tokens, and re-embedding 10% a month costs $0.10. The big line is answering: 100,000 questions, each sending 3,100 input tokens and getting 300 back, cost $460 a month on Claude Haiku 4.5. The total is about $460.44 a month, under half a cent per question.
Questions people ask
What is the most expensive part of a RAG system?
Usually answer generation, because every question sends the retrieved chunks to the model. Embedding documents is typically a small one-off cost, and vector storage is modest until you have many millions of vectors.
How can I reduce RAG costs?
Retrieve fewer or smaller chunks, re-rank to keep only the best ones, cache frequent questions, and use a smaller answer model where quality allows. Cutting retrieved context from 5 chunks to 3 reduces input on every question.
How much storage do embeddings use?
Dimensions × 4 bytes per vector as 32-bit floats: 1,536 dimensions is about 6 KB. Indexes and metadata add more. Quantized or lower-dimension embeddings reduce it.
What chunk size should I use?
Many systems use a few hundred tokens per chunk. Smaller chunks give more precise retrieval but more vectors; larger chunks carry more context into each answer, which raises generation cost.
Does this include re-ranking or hosting?
No. Add re-ranking API calls, application hosting and monitoring separately. The Vector Database Storage Calculator gives a more detailed storage estimate.