Home / Tools / RAG cost
RAG cost calculator Estimate retrieval-augmented generation cost: corpus embedding, per-query retrieval pack, and LLM generation. See which line dominates your monthly bill.
Cost calculator · Embedding cost · Context window · Cache savings · RAG budgeting
Generation model OpenAI. ChatGPT-4o Latest OpenAI. GPT-3.5 Turbo OpenAI. GPT-4 OpenAI. GPT-4.1 OpenAI. GPT-4.1 (2025-04-14) OpenAI. GPT-4.1 Mini OpenAI. GPT-4.1 Nano OpenAI. GPT-4 Turbo OpenAI. GPT-4o OpenAI. GPT-4o (2024-08-06) OpenAI. GPT-4o mini OpenAI. GPT-5 OpenAI. GPT-5 Mini OpenAI. GPT-5 Nano OpenAI. GPT-5.6 Sol OpenAI. GPT-5.6 Terra OpenAI. GPT-5.6 Luna OpenAI. GPT-6 Astra OpenAI. o1 OpenAI. o1-mini OpenAI. o1-preview OpenAI. o3 OpenAI. o3-mini OpenAI. o4-mini Anthropic. Claude 2.1 Anthropic. Claude 3.5 Haiku Anthropic. Claude 3.5 Sonnet Anthropic. Claude 3 Haiku Anthropic. Claude 3 Opus Anthropic. Claude 3 Sonnet Anthropic. Claude Fable 5.1 Anthropic. Claude Haiku 4.5 Anthropic. Claude Opus 5 Anthropic. Claude Opus 4.6 Anthropic. Claude Opus 4.8 Anthropic. Claude Sonnet 4.6 Anthropic. Claude Sonnet 5 Google. Gemini 1.0 Pro Google. Gemini 1.5 Flash Google. Gemini 1.5 Flash-8B Google. Gemini 1.5 Pro Google. Gemini 2.0 Flash Google. Gemini 2.0 Flash-Lite Google. Gemini 2.0 Pro Exp Google. Gemini 2.5 Flash Google. Gemini 2.5 Flash-Lite Google. Gemini 2.5 Pro Google. Gemini 3.1 Flash-Lite Google. Gemini 3.1 Pro Google. Gemini 3 Flash Google. Gemini 3.8 Flash Google. Gemma 3 27B DeepSeek. DeepSeek Chat DeepSeek. DeepSeek Coder DeepSeek. DeepSeek R1 DeepSeek. DeepSeek R1 Distill Qwen 32B DeepSeek. DeepSeek Reasoner DeepSeek. DeepSeek V3 DeepSeek. DeepSeek V4 Flash DeepSeek. DeepSeek V4 Pro DeepSeek. DeepSeek VL xAI. Grok 2 xAI. Grok 2 Mini xAI. Grok 2 Vision xAI. Grok 3 xAI. Grok 3 Mini xAI. Grok 4.6 xAI. Grok 4 xAI. Grok 4 Fast xAI. Grok Beta xAI. Grok Vision Beta Mistral. Codestral Mistral. Ministral 8B Mistral. Mistral Large (2407) Mistral. Mistral Large 3 Mistral. Mistral Medium 3 Mistral. Mistral Nemo Mistral. Mistral Saba Mistral. Mistral Small Mistral. Mistral Tiny Mistral. Open Mixtral 8x22B Mistral. Open Mixtral 8x7B Mistral. Pixtral 12B Groq. DeepSeek R1 Distill Llama 70B (Groq) Groq. Gemma 7B (Groq) Groq. Gemma 2 9B (Groq) Groq. Llama 3.1 8B Instant (Groq) Groq. Llama 3.2 11B Vision (Groq) Groq. Llama 3.2 1B (Groq) Groq. Llama 3.2 3B (Groq) Groq. Llama 3.3 70B (Groq) Groq. Llama 3.3 70B SpecDec (Groq) Groq. Llama 4 Scout (Groq) Groq. Llama Guard 3 8B (Groq) Groq. Mixtral 8x7B (Groq) Groq. QwQ 32B (Groq) Groq. Qwen3 32B (Groq) Cohere. Command A Cohere. Command Light Cohere. Command Nightly Cohere. Command R Cohere. Command R (08-2024) Cohere. Command R+ Cohere. Command R+ (08-2024) Cohere. Command R7B Together AI. DeepSeek R1 (Together) Together AI. DeepSeek V3 (Together) Together AI. Llama 3.1 70B (Together) Together AI. Llama 3.1 8B (Together) Together AI. Llama 3.2 3B (Together) Together AI. Llama 3.3 70B (Together) Together AI. Llama 4 Maverick (Together) Together AI. Mistral 7B (Together) Together AI. Mixtral 8x22B (Together) Together AI. Qwen2.5 72B (Together) Together AI. Qwen2.5 7B (Together) Together AI. Qwen2.5 Coder 32B (Together) Fireworks. DeepSeek R1 (Fireworks) Fireworks. DeepSeek V3 (Fireworks) Fireworks. Llama 3.1 70B (Fireworks) Fireworks. Llama 3.1 8B (Fireworks) Fireworks. Llama 3.2 3B (Fireworks) Fireworks. Llama 3.3 70B (Fireworks) Fireworks. Llama 4 Maverick (Fireworks) Fireworks. Mixtral 8x22B (Fireworks) Fireworks. MythoMax L2 13B (Fireworks) Fireworks. Qwen2.5 32B (Fireworks) Fireworks. Qwen2.5 72B (Fireworks) OpenAI · input $0.400/1M · output $1.60/1M · cached $0.100/1M · rates checked 2026-09-18
Corpus indexing Per query retrieval + answer Retrieved context ≈ 2,560 tokens (chunk × k).
Input tokens / query 3,040
Generation / query $0.001856
Query embed / query $0.0000016
Index once $0.04
Generation / month $55.68
Query embeds / month $0.048
Index amortized / month $0.004
Total / month $55.732 (100% generation) How to use this Pick the generation model you will call. Set corpus size and your real embedding $/1M from the embed provider. Set chunk size and top-k so retrieved tokens match production. Enter system, question, answer length, and monthly queries. Read which line dominates. Open GPT-4.1 Mini in cost calculator · RAG cost guide · Reduce RAG costs · Context window fit
Sources and references Official documentation used for definitions, counting methods, or rate cards. Always confirm critical budgets on the provider page.
FAQ
What are the cost components of a RAG system? Indexing (embed the corpus), tiny per-query embedding of the question, and generation (system + question + retrieved chunks + answer). Generation usually dominates.
Why is generation more expensive than embedding? Embedding rates are often cents per million tokens and mostly one-time. Generation bills retrieved context as input on every query at LLM rates.
How do I reduce RAG costs? Use a cheaper generation model when quality allows, lower top-k or chunk size, cache a stable system prompt, and cap answer length.
Are embedding rates from the TokenCalculator catalog? No. Edit the embedding $/1M field yourself from your embed provider. Generation uses published catalog rates for the selected LLM.