Home / Pricing / Long replies

Cheapest models for long replies

Ranked by output rate per 1M tokens. The example uses 1000 input tokens and 2000 output tokens, which is closer to support chat and writing than a short completion.

128 models · rates checked . Price rank only, not a quality score.

Ranked by published rates

#ModelProviderContextInput / 1MOutput / 1MLong example10k requests
1Llama 3.2 1B (Groq)Groq128K$0.040/1M$0.040/1M$0.00012$1.2Cost
2Llama 3.2 3B (Groq)Groq128K$0.060/1M$0.060/1M$0.00018$1.8Cost
3Llama 3.2 3B (Together)Together AI128K$0.060/1M$0.060/1M$0.00018$1.8Cost
4Gemma 7B (Groq)Groq8K$0.070/1M$0.070/1M$0.00021$2.1Cost
5Llama 3.1 8B Instant (Groq)Groq128K$0.050/1M$0.080/1M$0.00021$2.1Cost
6Llama 3.2 3B (Fireworks)Fireworks128K$0.100/1M$0.100/1M$0.0003$3.00Cost
7Gemini 1.5 Flash-8BGoogle1M$0.037/1M$0.150/1M$0.000337$3.37Cost
8Command R7BCohere128K$0.037/1M$0.150/1M$0.000337$3.37Cost
9Ministral 8BMistral128K$0.150/1M$0.150/1M$0.00045$4.5Cost
10Mistral NemoMistral128K$0.150/1M$0.150/1M$0.00045$4.5Cost
11Pixtral 12BMistral128K$0.150/1M$0.150/1M$0.00045$4.5Cost
12Llama 3.2 11B Vision (Groq)Groq128K$0.180/1M$0.180/1M$0.00054$5.4Cost
13Llama 3.1 8B (Together)Together AI128K$0.180/1M$0.180/1M$0.00054$5.4Cost
14Gemma 2 9B (Groq)Groq8K$0.200/1M$0.200/1M$0.0006$6.00Cost
15Llama Guard 3 8B (Groq)Groq8K$0.200/1M$0.200/1M$0.0006$6.00Cost
16Mistral 7B (Together)Together AI33K$0.200/1M$0.200/1M$0.0006$6.00Cost
17Llama 3.1 8B (Fireworks)Fireworks128K$0.200/1M$0.200/1M$0.0006$6.00Cost
18MythoMax L2 13B (Fireworks)Fireworks4K$0.200/1M$0.200/1M$0.0006$6.00Cost
19Mixtral 8x7B (Groq)Groq33K$0.240/1M$0.240/1M$0.00072$7.2Cost
20Mistral TinyMistral33K$0.250/1M$0.250/1M$0.00075$7.5Cost
21DeepSeek CoderDeepSeek128K$0.140/1M$0.280/1M$0.0007$7.00Cost
22DeepSeek VLDeepSeek4K$0.140/1M$0.280/1M$0.0007$7.00Cost
23Gemini 1.5 FlashGoogle1M$0.075/1M$0.300/1M$0.000675$6.75Cost
24Gemini 2.0 Flash-LiteGoogle1M$0.075/1M$0.300/1M$0.000675$6.75Cost
25Qwen2.5 7B (Together)Together AI33K$0.300/1M$0.300/1M$0.0009$9.00Cost
26Llama 4 Scout (Groq)Groq128K$0.110/1M$0.340/1M$0.00079$7.9Cost
27QwQ 32B (Groq)Groq128K$0.290/1M$0.390/1M$0.00107$10.7Cost
28GPT-5 NanoOpenAI128K$0.050/1M$0.400/1M$0.00085$8.5Cost
29GPT-4.1 NanoOpenAI1M$0.100/1M$0.400/1M$0.0009$9.00Cost
30Gemini 2.0 FlashGoogle1M$0.100/1M$0.400/1M$0.0009$9.00Cost
31Gemini 2.5 Flash-LiteGoogle1M$0.100/1M$0.400/1M$0.0009$9.00Cost
32Gemma 3 27BGoogle128K$0.200/1M$0.400/1M$0.001$10.00Cost
33DeepSeek ChatDeepSeek128K$0.280/1M$0.420/1M$0.00112$11.2Cost
34Grok 2 MinixAI131K$0.200/1M$0.500/1M$0.0012$12.00Cost
35Grok 3 MinixAI131K$0.300/1M$0.500/1M$0.0013$13.00Cost
36Qwen3 32B (Groq)Groq131K$0.290/1M$0.590/1M$0.00147$14.7Cost
37GPT-4o miniOpenAI128K$0.150/1M$0.600/1M$0.00135$13.5Cost
38Mistral SmallMistral128K$0.150/1M$0.600/1M$0.00135$13.5Cost
39Command RCohere128K$0.150/1M$0.600/1M$0.00135$13.5Cost
40Command R (08-2024)Cohere128K$0.150/1M$0.600/1M$0.00135$13.5Cost
41Mistral SabaMistral33K$0.200/1M$0.600/1M$0.0014$14.00Cost
42DeepSeek R1 Distill Qwen 32BDeepSeek128K$0.300/1M$0.600/1M$0.0015$15.00Cost
43Command LightCohere4K$0.300/1M$0.600/1M$0.0015$15.00Cost
44Open Mixtral 8x7BMistral33K$0.700/1M$0.700/1M$0.0021$21.00Cost
45Llama 3.3 70B (Groq)Groq128K$0.590/1M$0.790/1M$0.00217$21.7Cost
46Qwen2.5 Coder 32B (Together)Together AI33K$0.800/1M$0.800/1M$0.0024$24.00Cost
47Llama 4 Maverick (Together)Together AI1M$0.270/1M$0.850/1M$0.00197$19.7Cost
48Llama 4 Maverick (Fireworks)Fireworks1M$0.220/1M$0.880/1M$0.00198$19.8Cost
49Llama 3.1 70B (Together)Together AI128K$0.880/1M$0.880/1M$0.00264$26.4Cost
50Llama 3.3 70B (Together)Together AI128K$0.880/1M$0.880/1M$0.00264$26.4Cost
51CodestralMistral256K$0.300/1M$0.900/1M$0.0021$21.00Cost
52DeepSeek R1 (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
53DeepSeek V3 (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
54Llama 3.1 70B (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
55Llama 3.3 70B (Fireworks)Fireworks128K$0.900/1M$0.900/1M$0.0027$27.00Cost
56Mixtral 8x22B (Fireworks)Fireworks66K$0.900/1M$0.900/1M$0.0027$27.00Cost
57Qwen2.5 32B (Fireworks)Fireworks33K$0.900/1M$0.900/1M$0.0027$27.00Cost
58Qwen2.5 72B (Fireworks)Fireworks33K$0.900/1M$0.900/1M$0.0027$27.00Cost
59Llama 3.3 70B SpecDec (Groq)Groq8K$0.590/1M$0.990/1M$0.00257$25.7Cost
60DeepSeek R1 Distill Llama 70B (Groq)Groq128K$0.750/1M$0.990/1M$0.00273$27.3Cost
61DeepSeek V3DeepSeek128K$0.270/1M$1.10/1M$0.00247$24.7Cost
62GPT-5.6 LunaOpenAI1.1M$0.200/1M$1.20/1M$0.0026$26.00Cost
63DeepSeek V4 FlashDeepSeek1M$0.300/1M$1.20/1M$0.0027$27.00Cost
64Mixtral 8x22B (Together)Together AI66K$1.20/1M$1.20/1M$0.0036$36.00Cost
65Qwen2.5 72B (Together)Together AI33K$1.20/1M$1.20/1M$0.0036$36.00Cost
66Claude 3 HaikuAnthropic200K$0.250/1M$1.25/1M$0.00275$27.5Cost
67DeepSeek V3 (Together)Together AI128K$1.25/1M$1.25/1M$0.00375$37.5Cost
68Gemini 3.1 Flash-LiteGoogle1M$0.250/1M$1.50/1M$0.00325$32.5Cost
69GPT-3.5 TurboOpenAI16K$0.500/1M$1.50/1M$0.0035$35.00Cost
70Gemini 1.0 ProGoogle33K$0.500/1M$1.50/1M$0.0035$35.00Cost
71Mistral Large 3Mistral128K$0.500/1M$1.50/1M$0.0035$35.00Cost
72GPT-4.1 MiniOpenAI1M$0.400/1M$1.60/1M$0.0036$36.00Cost
73GPT-5 MiniOpenAI128K$0.250/1M$2.00/1M$0.00425$42.5Cost
74Command NightlyCohere128K$1.00/1M$2.00/1M$0.005$50.00Cost
75DeepSeek R1DeepSeek128K$0.550/1M$2.19/1M$0.00493$49.3Cost
76DeepSeek ReasonerDeepSeek128K$0.550/1M$2.19/1M$0.00493$49.3Cost
77Gemini 2.5 FlashGoogle1M$0.300/1M$2.50/1M$0.0053$53.00Cost
78Grok 3xAI131K$1.25/1M$2.50/1M$0.00625$62.5Cost
79Grok 4xAI256K$1.25/1M$2.50/1M$0.00625$62.5Cost
80Grok 4 FastxAI256K$1.25/1M$2.50/1M$0.00625$62.5Cost
81Gemini 3 FlashGoogle1M$0.500/1M$3.00/1M$0.0065$65.00Cost
82Gemini 3.8 FlashGoogle1.0M$0.750/1M$3.75/1M$0.00825$82.5Cost
83DeepSeek V4 ProDeepSeek1M$1.32/1M$3.96/1M$0.00924$92.4Cost
84Claude 3.5 HaikuAnthropic200K$0.800/1M$4.00/1M$0.0088$88.00Cost
85o1-miniOpenAI128K$1.10/1M$4.40/1M$0.0099$99.00Cost
86o3-miniOpenAI200K$1.10/1M$4.40/1M$0.0099$99.00Cost
87o4-miniOpenAI200K$1.10/1M$4.40/1M$0.0099$99.00Cost
88Claude Haiku 4.5Anthropic200K$1.00/1M$5.00/1M$0.011$110.00Cost
89Gemini 1.5 ProGoogle2M$1.25/1M$5.00/1M$0.01125$112.5Cost
90Gemini 2.0 Pro ExpGoogle2M$1.25/1M$5.00/1M$0.01125$112.5Cost
91Grok 4.6xAI500K$2.00/1M$6.00/1M$0.014$140.00Cost
92Mistral Large (2407)Mistral128K$2.00/1M$6.00/1M$0.014$140.00Cost
93Open Mixtral 8x22BMistral66K$2.00/1M$6.00/1M$0.014$140.00Cost
94DeepSeek R1 (Together)Together AI128K$3.00/1M$7.00/1M$0.017$170.00Cost
95Mistral Medium 3Mistral128K$1.50/1M$7.50/1M$0.0165$165.00Cost
96GPT-4.1OpenAI1M$2.00/1M$8.00/1M$0.018$180.00Cost
97GPT-4.1 (2025-04-14)OpenAI1M$2.00/1M$8.00/1M$0.018$180.00Cost
98o3OpenAI200K$2.00/1M$8.00/1M$0.018$180.00Cost
99GPT-5OpenAI400K$1.25/1M$10.00/1M$0.02125$212.5Cost
100Gemini 2.5 ProGoogle1M$1.25/1M$10.00/1M$0.02125$212.5Cost
101Claude Sonnet 5Anthropic1M$2.00/1M$10.00/1M$0.022$220.00Cost
102Grok 2xAI131K$2.00/1M$10.00/1M$0.022$220.00Cost
103Grok 2 VisionxAI33K$2.00/1M$10.00/1M$0.022$220.00Cost
104GPT-4oOpenAI128K$2.50/1M$10.00/1M$0.0225$225.00Cost
105GPT-4o (2024-08-06)OpenAI128K$2.50/1M$10.00/1M$0.0225$225.00Cost
106Command ACohere256K$2.50/1M$10.00/1M$0.0225$225.00Cost
107Command R+Cohere128K$2.50/1M$10.00/1M$0.0225$225.00Cost
108Command R+ (08-2024)Cohere128K$2.50/1M$10.00/1M$0.0225$225.00Cost
109GPT-5.6 TerraOpenAI1.1M$2.00/1M$12.00/1M$0.026$260.00Cost
110Gemini 3.1 ProGoogle1M$2.00/1M$12.00/1M$0.026$260.00Cost
111Claude 3.5 SonnetAnthropic200K$3.00/1M$15.00/1M$0.033$330.00Cost
112Claude 3 SonnetAnthropic200K$3.00/1M$15.00/1M$0.033$330.00Cost
113Claude Sonnet 4.6Anthropic1M$3.00/1M$15.00/1M$0.033$330.00Cost
114ChatGPT-4o LatestOpenAI128K$5.00/1M$15.00/1M$0.035$350.00Cost
115Grok BetaxAI131K$5.00/1M$15.00/1M$0.035$350.00Cost
116Grok Vision BetaxAI8K$5.00/1M$15.00/1M$0.035$350.00Cost
117GPT-5.6 SolOpenAI1.1M$4.00/1M$20.00/1M$0.044$440.00Cost
118Claude 2.1Anthropic200K$8.00/1M$24.00/1M$0.056$560.00Cost
119Claude Opus 5Anthropic1M$5.00/1M$25.00/1M$0.055$550.00Cost
120Claude Opus 4.6Anthropic1M$5.00/1M$25.00/1M$0.055$550.00Cost
121Claude Opus 4.8Anthropic1M$5.00/1M$25.00/1M$0.055$550.00Cost
122GPT-4 TurboOpenAI128K$10.00/1M$30.00/1M$0.07$700.00Cost
123GPT-6 AstraOpenAI1.1M$10.00/1M$50.00/1M$0.11$1100.00Cost
124Claude Fable 5.1Anthropic1M$10.00/1M$50.00/1M$0.11$1100.00Cost
125o1OpenAI200K$15.00/1M$60.00/1M$0.135$1350.00Cost
126o1-previewOpenAI128K$15.00/1M$60.00/1M$0.135$1350.00Cost
127GPT-4OpenAI8K$30.00/1M$60.00/1M$0.15$1500.00Cost
128Claude 3 OpusAnthropic200K$15.00/1M$75.00/1M$0.165$1650.00Cost

Related tools

LLM cost calculator · Cheapest LLM API · Compare model prices · Cache savings · Batch pricing

Related guides

How LLM API pricing works · How prompt cost is calculated · Prompt caching explained · Exact vs Approx

Sources and references

Official documentation used for definitions, counting methods, or rate cards. Always confirm critical budgets on the provider page.

Pricing rank FAQ

Why rank by output?
Long answers are billed on output tokens, which are usually more expensive than input. A cheap input rate can still be a costly chat model.
What is the long example?
1000 input tokens and 2000 output tokens at standard rates, no cache and no batch.
Can I change the mix?
Yes. Open Cost on any row and set your own output size and monthly volume.