Compare stack costs
Price AI models or tools in any stack category on the same usage. Every comparison is cost-only and states what could not be modeled.
Cost only. Every model below is priced on the exact same workload and sorted by price. This is not a quality, capability, or suitability ranking, and it does not imply these models are interchangeable.
| Model | Provider | Cost per request | Monthly cost | Source |
|---|---|---|---|---|
| GPT-5 nano | OpenAI | $0.000550 | $110.00 | Verified 2026-08-15 |
| DeepSeek V4 Flash | DeepSeek | $0.000700 | $140.00 | Verified 2026-08-15 |
| Gemini 2.5 Flash-Lite | $0.000700 | $140.00 | Verified 2026-08-15 | |
| GPT-4.1 nano | OpenAI | $0.000700 | $140.00 | Verified 2026-08-15 |
| GPT-4o mini | OpenAI | $0.001050 | $210.00 | Verified 2026-08-15 |
| Mistral Small 4 | Mistral AI | $0.001050 | $210.00 | Verified 2026-08-15 |
| GPT-5.6 Luna | OpenAI | $0.001800 | $360.00 | Verified 2026-08-15 |
| DeepSeek V4 Pro | DeepSeek | $0.002175 | $435.00 | Verified 2026-08-15 |
| Gemini 3.1 Flash-Lite | $0.002250 | $450.00 | Verified 2026-08-15 | |
| GPT-5 mini | OpenAI | $0.002750 | $550.00 | Verified 2026-08-15 |
| GPT-4.1 mini | OpenAI | $0.002800 | $560.00 | Verified 2026-08-15 |
| Mistral Large 3 | Mistral AI | $0.003000 | $600.00 | Verified 2026-08-15 |
| Gemini 2.5 Flash | $0.003400 | $680.00 | Verified 2026-08-15 | |
| Gemini 3.5 Flash-Lite | $0.003400 | $680.00 | Verified 2026-08-15 | |
| Gemini 3.7 Flash | $0.006000 | $1,200.00 | Verified 2026-08-15 | |
| Claude Haiku 4.5 | Anthropic | $0.008000 | $1,600.00 | Verified 2026-08-15 |
| Mistral Medium 3.5 | Mistral AI | $0.012000 | $2,400.00 | Verified 2026-08-15 |
| Gemini 3.5 Flash | $0.013500 | $2,700.00 | Verified 2026-08-15 | |
| GPT-5 | OpenAI | $0.013750 | $2,750.00 | Verified 2026-08-15 |
| GPT-5.1 | OpenAI | $0.013750 | $2,750.00 | Verified 2026-08-15 |
| GPT-4.1 | OpenAI | $0.014000 | $2,800.00 | Verified 2026-08-15 |
| o3 | OpenAI | $0.014000 | $2,800.00 | Verified 2026-08-15 |
| Claude Sonnet 5 | Anthropic | $0.016000 | $3,200.00 | Verified 2026-08-15 |
| GPT-4o | OpenAI | $0.017500 | $3,500.00 | Verified 2026-08-15 |
| GPT-5.6 Terra | OpenAI | $0.018000 | $3,600.00 | Verified 2026-08-15 |
| GPT-5.2 | OpenAI | $0.019250 | $3,850.00 | Verified 2026-08-15 |
| Claude Sonnet 4.5 | Anthropic | $0.024000 | $4,800.00 | Verified 2026-08-15 |
| Claude Sonnet 4.6 | Anthropic | $0.024000 | $4,800.00 | Verified 2026-08-15 |
| Claude Opus 4.5 | Anthropic | $0.040000 | $8,000.00 | Verified 2026-08-15 |
| Claude Opus 4.6 | Anthropic | $0.040000 | $8,000.00 | Verified 2026-08-15 |
| Claude Opus 4.7 | Anthropic | $0.040000 | $8,000.00 | Verified 2026-08-15 |
| Claude Opus 4.8 | Anthropic | $0.040000 | $8,000.00 | Verified 2026-08-15 |
| Claude Opus 5 | Anthropic | $0.040000 | $8,000.00 | Verified 2026-08-15 |
| GPT-5.6 Sol | OpenAI | $0.045000 | $9,000.00 | Verified 2026-08-15 |
Not included in these estimates
- Excludes taxes, contractual or volume discounts, and regional pricing.
- Excludes batch-processing discounts.
- Excludes server-side tool calls, web search, and other non-token charges.
- Text tokens only. Image, audio, and video pricing is not modeled.
- Cached rate shown is the cache-read (hit) price. Cache-write charges are not modeled.
- Cache writes, batch discounts, regional platform premiums, tool-specific fees, and taxes are not modeled.
- Cache writes, batch discounts, US-only inference premiums, tool-specific fees, and taxes are not modeled.
- Cache writes, batch discounts, data residency premiums, tool-specific fees, and taxes are not modeled.
- Cache writes, fast mode, batch discounts, data residency premiums, tool-specific fees, and taxes are not modeled.
- Cached rate shown is the cache-read (hit) price. Cache-write charges (1.25x for 5 minutes, 2x for 1 hour) are not modeled.
- Excludes fast mode premium pricing and the 1.1x US data-residency multiplier.
- Excludes the 1.1x US data-residency multiplier.
- Input rate is the cache-miss price; the cached rate is the cache-hit price.
- The vendor announced peak / off-peak billing effective 16 August 2026. Time-of-day pricing is not modeled; these are the rates published on the verification date.
- Thinking tokens are billed as output tokens; audio, image and video generation, cache storage, batch pricing, grounding, maps, tool-use fees, regional pricing, and taxes are not modeled.
- Audio, image and video generation, cache storage, batch pricing, grounding, maps, tool-use fees, regional pricing, and taxes are not modeled.
- Rates are the promotional prices published through 31 December 2026. The page lists higher rates ($1.50 in / $7.50 out per million) from 1 January 2027; the future price is not modeled.
- No cached-input rate is published for this model, so prompt caching is not modeled.
- Batch, priority processing, long-context fine-tuning premiums, image input, tool-call fees, and taxes are not modeled.
- Batch, priority processing, image input, audio, fine-tuning, tool-call fees, and taxes are not modeled.
- Batch, flex, priority processing, image input, tool-call fees, fine-tuning, and taxes are not modeled.
- Prompts above 272K input tokens use higher rates and are not modeled.
- Cache writes, taxes, contractual or volume discounts, and regional pricing are not modeled.
- Reasoning tokens are billed as output tokens; batch, flex, priority processing, image input, tool-call fees, and taxes are not modeled.