Compare AI Model Token Costs & Inference Ratios
Stop guessing your monthly LLM bill. We normalize real-time prompt, completion, and cached token pricing across top models—from Claude Sonnet 5 and Gemini 3.1 Pro to DeepSeek V4 and o3-mini.
Interactive LLM Monthly Bill Estimator
Adjust your monthly token volume to see the true cost difference across every model.
Llama 4 Scout
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B.
Qwen3 Coder Next
Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows.
Llama 4 Maverick
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion act...
DeepSeek V4 Flash Vision Exp
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while matching the base model on text ...
GPT-5.6 Luna
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series.
Devstral 2
Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding.
Gemini 3.5 Flash Lite
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities.
DeepSeek V4 Pro
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
DeepSeek R1
May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens.
Gemini 3.8 Flash
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
OpenAI o3 Mini High
OpenAI o3-mini-high is the same model as o3-mini with reasoning_effort set to high.
Claude Haiku 4.5
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models.
Mistral Medium 3.5
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI.
Grok 4.6
Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Qwen3.8 Max
Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview.
OpenAI o3
o3 is a well-rounded and powerful model across domains.
Claude Sonnet 5
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work.
Gemini 3.1 Pro (Preview)
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage a...
GPT-5.6 Terra
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier.
Claude Opus 5
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work.
Full Token Pricing Benchmark Table
Comprehensive breakdown of input rates, output rates, prompt cache discounts, and context limits.
| Model & Vendor | Category | Input / 1M | Output / 1M | Cached Read / 1M | Context Window | CostRatio Score |
|---|---|---|---|---|---|---|
| Llama 4 Scout Meta (Hosted) | Open Weights | $0.100 | $0.300 | — | 1.3M | A+ |
| Qwen3 Coder Next Qwen / Alibaba | Coding & Specialized | $0.120 | $0.800 | $0.070 | 262K | A+ |
| Llama 4 Maverick Meta (Hosted) | Open Weights | $0.200 | $0.696 | — | 1M | A |
| DeepSeek V4 Flash Vision Exp DeepSeek | Budget & Fast | $0.220 | $0.660 | $0.0070 | 1M | A |
| GPT-5.6 Luna OpenAI | Budget & Fast | $0.200 | $1.20 | $0.020 | 1.1M | A |
| Devstral 2 Mistral AI | Coding & Specialized | $0.400 | $2.00 | $0.040 | 262K | A |
| Gemini 3.5 Flash Lite Google Cloud / AI Studio | Budget & Fast | $0.300 | $2.50 | $0.030 | 1M | B |
| DeepSeek V4 Pro DeepSeek | Budget & Fast | $0.580 | $1.74 | $0.019 | 1M | B |
| DeepSeek R1 DeepSeek | Reasoning & Thinking | $0.500 | $2.15 | $0.350 | 164K | A+ |
| Gemini 3.8 Flash Google Cloud / AI Studio | Budget & Fast | $0.750 | $3.75 | $0.075 | 1M | B |
| OpenAI o3 Mini High OpenAI | Reasoning & Thinking | $1.10 | $4.40 | $0.550 | 200K | A |
| Claude Haiku 4.5 Anthropic | Budget & Fast | $1.00 | $5.00 | $0.100 | 200K | C |
| Mistral Medium 3.5 Mistral AI | Frontier Flagship | $1.50 | $7.50 | — | 262K | A+ |
| Grok 4.6 xAI | Frontier Flagship | $2.00 | $6.00 | $0.500 | 500K | A+ |
| Qwen3.8 Max Qwen / Alibaba | Open Weights | $2.00 | $6.00 | $0.250 | 1M | C |
| OpenAI o3 OpenAI | Reasoning & Thinking | $2.00 | $8.00 | $0.500 | 200K | A |
| Claude Sonnet 5 Anthropic | Frontier Flagship | $2.00 | $10.00 | $0.200 | 1M | A |
| Gemini 3.1 Pro (Preview) Google Cloud / AI Studio | Frontier Flagship | $2.00 | $12.00 | $0.200 | 1M | A |
| GPT-5.6 Terra OpenAI | Frontier Flagship | $2.00 | $12.00 | $0.200 | 1.1M | A |
| Claude Opus 5 Anthropic | Reasoning & Thinking | $5.00 | $25.00 | $0.500 | 1M | B |
Frequently Asked Questions
Everything you need to know about AI inference pricing and token economics.
What is the cheapest AI model API in 2026?
As of 2026, the cheapest mainstream LLM APIs are DeepSeek V4 Flash ($0.065/1M input, $0.18/1M output), Llama 4 Scout ($0.10/1M input, $0.30/1M output), and Gemini 3.5 Flash Lite ($0.30/1M input, $2.50/1M output).
How much cheaper is DeepSeek R1 than OpenAI o1?
DeepSeek R1 charges approximately $0.70 per 1M input tokens and $2.50 per 1M output tokens. OpenAI o1 charges $15.00 per 1M input and $60.00 per 1M output. This makes DeepSeek R1 roughly 96% cheaper while delivering competitive reasoning performance on math and coding benchmarks.
What is a blended token cost?
A blended token cost calculates the average price per 1M tokens assuming a standard 3:1 input-to-output ratio. Since output tokens are typically 3x to 5x more expensive than prompt tokens, blended pricing gives a more realistic picture of your actual monthly bill.
What is prompt caching and how does it reduce LLM costs?
Prompt caching stores frequently repeated context (such as system instructions, document repositories, or codebases) on the provider's servers. Providers like Anthropic, OpenAI, DeepSeek, and Google discount cached prompt reads by 50% to 90%, drastically lowering costs for conversational agents and RAG applications.
Need Cloud Storage or Zero-Egress Object Storage?
Compare $/TB cost ratios across Google One, iCloud, Dropbox, Wasabi, and Backblaze B2.