Llama 4 Maverick Pricing & Token Economics
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion act...
Normalized Token Pricing Matrix
Real-time input, completion, and cache rates normalized to $/1M tokens.
Cost for prompts, instructions, document retrieval, and system context sent into the model.
Cost for generated text, code, tool calls, and structured JSON responses returned by the model.
Standard rates apply. No separate prompt cache discount published via public API routes.
Standard industry benchmark assuming 75% prompt reads and 25% completion generation.
Llama 4 Maverick Monthly Bill Estimator
Estimated monthly costs across four standardized enterprise and developer workloads.
| Workload Scenario | Monthly Token Volume | Standard Bill | With Prompt Caching |
|---|---|---|---|
| Customer Support Chatbot 10M prompt tokens + 2.5M output tokens per month | 12.5M tokens (10M in / 2.5M out) | $3.51/mo | — |
| Document Analysis & RAG 30M prompt tokens (PDFs/context) + 3M output summaries | 33.0M tokens (30M in / 3M out) | $7.58/mo | — |
| Coding Agent / IDE Swarm 60M prompt tokens (repo context) + 15M code completions | 75.0M tokens (60M in / 15M out) | $21.04/mo | — |
| Enterprise Agentic Workflow 120M prompt tokens (reasoning loops) + 30M output actions | 150.0M tokens (120M in / 30M out) | $42.08/mo | — |
Need custom token inputs or prompt volumes?
Use our interactive multi-model token calculator to test exact prompt and completion sliders.
Architecture & Capabilities
Technical parameters, supported modalities, and optimal developer use cases.
🎯 Recommended Use Cases
Self-hostable enterprise intelligence, private on-prem workflows, fine-tuning
Supported Modalities
📐 Context & Output Limits
- Maximum Context Window: 1,048,576 tokens (1.0M)
- Maximum Output Length: 16,384 tokens
- Speed / Latency Tier: Fast
- Official Provider: Meta (Hosted)
Frequently Asked Questions
Everything you need to know about Llama 4 Maverick API pricing and token calculations.
How much does Llama 4 Maverick cost per 1M tokens?
Llama 4 Maverick charges $0.19 per 1M input tokens and $0.65 per 1M output tokens. For standard balanced workloads (3:1 input:output ratio), the effective blended rate is $0.304/1M tokens.
Does Llama 4 Maverick support prompt caching?
Standard prompt rates apply. Llama 4 Maverick does not offer a separate discounted prompt caching tier via standard public endpoints.
What is Llama 4 Maverick's maximum context window?
Llama 4 Maverick supports up to 1,048,576 tokens (1.0M) in its context window, allowing you to ingest large files and multi-turn chat history with up to 16,384 tokens generated per response.
How does Llama 4 Maverick compare to other models?
With a CostRatio rating of A, Llama 4 Maverick provides competitive token economics for open weights tasks. Review the peer comparison table above to compare input and completion rates directly against alternative models.
Explore all latest-generation models and interactive calculators
View Full AI Model Matrix →