Anthropic Budget & Fast Live Rates β€’ September 2026

Claude Haiku 4.5 Pricing & Token Economics

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models.

Context Window 200K
Max Completion 64K Tokens
Inference Speed Ultra Fast
CostRatio Score C

Normalized Token Pricing Matrix

Real-time input, completion, and cache rates normalized to $/1M tokens.

πŸ“₯ Prompt Tokens
$ 1.00 / 1M tokens

Cost for prompts, instructions, document retrieval, and system context sent into the model.

~$0.0010 per 1K tokens
πŸ“€ Completion Tokens
$ 5.00 / 1M tokens

Cost for generated text, code, tool calls, and structured JSON responses returned by the model.

~$0.0050 per 1K tokens
⚑ Prompt Caching
$ 0.100 / 1M tokens

Massive 90% savings on repeated context, documents, and system instructions.

Cached Read Discount
βš–οΈ 3:1 Blended Benchmark
$ 2.000 / 1M tokens

Standard industry benchmark assuming 75% prompt reads and 25% completion generation.

Weighted Benchmark

Claude Haiku 4.5 Monthly Bill Estimator

Estimated monthly costs across four standardized enterprise and developer workloads.

Workload Scenario Monthly Token Volume Standard Bill With Prompt Caching
πŸ’¬
Customer Support Chatbot
10M prompt tokens + 2.5M output tokens per month
12.5M tokens (10M in / 2.5M out) $22.50/mo $16.20/mo -28%
πŸ“‘
Document Analysis & RAG
30M prompt tokens (PDFs/context) + 3M output summaries
33.0M tokens (30M in / 3M out) $45.00/mo $26.10/mo -42%
πŸ’»
Coding Agent / IDE Swarm
60M prompt tokens (repo context) + 15M code completions
75.0M tokens (60M in / 15M out) $135.00/mo $97.20/mo -28%
πŸ€–
Enterprise Agentic Workflow
120M prompt tokens (reasoning loops) + 30M output actions
150.0M tokens (120M in / 30M out) $270.00/mo $194.40/mo -28%

Need custom token inputs or prompt volumes?

Use our interactive multi-model token calculator to test exact prompt and completion sliders.

Open Interactive Token Calculator β†’

Architecture & Capabilities

Technical parameters, supported modalities, and optimal developer use cases.

🎯 Recommended Use Cases

Sub-agents, real-time chat, fast triage, and bulk data processing

Supported Modalities

πŸ“ Text πŸ‘οΈ Vision & Images πŸ› οΈ Function / Tool Calling

πŸ“ Context & Output Limits

  • Maximum Context Window: 200,000 tokens (200K)
  • Maximum Output Length: 64,000 tokens
  • Speed / Latency Tier: Ultra Fast
  • Official Provider: Anthropic

Frequently Asked Questions

Everything you need to know about Claude Haiku 4.5 API pricing and token calculations.

How much does Claude Haiku 4.5 cost per 1M tokens?

Claude Haiku 4.5 charges $1.00 per 1M input tokens and $5.00 per 1M output tokens. For standard balanced workloads (3:1 input:output ratio), the effective blended rate is $2.000/1M tokens.

Does Claude Haiku 4.5 support prompt caching?

Yes. Cached prompt reads are billed at $0.1000 per 1M tokens, saving you up to 90% on repeated context, RAG documents, and system instructions.

What is Claude Haiku 4.5's maximum context window?

Claude Haiku 4.5 supports up to 200,000 tokens (200K) in its context window, allowing you to ingest large files and multi-turn chat history with up to 64,000 tokens generated per response.

How does Claude Haiku 4.5 compare to other models?

With a CostRatio rating of C, Claude Haiku 4.5 provides competitive token economics for budget & fast tasks. Review the peer comparison table above to compare input and completion rates directly against alternative models.

Explore all latest-generation models and interactive calculators

View Full AI Model Matrix β†’