Claude Sonnet 4.6
Claude Sonnet 4.6 is a verified AI model profile covering official specifications, pricing or access, capabilities, practical use…
Claude Haiku 4.5 is Anthropic's fastest active Claude model with 200K context, 64K output, extended thinking and $1/M input pricing.
Claude Haiku 4.5 is Anthropic's fastest active Claude model, designed for high-volume, interactive and cost-sensitive workloads that still need strong reasoning.
Claude Haiku 4.5 is Anthropic’s fastest active Claude model and the cost-oriented option in a lineup now dominated by 1M-context Sonnet, Opus and Fable models.
Haiku exists for a different optimization target than Claude’s larger models. Instead of maximizing deep reasoning regardless of latency, it prioritizes responsive interaction and high-volume processing while retaining enough intelligence for tool use, extraction, classification and many support tasks. Anthropic describes it as the fastest model with near-frontier intelligence, making it especially relevant to products where users notice every extra second of response time.
The model provides a 200K-token context window and up to 64K output tokens. It accepts text and image input and returns text. Haiku 4.5 supports manual extended thinking, which developers can enable with an explicit thinking budget for more difficult reasoning. This differs from newer Claude 4.6/5 models, which use adaptive thinking. Tool use and structured outputs are supported, while the reliable knowledge cutoff is February 2025.
A high-volume support system can use Haiku for intent classification, retrieval synthesis and routine responses. In a multi-agent architecture, Haiku can operate as the fast executor that handles repetitive sub-tasks while Sonnet or Opus acts as an advisor for difficult decisions. It also works well for interactive forms, extraction pipelines and lightweight vision tasks where the 200K context window is sufficient.
Anthropic lists Claude Haiku 4.5 as Active (latest) and supports it through the Claude API and major cloud platforms. Its dated model ID, claude-haiku-4-5-20251001, gives applications a fixed version target, while aliases are available on supported platforms. Anthropic’s current lifecycle table indicates it remains active at least through the near term.
Pricing is $1/M input and $5/M output. Cache reads cost $0.10/M, and Batch API jobs receive a 50% input/output discount. This makes Haiku materially cheaper than Sonnet 5 at $2/$10 and much cheaper than Opus/Fable. For workloads with large repeated prompts, caching plus Haiku can further reduce effective cost.
Choose Haiku when speed, throughput and cost are primary and the workload is easy to evaluate or escalate. Choose Sonnet 5 for harder coding and agentic work or when the 1M context window is important. Use Opus/Fable only when the task value justifies premium reasoning. Haiku is often the right worker model in a routing architecture rather than the only model in the system.
Haiku’s February 2025 reliable knowledge cutoff is notably older than current top Claude models, so fresh information should come from retrieval or tools. Manual extended thinking also behaves differently from adaptive thinking on newer Claude generations. Developers migrating prompts between model classes should re-test sampling parameters, cache behavior, tool versions and latency rather than assuming configuration is interchangeable.
In addition, its main capabilities include 200K context, 64K output, Fast latency, Extended thinking, Vision and Tool use. For example, common use cases include High-volume apps, Support, Sub-agents, Extraction, Classification and Interactive assistants.
In practice, this model may suit Realtime applications, High-volume processing, Fast support assistants, Sub-agents, Cost-sensitive reasoning and Classification. Also, notable strengths include Anthropic's fastest current model, Low Claude API cost, Strong reasoning for its tier and Large 64K output limit. However, review trade-offs such as Less capable than Sonnet/Opus on hard multi-step tasks, Extended thinking affects cache behavior and Older cutoff makes current-data grounding important before adopting it.
Meanwhile, $1 per 1M input tokens and $5 per 1M output tokens. Cache reads cost $0.10/M and Batch API input/output receives a 50% discount. At $1/M input and $5/M output, Haiku 4.5 is half the standard Sonnet 5 token price and is well suited to sub-agents and high-throughput systems.
Use the official model website, official documentation, pricing or release source and additional primary source to confirm current availability, limits and pricing. Product details can change after publication, so rely on primary documentation for final decisions.
Next, continue your research in the AI models directory, Anthropic models and General Purpose Language Model models. Compare providers, pricing, modalities and practical limitations side by side to choose the right model for your workflow.
Classify and summarize these incoming support requests quickly.Act as a fast sub-agent and verify these facts against the provided context.Extract the requested data and return it in the required structure.$1 per 1M input tokens and $5 per 1M output tokens. Cache reads cost $0.10/M and Batch API input/output receives a 50% discount.
A practical high-throughput Claude model for applications that need more intelligence than basic classifiers without paying Sonnet or Opus rates.