Friday, October 9, 2026

Today’s Edition

AI Intel Report

MARKETS —

Frontier Models

Claude Haiku 5.5 Slashes Small-Model Costs 75% in Agentic AI Price War

Anthropic's new small model undercuts its predecessor by roughly 75% on average, pairing a 1 million-token context window with agentic coding and computer-use gains.

9 MIN READ
Wide shot of a foggy data center campus with a technician walking, representing Anthropic's cloud infrastructure.
Illustration: AI Intel Report

Claude Haiku 5.5 is Anthropic's cheapest, fastest, and most capable small model, released Oct. 7, 2026, at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens.

Claude Haiku 5.5, released Oct. 7, 2026, is the first Anthropic small model to include adaptive thinking, a 1 million-token context window, and a 128,000-token output limit. The company said requests with prompts up to 100,000 tokens make up about 90% of traffic to its previous Haiku model, and it set the new model's list price at $0.10 per million input tokens and $0.50 per million output tokens for that range. Prompts longer than 100,000 tokens are priced 50% lower, according to Anthropic. The launch intensifies an AI price war in which Google and Mistral have made low-cost small models a centerpiece of their enterprise agent strategies.

Amazon Web Services, in a blog post announcing availability on Amazon Bedrock, called Claude Haiku 5.5 the fastest and most efficient model in the Claude 5.5 family, built for subagents and high-volume, cost-sensitive work. AWS said the model costs around 75% less than Claude Haiku 4.5 for most tasks. The combination of lower price and faster inference targets the repetitive, parallel operations that dominate agent architectures, where a single task can fan out into hundreds of model calls.

How does Claude Haiku 5.5 fit into Anthropic's model lineup?

Claude Haiku 5.5 sits at the bottom of Anthropic's Claude 5.5 family, beneath Claude Sonnet 5.5 and Claude Opus 5.5. Anthropic describes Haiku as a small model, but the label refers to cost and speed rather than raw capability. The model is designed to run as a subagent alongside Opus 5.5 and Sonnet 5.5, handling high-volume subtasks such as data extraction, code triage, and UI interaction while a larger model manages the overall plan. Anthropic said Haiku 5.5 is its most capable Haiku model across coding, tool use, computer use, and agentic tasks.

  1. A 1 million-token context window, giving subagents access to long documents and large codebases in a single call.
  2. Up to 128,000 output tokens, enough for structured documents and long generated responses.
  3. Adaptive thinking with effort controls, a first for the Haiku line, allowing developers to trade latency for reasoning depth.
  4. Pricing of $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens.
  5. Availability on Claude Platform, Amazon Bedrock, Google Cloud, and Microsoft Foundry.

Anthropic said Haiku 5.5 is especially good value when used for tasks with prompts up to 100,000 tokens, which make up around 90% of requests to its previous Haiku model. That statistic anchors the low price to the actual usage pattern of the model family. Developers are not paying a headline rate that applies only to toy examples; the $0.10 input price covers the majority of production traffic Anthropic already sees on Haiku.

What is adaptive thinking and why does it matter for agents?

Adaptive thinking is Anthropic's mechanism for letting a model decide how much internal reasoning to perform before answering. Effort controls expose that mechanism to developers, who can set a lower effort level for simple lookups and a higher level for multi-step tasks. The feature first appeared in larger Claude models; Haiku 5.5 is the first Haiku to include it. For agent developers, the practical consequence is that a single subagent can serve both fast-path and slow-path requests without being swapped for a different model. That flexibility reduces orchestration complexity and lets teams tune cost per agent turn rather than choosing between two separate models.

What benchmarks show the performance gains?

Anthropic reported major gains over Claude Haiku 4.5 on agentic benchmarks. On OSWorld 2.1, a computer-use evaluation, Haiku 5.5 scored 72.4% versus 15.7% for the prior model. On Terminal-Bench 4.0, an agentic coding benchmark, it scored 39.2% versus 0.0%. Anthropic also said Haiku 5.5 showed gains on knowledge work benchmarks. The published numbers are large enough to change procurement decisions, because they suggest that small models can now handle computer-use and coding tasks that were previously out of reach.

MetricClaude Haiku 5.5Claude Haiku 4.5Reported gain
OSWorld 2.1 computer use72.4%15.7%+56.7 points
Terminal-Bench 4.0 agentic coding39.2%0.0%+39.2 points
Average operating cost vs. Haiku 4.5About 75% lowerBaselineReported by AWS

The OSWorld result is particularly significant for the agent market because computer use has been a weak spot for small models. A score of 72.4% means Haiku 5.5 can automate desktop and browser workflows that earlier small models could not complete reliably. Terminal-Bench's jump from 0.0% to 39.2% reflects the addition of agentic coding capabilities that were effectively absent from Haiku 4.5. For teams building AI teammates, these are the benchmarks that separate a demo from a production workload.

Why is small-model pricing central to the agentic AI market?

Agentic workloads change the economics of inference because they multiply the number of calls per task. A single agent run can involve planning, tool selection, code execution, and verification, each requiring one or more model invocations. When those operations scale across thousands of concurrent sessions, a fraction of a cent per call becomes a meaningful line item. Small-model pricing therefore determines whether autonomous workflows are viable at high volume. Anthropic's $0.10 input price undercuts its own prior Haiku tier and puts direct pressure on Google and Mistral, both of which have emphasized cheaper models for agent pipelines.

Haiku 5.5's pricing also changes the calculus for developers who previously chose third-party small models to avoid Anthropic's premium tiers. At $0.10 per million input tokens, the model is priced for tasks that run constantly in the background, such as log analysis, document classification, and UI automation. Anthropic said prompts up to 100,000 tokens account for roughly 90% of requests to its previous Haiku model, which suggests the low price applies to the vast majority of real-world usage. The 50% lower price for longer prompts adds another layer of flexibility for retrieval-heavy workloads.

We’re very impressed with Claude Haiku 5.5, particularly its speed. We ran it through our eval suite for AI Teammates, our AI agent product, covering use cases like triaging bugs, setting up projects, and searching large portfolios to surface high-risk or overdue work. Compared with the model we use today, we saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn. It’s a noticeably snappier experience.

We’re very impressed with Claude Haiku 5.5, particularly its speed. We ran it through our eval suite for AI Teammates, our AI agent product, covering use cases like triaging bugs, setting up projects, and searching large portfolios to surface high-risk or overdue work. Compared with the model we use today, we saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn. It’s a noticeably snappier experience.

Where is Claude Haiku 5.5 available?

Claude Haiku 5.5 is available on the Claude Platform, Amazon Bedrock, Google Cloud, and Microsoft Foundry. AWS said the model is available on Amazon Bedrock and Claude Platform on AWS. The multi-cloud distribution is consistent with Anthropic's strategy of selling through the major infrastructure providers rather than requiring customers to move workloads to a proprietary cloud. For teams already using Bedrock's agent tooling, the integration means Haiku 5.5 can be swapped into existing pipelines without changing orchestration code.

The availability on Microsoft Foundry and Google Cloud gives Anthropic a presence in the two largest enterprise cloud ecosystems alongside AWS. That reach matters for the subagent use case, because the model is designed to be called repeatedly from within a larger application. Developers can route high-frequency subtasks to Haiku 5.5 while reserving Opus 5.5 and Sonnet 5.5 for planning and final reasoning. Cloud providers also benefit: a cheaper, faster Haiku makes agent products built on their platforms more economical at scale.

What does the price cut mean for enterprises and competitors?

For enterprises, the practical effect is a lower ceiling on per-agent operating cost. An application that generates 1 million input tokens and 100,000 output tokens per day would spend about $150 per day at the new list price for the prompt range up to 100,000 tokens, based on Anthropic's published rates. The 75% average reduction cited by AWS does not apply uniformly to every workload, but it gives finance teams a concrete number when modeling agent rollout costs. Lower inference cost also changes the break-even point for automating tasks that were previously too marginal to justify a model call.

For Anthropic, the pricing strategy trades margin on individual calls for volume and ecosystem lock-in. Haiku is the entry point that brings developers into the Claude API, and a cheaper Haiku makes it easier for startups to build agent products without passing high inference costs to customers. The same logic has driven Google and Mistral to compress prices on their smallest models, creating a market in which the default assumption is that agent subtasks run on inexpensive models. AWS's launch post emphasized the model's fit for subagents and high-volume, cost-sensitive work, signaling that the next phase of competition will be about total cost across a full agent workflow.

What are the limitations and what comes next?

The new model does not replace the reasoning-heavy work handled by Opus 5.5 and Sonnet 5.5. Anthropic positions Haiku 5.5 as the fast, cheap component of a multi-model agent architecture, not as a general-purpose flagship. Teams that need deep planning, complex code generation, or long-horizon autonomy will still route those calls to larger models. The pricing boundary at 100,000 tokens also means developers with very long prompts should verify which tier applies to their traffic, even though Anthropic prices longer prompts 50% lower.

Shipping adaptive thinking and a 1 million-token context window in the cheapest model suggests Anthropic is treating capability parity as a family-wide requirement rather than a flagship feature. The company has released Claude Opus 5.5, Claude Sonnet 5.5, and now Claude Haiku 5.5 within the same model generation, and each subsequent release has carried over capabilities that once distinguished the top tier. That pattern points to a roadmap in which model tiers differ mainly by scale, speed, and price, not by the presence or absence of core reasoning features. The most immediate competitive effect is on the price war: rivals now have to match a $0.10-per-million-token input price while also approaching the computer-use and agentic coding scores Anthropic reported.

  1. Whether Google and Mistral announce matching per-token prices for their smallest enterprise models.
  2. Whether independent benchmark runs confirm the OSWorld 2.1 and Terminal-Bench 4.0 scores.
  3. Whether enterprises shift default agent subagents to Haiku 5.5 on Bedrock, Google Cloud, or Foundry.
  4. Whether Anthropic extends the lower pricing for prompts over 100,000 tokens to the rest of the Claude 5.5 family.

Frequently asked

What is Claude Haiku 5.5?

Claude Haiku 5.5 is Anthropic's fastest, cheapest, and most capable small model, released Oct. 7, 2026. It is priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens and includes a 1 million-token context window.

How much does Claude Haiku 5.5 cost compared with Claude Haiku 4.5?

Anthropic and Amazon Web Services say Claude Haiku 5.5 costs around 75% less than Claude Haiku 4.5 on average. The list price for prompts up to 100,000 tokens is $0.10 per million input tokens and $0.50 per million output tokens, with prompts longer than 100,000 tokens priced 50% lower.

Where is Claude Haiku 5.5 available?

Claude Haiku 5.5 is available on the Claude Platform, Amazon Bedrock, Google Cloud, and Microsoft Foundry. AWS said the model is available on Amazon Bedrock and Claude Platform on AWS.

What benchmarks show Claude Haiku 5.5's gains?

Anthropic reported an OSWorld 2.1 computer-use score of 72.4% versus 15.7% for Claude Haiku 4.5, and a Terminal-Bench 4.0 agentic coding score of 39.2% versus 0.0% for the prior model.

Sources

  1. Anthropic — Claude Haiku 5.5 pricing, benchmark scores, context window, adaptive thinking, availability, and Anthropic's statement that it is the cheapest, fastest, and most capable small model Anthropic has released; also includes Aaron Vinh's evaluation comments.
  2. Amazon Web Services — Availability of Claude Haiku 5.5 on Amazon Bedrock and Claude Platform on AWS; AWS characterization of the model as the fastest and most efficient in the Claude 5.5 family; 75% average cost reduction versus Claude Haiku 4.5.
  3. CNET — Anthropic launched Claude Haiku 5.5, its fastest and cheapest model to date at $0.10 per million input tokens and $0.50 per million output (under 100k tokens). It targets high-volume tasks with major gains in computer…