Anthropic dropped its newest small model on Wednesday, and the pitch is simple: do 75 percent more for 75 percent less. Claude Haiku 5.5, released October 7, is the company’s cheapest, fastest small model to date. It undercuts both its predecessor and OpenAI’s budget rival on price while beating them on several benchmarks.
“Claude Haiku 5.5 is designed for high-volume, cost-sensitive tasks,” Anthropic said in the launch announcement. “It reliably handles quick and repetitive workloads like summaries, compactions, database queries, and classification requests. It pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work. And, since it’s also our fastest model to date, it works especially well for speed-sensitive tasks like live customer support and browser use.”
This is a straight-up price war move. And it’s aimed squarely at developers running the same prompt thousands of times a day.
What Claude Haiku 5.5 actually costs
The pricing is tiered by prompt length, which matters more than it sounds. For prompts up to 100,000 tokens (roughly 90 percent of requests to the previous Haiku model, according to Anthropic), input costs $0.10 per million tokens and output costs $0.50 per million. Cache reads are $0.01 per million, cache writes $0.125.
Cross that 100,000-token line and the meter jumps: input goes to $0.50 per million, output to $2.50, with cache reads at $0.05 and writes at $0.625. The headline number Anthropic puts on the launch page: Haiku 5.5 is priced 90 percent lower than Haiku 4.5 for requests under 100,000 tokens, 50 percent lower above that, and “on average, it now costs around 75 percent less to run.”
The $0.10 input figure isn’t random. It matches OpenAI’s GPT-6 Luna exactly, which makes this launch as much a competitive statement as a product update.
A first for Haiku: adjustable effort
Haiku 5.5 is the first small model from Anthropic with an adjustable effort setting. Developers can dial the trade-off between cost and intelligence directly, instead of having the model decide how hard to think.
There’s also a footnote on the “fastest model” claim that Anthropic is refreshingly upfront about: Haiku 5.5 is fastest at each model’s standard speed, but it still trails Opus-class models running in Fast Mode. Fast is relative. Context matters.
Anthropic ships the model under the ID claude-haiku-5-5, and it’s live now on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure. The model outputs text only and has a knowledge cutoff of June 2026.
The benchmark scores
Small models live or die on whether they can keep up with bigger ones on real tasks. Here’s what Anthropic published:
- OSWorld 2.1 (offline subset): 72.4 percent, up from 15.7 percent on Haiku 4.5. That’s not an incremental bump; it’s a different category
- Terminal-Bench 4.0: 39.2 percent, where the previous Haiku scored zero
- FrontierCode 1.1: 46.4 percent
- Humanity’s Last Exam: 45.9 percent without tools, 57.4 percent with tools
Independent comparisons put those numbers above OpenAI’s GPT-6 Luna at the same price point: 72.4 versus 48.9 percent on OSWorld 2.1, and 39.2 versus 16.4 on Terminal-Bench 4.0.
Early partner results add color. Cursor claims roughly 10x lower cost than Haiku 4.5 on shorter requests. GitHub Copilot’s team says Haiku 5.5 “matched Claude Sonnet 5 on many coding tasks while using fewer tokens and steps” inside VS Code. Devin reported 58.4 percent on FrontierCode 1.1, ahead of Sonnet 5 at roughly one-eighth the cost per task. And Asana’s staff software engineer Aaron Vinh said the company saw “over a 30 percent reduction in latency for task completions and up to 2.5x faster inference per agent turn.”
Two more changes came with the launch
Haiku 5.5 didn’t arrive alone. Anthropic also halved cache read pricing for Claude Sonnet 5.5, from $0.20 to $0.10 per million tokens, which the company says makes Sonnet about 20 percent cheaper on most long-running, multi-step agentic work. Cheaper context is a subsidy for agent architectures, and Anthropic knows it.
The third change: monthly API credits for paid subscribers. Max 5x plans get $100 per month, Max 20x gets $200, and Team subscribers get up to $500 pooled across users, usable on any model and in third-party harnesses. It’s Anthropic putting its own money behind developers building agents on its platform.
The Python and TypeScript SDKs got a same-day update too, with computer-use and browser-use toolsets built in. The SDK now runs the action loop and routes clicks and keystrokes to drivers from Browserbase, E2B, Daytona, or browser_use, so developers don’t have to hand-roll that plumbing.
Why this matters
The model race is quietly becoming a unit-economics race. Frontier models grab headlines, but the money is in the thousands of small, repeated jobs: classifying tickets, summarizing threads, compacting context, answering support chats — where a dime per million tokens changes the P&L.
Anthropic is betting that developers care more about what a model costs at scale than what it scores on the hardest evals. With Haiku 5.5 matching GPT-6 Luna’s price while beating it on the benchmarks developers actually run, OpenAI now has a real pricing problem at the low end.
There’s a second, subtler story here. Pairing a dirt-cheap fast model as a “subagent” under a big model like Opus 5.5 is becoming the standard architecture: the heavyweight thinks, the lightweight executes. Anthropic is now selling both halves of that stack, and pricing the lightweight half to own the pattern.
FAQ
What is Claude Haiku 5.5? Claude Haiku 5.5 is Anthropic’s newest small AI model, released October 7, 2026. It handles high-volume, cost-sensitive tasks like summaries, database queries, and classification, and works as a subagent under Opus 5.5 or Sonnet 5.5 for coding work.
How much does Claude Haiku 5.5 cost? For prompts up to 100,000 tokens: $0.10 per million input tokens, $0.50 per million output tokens. Above 100,000 tokens, rates quintuple. On average it costs about 75 percent less to run than Haiku 4.5.
Is Haiku 5.5 actually the fastest Claude? At standard speed, yes, per Anthropic’s own footnote. But it doesn’t beat Opus models in Fast Mode, and “fast” here also includes cheaper cache and latency gains, not just raw tokens per second.
Can developers use Claude Haiku 5.5 outside Anthropic’s platform? Yes. It’s available at launch on the Claude Platform, AWS, Google Cloud, and Microsoft Azure, plus day-one integrations with Cursor, GitHub Copilot, and Devin.
Sources: Anthropic (official launch announcement, @claudeai, @ClaudeDevs), 9to5Mac, Gadgets360, Unite.AI, TestingCatalog, Latent.Space, AIWeekly.
