On October 7, Anthropic quietly launched Claude Haiku 5.5, with AWS's official blog announcing its arrival on Bedrock the same day and QbitAI publishing a detailed review the next morning. This "small cup" model simultaneously drove prices and benchmark positions to new lows: on raw scores it clears what were previously seen as the small-model kill lines — DeepSeek V4.1 Flash and GLM-5.3-Flash — becoming the new gatekeeper, while its official benchmarks outscore GPT-6 Luna item by item. Yahoo Finance's verdict was blunt: the AI pricing war is intensifying.

Pricing: down 90% in one year

Pricing is the headline. For prompts within 100K tokens — about 90% of all Haiku 4.5 requests — input and output prices are cut by 90%; the portion beyond 100K tokens still gets 50% off. Haiku 4.5, released a year ago, cost 10 times as much. After the cuts, Haiku 5.5 is now cheaper than DeepSeek V4.1 Flash on everything except cache-hit input.

Key numbers at a glance

· OSWorld 2.1 (real computer use): Low effort 42.0% accuracy at $0.07 per task; Max effort 72.4% at $0.61; Haiku 4.5 Max managed only 15.7% at $1.45

· GDPval-AA v2.1 (real work across 44 professions): Low Elo 1125 at $0.01; Max Elo 1620 at $0.87; the previous Max scored just 735

· Terminal-Bench 4.0: Haiku 5.5 scores 39.2% versus Sonnet 5.5's 70.6%

· Sonnet 5.5 cache reads drop from $0.20 to $0.10 per million tokens — Anthropic says most agent tasks get about 20% cheaper

Small cup stays small: it won't replace Sonnet

Beyond benchmarks, keep a cool head: Anthropic itself recommends Sonnet 5.5 or Opus 5.5 for complex agentic coding. The 39.2% vs 70.6% gap on Terminal-Bench 4.0 shows that multi-step coding, cross-file refactoring and long-horizon autonomous planning remain big-cup territory. Haiku 5.5 shines at the execution layer — tasks already decomposed, with clear acceptance criteria, that can run in parallel. Cognition's Devin pairs Opus 5.5 as the main model with Haiku 5.5 sub-agents, reaching 66.2% on FrontierCode — better than either model alone — a textbook case of orchestration beating a single point.

The tokenizer catch and five breaking changes

One caveat: Haiku 5.5 switches to the same tokenizer as Sonnet 5.5 and Opus 5.5, consuming more tokens for identical tasks, so real-world savings lag the sticker discount. The API also has five breaking changes: budget_tokens now errors out and must move to adaptive thinking plus effort parameters; temperature, top_p and top_k are all locked to defaults; assistant message pre-filling is gone; the computer-use tool moves from computer_20250124 to computer_toolset_20260801; and the first response block may be a thinking block, so parsers must filter by type. This is also the first Haiku to support five effort levels, from low to max.

The dividends of a price war go to engineering teams that invested early in evaluation, routing and migration abstractions — the "just swap the model name" era is over.

The Claude family is now a three-tier lineup: Opus 5.5 for hard reasoning, Sonnet 5.5 for general execution, Haiku 5.5 for volume and speed. Our Zenith-Act platform follows the same philosophy in model routing: schedule dynamically by task complexity and cost, and let small models absorb high-frequency execution work — that is how you genuinely push down the marginal cost of AI. (Compiled from QbitAI, Anthropic's official site, the AWS blog, Yahoo Finance and other public reports)