
Anthropic has introduced Claude Haiku 5.5, a smaller model aimed at the work that can quietly make an AI product expensive: thousands of summaries, classifications, support replies and background requests every day. It is the first Haiku model with an adjustable effort setting, so developers can ask it to spend less computing time on a straightforward job and more on one that needs deeper reasoning.
That choice matters when the same application handles very different requests. A short label for an incoming email does not need the same effort as a tricky customer query. Sending both to a powerful model at its highest setting can add cost and delay without improving the simpler answer. Haiku 5.5 gives teams another way to tune that trade-off, although they will still need to test whether lower effort preserves the accuracy their own tasks require.
For prompts of up to 100,000 tokens, Anthropic lists a price of $0.10 per million input tokens and $0.50 per million output tokens. Cache reads cost $0.01 per million tokens and cache writes $0.125. The rates rise for prompts longer than 100,000 tokens: $0.50 for input and $2.50 for output, with higher cache charges as well. That distinction is important for developers feeding long documents or large histories into the model. A low headline rate will not describe every production workload.
Anthropic estimates Haiku 5.5 costs about 75% less per task on average than Haiku 4.5. That is its own estimate across tasks, not a guarantee that every customer’s bill will fall by that amount. Token use, prompt length, caching and the effort setting all affect the final cost. The company also says the new model is faster and more capable than its previous Haiku release, but its published benchmark results should be read as vendor-reported tests rather than a substitute for real-world trials.
The model is available now through Anthropic’s developer platform and through AWS, Google Cloud and Microsoft Azure. That gives businesses already using those clouds a way to test it without moving their applications to a new provider. The more interesting question is not whether Haiku can replace a larger Claude model everywhere. It is whether a business can reserve the expensive model for difficult cases while a cheaper one deals reliably with the routine volume.
There is a related saving for developers using Sonnet. Anthropic has cut the cache-read rate for Sonnet 5.5 from $0.20 to $0.10 per million tokens. Caching lets an application reuse parts of a prompt rather than paying the full input price each time, which can matter in tools that repeatedly send the same instructions or documents. The change adds a new pricing dimension to Sonnet 5.5’s earlier speed improvements.
Small models are often discussed as if they are a compromise made only to save money. At scale, they can be the difference between a useful AI feature and one that is too costly to leave switched on. Haiku 5.5 gives developers a sharper set of controls, but the sensible deployment will still be measured; compare quality, latency and the actual cost of completed tasks before routing millions of requests through it.







