Claude Haiku 5.5 Cuts API Prices by up to 90% — With a Catch
Anthropic launched Claude Haiku 5.5 on October 7, 2026 with the most aggressive price cut in its history: $0.10 per million input tokens and $0.50 output for prompts up to 100K tokens. Against Haiku 4.5 ($1 and $5), that is a 90% drop. The number puts the model at exactly the same level as OpenAI’s GPT-6 Luna — and that is where the math gets interesting.
Quick answer: what does Haiku 5.5 cost?
$0.10 per million input tokens and $0.50 output on prompts up to 100K tokens; above that, $0.50 and $2.50. It keeps a 1M-token context window and up to 128K output tokens. Anthropic itself, however, estimates real savings at around 75%, not 90% — because the new tokenizer consumes more tokens per task.
The full pricing table
| Item | Up to 100K tokens | Above 100K |
|---|---|---|
| Input | $0.10 / million | $0.50 / million |
| Output | $0.50 / million | $2.50 / million |
| Cache reads | $0.01 / million | $0.05 / million |
| Cache writes | $0.125 / million | $0.625 / million |
Batch processing takes another 50% off. For anyone running classification, extraction or moderation at volume, combining cache and batch brings the cost to a tenth of what it was in 2025.
The benchmarks
On Anthropic’s published numbers, the generational jump is large:
| Benchmark | Haiku 5.5 | GPT-6 Luna | Haiku 4.5 |
|---|---|---|---|
| OSWorld | 72.4% | 48.9% | 15.7% |
| Terminal-Bench 4.0 | 39.2% | 16.4% | 0.0% |
| GDPval-AA v2.1 | 1620 | 1437 | 735 |
The tokenizer catch
Here is the part that changes the buying decision. Haiku 5.5 uses a new tokenizer that splits the same text into more pieces. Lower price per token, higher token count — and the two partly cancel out.
Worse: independent data from Artificial Analysis shows the model generates far more reasoning and output tokens per task. The result is an average cost per task of $0.21, against $0.07 for GPT-6 Luna — three times more expensive at identical per-token pricing. In exchange, Haiku scores 43 on the same analysis’s intelligence index, against 38 for Luna.
Why this matters
The practical lesson applies to any AI API: price per token is not price per task. Before migrating, run your own use case through both models and compare the total cost of 100 real requests, not the price sheet. If your task is short and specific (classify, extract a field, translate), Luna tends to come out cheaper; if it needs multi-step reasoning, Haiku delivers more accuracy for the extra cost.
There is a second-order effect worth watching. When the cheapest tier of a frontier lab drops 90% in a single release, the floor moves for everyone: products that were uneconomical at $1 per million tokens — per-message moderation, inline translation of every comment, summarizing an entire support archive nightly — become ordinary line items. The constraint stops being cost and becomes latency and accuracy.
And the 1-million-token context window matters more than the headline price for one specific pattern: feeding a whole codebase, contract or document set in a single call instead of engineering retrieval around it. At $0.10 per million input tokens, brute force is now cheaper than cleverness for a surprising number of jobs.
Frequently asked questions
How much does Claude Haiku 5.5 cost?
$0.10 per million input tokens and $0.50 output for prompts up to 100K tokens. Above that, $0.50 and $2.50.
Is the saving really 90%?
On price per token, yes. In practice Anthropic itself estimates around 75%, because the new tokenizer consumes more tokens per task.
Is it cheaper than GPT-6 Luna?
Per token, the price is identical. Per task, independent analysis puts Haiku at $0.21 against $0.07 for Luna, because Haiku generates more reasoning tokens.
How large is the context window?
1 million tokens of context, with up to 128K tokens of output.
At DigitalRadar, we read the whole price sheet — footnotes included. Stay on the radar.