Haiku 5.5 vs GPT-6 Luna: Same Token Price, 3x Different Cost
Since October 7, 2026, Claude Haiku 5.5 and GPT-6 Luna cost exactly the same on the price sheet: $0.10 per million input tokens and $0.50 output. Identical price, identical category, identical audience. And yet running the same task on one versus the other can cost three times more. This analysis explains why — and when each one earns its keep.
Quick answer: which is cheaper?
GPT-6 Luna, on most tasks. Despite identical per-token pricing, independent analysis from Artificial Analysis measures an average cost of $0.21 per task on Haiku 5.5 against $0.07 on Luna, because Haiku generates far more reasoning and output tokens. In exchange, Haiku delivers more quality: 43 against 38 on the same analysis’s intelligence index, and a wide margin on agent benchmarks.
Head to head
| Criterion | Claude Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Input (per million) | $0.10 | $0.10 |
| Output (per million) | $0.50 | $0.50 |
| Average cost per task | $0.21 | $0.07 |
| Intelligence index | 43 | 38 |
| OSWorld | 72.4% | 48.9% |
| Terminal-Bench 4.0 | 39.2% | 16.4% |
| GDPval-AA v2.1 | 1620 | 1437 |
| Context window | 1 million tokens | — |
Why the per-token price misleads
Two things happen at once on Haiku 5.5. The first is the new tokenizer, which splits the same text into more pieces — so the same prompt already enters costing more. The second, and heavier, is reasoning volume: the model generates far more intermediate tokens before answering.
That second point explains nearly the whole gap. You do not pay for the text you read; you pay for everything the model wrote to get there. A model that “thinks” more gets more right and costs more — and no price sheet shows that.
When each one wins
Pick GPT-6 Luna when the task is short, repetitive and has an objective answer: classify a message, extract a field from a document, translate a passage, tag sentiment, moderate a comment. In those cases Haiku’s extra reasoning does not improve the result — it only grows the bill.
Pick Claude Haiku 5.5 when the task involves chained steps or tool use: driving a browser, running terminal commands, following a flow with decisions. The 72.4% on OSWorld against 48.9%, and 39.2% on Terminal-Bench against 16.4%, are not fine-tuning differences — they are the difference between the task completing or not. Paying 3x for a run that works is cheaper than paying 1x for one that fails and needs human rework.
How to test for your case
- Take 100 real requests from your product, not synthetic examples.
- Run the same 100 through both models with the same prompt.
- Measure three numbers: total cost, success rate and latency.
- Compute cost per success, not per call — it is the only metric that matters.
- Repeat with caching and batching on; both models move significantly with them.
Verdict
There is no absolute winner here, which is rare in a comparison. Luna is the default choice for economics at volume; Haiku is the choice when the task has steps and failure is expensive. The common — and costly — mistake is choosing from the price sheet, which in this case is identical and tells you nothing.
Frequently asked questions
Which model is cheaper per task?
GPT-6 Luna, at an average $0.07 against $0.21 for Haiku 5.5, according to Artificial Analysis — despite identical per-token pricing.
Which performs better?
Haiku 5.5, at 43 against 38 on the intelligence index and a wide lead on OSWorld (72.4% vs 48.9%) and Terminal-Bench (39.2% vs 16.4%).
Why does cost per task differ so much?
Because Haiku 5.5 generates more reasoning tokens before answering, and the new tokenizer splits text into more pieces. You pay for all of it.
Can I use both?
Yes, and it is usually the best architecture: Luna for simple high-volume tasks and Haiku for anything involving tools or multiple steps.
At DigitalRadar, we compare on real cost, not the price sheet. Stay on the radar.
Reported from the sources cited and verified before publishing. See our Editorial Policy.