Claude Haiku 5.5 leads Haiku 4.5 and GPT-6 Luna on shared benchmarks in Anthropic’s published comparison; Sonnet 5.5 remains ahead. Coding and agent results are promising, but workload-level speed and cost still need testing.
Cheap tokens do not guarantee cheap completed tasks. Our research found equal invoice accuracy at lower reported cost, while cheaper research-agent outputs missed the evaluator’s quality threshold. Retries and corrections can erase API cost savings.
Get more from your API budget with LinkModel’s pay-as-you-go pricing and savings of up to 30% versus official API rates. Its pricing offers vary by model, so compare API rates and choose a lower-cost option for your workload.

Claude Haiku 5.5 Benchmark Results at a Glance
Reviewed October 9, 2026. Our research examines official evaluations, published task tests, and public user questions. We did not independently run these API tests. External results retain their original configurations and limitations.
The Anthropic benchmark comparison supplies the published scores below. These are reported evaluation results, rather than independently reproduced measurements.
Coding and Computer-Use Scores
| Benchmark | Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Terminal-Bench 4.0 | 39.2% | 16.4% |
| FrontierCode 1.1, Main | 46.4% | 42.4% |
| OSWorld 2.1, offline subset | 72.4% | 48.9% |
| Chartography, no tools | 46.4% | 29.1% |
| Benchmark | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|
| Terminal-Bench 4.0 | 0.0% | 70.6% |
| FrontierCode 1.1, Main | Not reported | 52.1%, Xhigh |
| OSWorld 2.1, offline subset | 15.7% | 83.9% |
| Chartography, no tools | 6.4% | 61.6% |
Keep the benchmark versions and conditions attached to the numbers. A missing result is not a zero, and a zero on one evaluation does not establish that a model cannot perform any related task.
For an existing Haiku application, the comparison justifies testing an upgrade. For work already handled by Sonnet, require evidence that Haiku can meet the same task requirements before changing the route.
Knowledge-Work and Reasoning Scores
| Model | GDPval-AA v2.1 | AA-Briefcase v1.1 |
|---|---|---|
| Haiku 5.5 | 1620 | 1578 |
| Haiku 4.5 | 735 | 614 |
| GPT-6 Luna | 1437 | 1336 |
| Sonnet 5.5 | 1840 | 1824 |
| Model | HLE, no tools | HLE, with tools |
|---|---|---|
| Haiku 5.5 | 45.9% | 57.4% |
| Haiku 4.5 | 10.2% | 18.7% |
| GPT-6 Luna | Not reported | Not reported |
| Sonnet 5.5 | 56.9% | 64.5% |
Do not average these figures into a single accuracy score. The knowledge-work values are not percentages, while the percentage-based evaluations cover different task sets.
Tool-assisted and unassisted reasoning also represent different conditions. Their results help identify what to investigate; they do not establish one model’s accuracy across all professional work.

Claude Haiku 5.5 Coding Benchmarks: From Scores to Working Changes
Coding agents must investigate a problem, implement a change, verify it, and recover when an assumption fails. Producing plausible code is only part of that process.
The published results make Haiku worth evaluating for bounded development assignments. Longer engineering workflows need separate checks because mistakes can accumulate across tool calls.

Focused Coding Tasks Are Easier to Validate
Begin with assignments whose completion you can inspect directly.
| Assignment | Acceptance check |
|---|---|
| Fix a contained defect | Reproduce the failure and verify the fix |
| Update a test | Confirm it catches the intended behavior |
| Investigate repository behavior | Check the supporting files and findings |
| Make a small implementation change | Verify requirements and regressions |
| Summarize code for another agent | Confirm accuracy and missing context |
Judge the completed change, including failed attempts and reviewer repair. A quick patch that needs substantial correction may offer less value than a slower result that passes.
For broad refactoring or an unfamiliar production defect, compare Haiku with a stronger coding model under the same requirements. Track whether each configuration maintains progress and recognizes when its approach is failing.
Creative-Coding Case: Full-Session Expense Matters
Our research included a voxel-pagoda coding comparison reporting API-equivalent expense.
| Model | Reported effort | Estimated API-equivalent cost |
|---|---|---|
| Haiku 5.5 | Xhigh | $24.96 |
| GPT-6 Luna | Xhigh | $1.96 |
These are estimated API-equivalent costs, not verified paid bills. The available record did not establish complete prompts, matched reasoning budgets, repeat counts, or a controlled quality comparison.
The figures cannot support a general claim that Haiku is more expensive. They identify a practical measurement requirement: record the whole coding session, including accumulated input, generated output, retries, and whether the task finished.
An identical effort label across providers does not establish an identical compute budget. Compare observed usage and accepted results instead.
Claude Haiku 5.5 Agent Benchmarks: Research Quality and Reliable Handoffs
Computer-use evaluations assess performance in a specified environment. A production agent must also preserve the goal, interpret tool results, collect sufficient evidence, and recognize incomplete work.
For research agents, successful tool execution does not establish a useful report. The final output must answer the brief with accurate, relevant support.
Four-Brief Research-Agent Comparison
Our review identified a research-agent evaluation covering brand visuals, forum quotations, company background, and advertising examples.
The evaluator compared Haiku with DeepSeek, specifically Haiku 5.5 and DeepSeek V4.1 Flash, using the same search and retrieval tools, including Exa, Brave, webpage retrieval, and image-related tools.
| Haiku effort | Wins / ties / losses against DeepSeek |
|---|---|
| Low | 1 / 0 / 3 |
| Medium | 1 / 1 / 2 |
| High | 0 / 1 / 3 |
| Configuration | Reported time | Reported cost |
|---|---|---|
| Haiku low | 480 seconds | $0.09 |
| Haiku medium | 555 seconds | $0.12 |
| Haiku high | 992 seconds | $0.31 |
| DeepSeek low | 721 seconds | $0.35 |
The evaluator retained DeepSeek for the workflow. Lower reported Haiku costs did not compensate for the assessed quality differences.

This was a small, self-reported comparison. Our research record contains available summaries rather than the complete experimental record. The full grading rubric was unavailable, and the aggregation of time and cost was insufficiently clear to calculate per-task averages.
The case supports investigating quality and cost together. It does not establish a general ranking between the two models.
What to Check in a Research Subagent
A research subagent should be evaluated on:
- Coverage of the actual brief.
- Accuracy of quotations and factual details.
- Suitability of sources for important claims.
- Traceability between evidence and conclusions.
- Disclosure of missing or unresolved information.
Access to the same tools does not guarantee the same research quality. Models may select different searches, read different pages, or stop before collecting enough evidence.
Use specific deliverables wherever possible. A request to find a financial figure and its supporting passage is easier to validate than an open-ended instruction to research an entire company.
Separate Subagent Performance from System Performance
A multi-model workflow may combine a stronger planning model, smaller execution models, tools, and recovery logic. Its overall score reflects that system.
Evaluate Haiku’s assigned contribution separately. Check whether its handoff preserves failed searches, missing evidence, and uncertainty. Otherwise, incomplete findings can become confident downstream conclusions.
A combined system result should not be presented as a standalone Haiku score.
Claude Haiku 5.5 Speed: What Is Known and What Remains Unmeasured
The reviewed evidence supports evaluating Haiku for latency-sensitive work, but it does not establish a universal completion-time advantage. Anthropic’s standard-speed positioning includes an exception for Opus models in Fast Mode. Application-level results still depend on effort, routing, tools, output length, and retries.
Measure Speed at the Right Stage
| Measurement | What it answers |
|---|---|
| Time to first token | How soon does the response begin? |
| Output generation speed | How quickly does text arrive? |
| End-to-end completion time | How long does the workflow take? |
| Tail latency | How slow are less predictable requests? |
| Time to accepted result | How long after retries and correction? |
First-token latency matters for conversation. Successful throughput matters for batch processing. A coding agent needs to finish and verify its change.
A fast incomplete response is not a fast completed task. Measure acceptance alongside latency so failures cannot make a configuration appear artificially efficient.
Published Speed Observations Have Limited Scope
In the DataCamp invoice evaluation, the author reported faster completion and less than a quarter of Haiku 4.5’s output tokens. The report did not supply a standardized latency distribution.
Our research also reviewed a creative-output discussion describing an approximately 40-minute wait at maximum effort. A token count in that discussion was questioned as potentially anomalous, and the output images were not verified in the research record.
That observation is a reason to investigate long-running behavior, rather than a stable measurement of model speed. It should not be combined with the invoice test to calculate general throughput.
The important remaining gap is a matched-route comparison of first-token latency, generation speed, and tail latency under representative workloads.
Claude Haiku 5.5 Invoice Benchmark: Equal Accuracy, Different Costs
The invoice evaluation provides a concrete example of matching task decisions at different reported costs.
Each model received 24 invoices with purchase orders and delivery notes, then selected pay, hold, or return according to a fixed policy.
Published Invoice Results and Configurations
The author’s test report records the following results. Both models reached the sample’s accuracy ceiling.
| Model | Correct decisions | Reported full-run cost |
|---|---|---|
| Haiku 5.5 | 24/24 | $0.0098 |
| Haiku 4.5 | 24/24 | $0.25 |
Haiku 5.5 used high effort. Haiku 4.5 used a 16,000-token thinking budget, so the reasoning configurations differed.
The cases included price-tolerance boundaries, split delivery lines, mixed date formats, incorrect order numbers, and vendor notes arguing for payment.
In one example, 80 boxes were billed but only 68 had arrived. Both models correctly selected hold despite the request for full payment.

Why Perfect Scores Do Not Establish Equal Capability
The test distinguishes reported expense, not accuracy improvement. Haiku 4.5 already answered every case correctly.
It also does not establish equal capability ceilings. Harder cases or repeated runs could reveal differences that this sample did not expose.
For deployment, extend testing to incomplete paperwork, ambiguous policies, conflicting records, and consequential errors. Keep incorrect financial decisions separate from harmless presentation mistakes.
The invoice and research-agent cases produce different practical conclusions because their acceptance criteria differ. A model can be economical for fixed-policy decisions and still unsuitable for a source-sensitive research workflow.
Claude Haiku 5.5 Effort Settings: When Additional Reasoning Helps
Higher effort should earn its place through better accepted results. Extra waiting and usage are worthwhile when they resolve failures that matter.
In the four-brief research comparison, the high-effort configuration recorded no wins and one tie, alongside greater reported time and cost than the lower settings. That observation is task-specific, but it rejects the assumption that the highest setting must deliver the best practical outcome.

Compare Effort Under Consistent Conditions
Keep the inputs, prompt, tools, and grading rules steady. Record acceptance, error type, completion time, usage, retries, and reviewer intervention.
Additional reasoning may help with an ambiguous decision. It may add little to a classification task that already passes reliably.
A longer explanation is not proof of improvement. Check whether the configuration fixes the actual failure.
Escalate Persistent Failures
Repeated errors may originate in unclear instructions, weak retrieval, or an assignment that is too broad.
Before increasing effort again, inspect the failure. Improve the evidence, clarify the rule, or divide the task into checkable parts. If the result still fails, route it to a stronger model or human review.
This approach makes effort a measured operational choice rather than a proxy for quality.
Claude Haiku 5.5 Benchmark Value: Cost per Accepted Task
Published input/output rates are $0.10/$0.50 per million tokens for prompts up to 100,000 tokens, and $0.50/$2.50 above that threshold. The threshold is a pricing boundary, not context capacity.

Why Token Rates Do Not Determine Workflow Economics
The reviewed cases describe different outcomes:
- Invoice processing reported lower expense at equal measured decisions.
- Research-agent configurations were cheaper but did not satisfy the evaluator.
- Creative coding raised questions about accumulated session expense.
These results cannot be averaged. Their tasks, configurations, quality standards, and cost measurements differ.
Their shared lesson is to measure the expense required to obtain usable work.
Calculate Cost per Accepted Task
Cost per accepted task = total workflow expense ÷ accepted results.
Include unsuccessful attempts in the expense. Where relevant, record paid tools, retries, escalation, and human correction alongside model usage.
For long sessions, track growing inputs and repeated processing. For cached workflows, compare observed cache behavior through the same route rather than assuming savings from a headline rate.
This metric prevents inexpensive failures from appearing more economical than accepted results.
Claude Haiku 5.5 Third-Party Rankings: Estimated Scores and Evidence
Our research recorded an October 7, 2026 snapshot of the BenchLM model page. The composite result was marked Estimated.
| Recorded field | Snapshot value |
|---|---|
| Composite score | 66.32/100 |
| Overall position | 28/216 |
| Agentic position | 18/122 |
| Coding position | 20/146 |
These values describe the dated research record. They are not a claim that the live ranking remains unchanged.
Use Composite Rankings to Shortlist Candidates
A composite score is not a measured percentage of real tasks completed correctly. Category positions also use different comparison pools.
The research record identified inconsistent evidence counts: five visible records in one area and seventeen referenced elsewhere. It did not provide enough information to reconstruct the composite score.
Use an estimated ranking as a screening tool, then inspect the underlying evidence. Missing categories, provider-reported scores, and independently reproduced tests require different interpretations.
Evaluating Claude Haiku 5.5 Before Migration
Choose an initial workload with representative inputs, explicit success criteria, and an output you can inspect.
| Workload | Primary acceptance check |
|---|---|
| Fixed-rule document decisions | Correct application of policy |
| Focused coding | Working change and meaningful verification |
| Research subagents | Source support and adequate coverage |
| Summaries and extraction | Fidelity to supplied material |
| Interactive assistance | Correct completion within latency limits |
| Long agent sessions | Recovery, completion, and total expense |
Separate Workflow Ideas from Validated Results
Our research also found reports of document-to-JSON extraction, conversational collection of custom-order requirements, prompt-evaluation pipelines, and proposed image-recognition trials.
Those scenarios lacked confirmed Haiku 5.5 measurements or complete version information. They identify possible evaluation targets rather than validated outcomes for this model.
This distinction keeps useful ideas in the comparison without turning them into unsupported performance claims.
Keep a Defensible Evaluation Record
For each case, preserve the model, provider route, date, effort, prompt version, usage, tool expense, latency, retries, and acceptance decision.
Repeat important cases and retain failures. Compare against the current workflow, including its review requirements.
Migrate when the candidate consistently meets the required standard at acceptable cost and latency. Keep an escalation path for cases that fail.
Conclusion
Claude Haiku 5.5’s published benchmarks make it a credible candidate for focused coding, computer-use workflows, and high-volume document tasks. The cases explain how to turn that promise into a decision: invoice processing matched decisions at lower reported cost, research agents showed that cheaper output can miss the quality threshold, and long coding sessions highlighted accumulated expense. Choose the configuration that reliably delivers accepted work within your time and cost limits. Representative inputs, consistent comparisons, and complete records provide a stronger basis for that choice than a headline score alone.
Frequently Asked Questions
Does Haiku 5.5 Beat GPT-6 Luna or Replace Sonnet?
Haiku leads Luna on their shared published comparison rows, while Sonnet remains ahead. These results support workload-specific evaluation. They do not establish that Haiku can replace either model across every application.
Is Haiku 5.5 Fast, or Does It Overthink and Produce Verbose Answers?
Our review of user questions identified concerns about overthinking and verbosity, but it did not establish general rates for either behavior. Compare effort settings by accepted results and full completion time rather than response onset alone.
Do the Benchmarks Establish Low Hallucination or Tool-Call Error Rates?
No. The reviewed scores do not supply general hallucination or tool-error rates for your application. Test factual support, uncertainty handling, invalid actions, recovery, and subagent handoffs directly.
Will Haiku 5.5 Reduce Long-Agent Costs After Caching and Tokenizer Changes?
Possibly, but the evidence does not establish a universal saving. Request size, tokenization, cache behavior, effort, and retries affect expense. Compare observed usage and accepted results through the same route; our research did not establish a uniform cached-cost advantage over competing models.
Should I Upgrade from Haiku 4.5 or Add Haiku Subagents?
Both are reasonable options to evaluate. The invoice case supports investigating fixed-rule work, while the research comparison shows why subagent quality needs separate checks. Confirm compatibility, acceptance rates, latency, and total expense before changing production traffic.
