Claude Haiku 5.5 Benchmarks: Coding, Agents & Speed

Does Claude Haiku 5.5 turn better benchmarks into cheaper results? Compare coding scores, agent tests, speed evidence, and cost per accepted task.

Claude Haiku 5.5 Benchmarks: Coding, Agents & Speed

Claude Haiku 5.5 leads Haiku 4.5 and GPT-6 Luna on shared benchmarks in Anthropic’s published comparison; Sonnet 5.5 remains ahead. Coding and agent results are promising, but workload-level speed and cost still need testing.

Cheap tokens do not guarantee cheap completed tasks. Our research found equal invoice accuracy at lower reported cost, while cheaper research-agent outputs missed the evaluator’s quality threshold. Retries and corrections can erase API cost savings.

Get more from your API budget with LinkModel’s pay-as-you-go pricing and savings of up to 30% versus official API rates. Its pricing offers vary by model, so compare API rates and choose a lower-cost option for your workload.

linkmodel.png

Claude Haiku 5.5 Benchmark Results at a Glance

Reviewed October 9, 2026. Our research examines official evaluations, published task tests, and public user questions. We did not independently run these API tests. External results retain their original configurations and limitations.

The Anthropic benchmark comparison supplies the published scores below. These are reported evaluation results, rather than independently reproduced measurements.

Coding and Computer-Use Scores

BenchmarkHaiku 5.5GPT-6 Luna
Terminal-Bench 4.039.2%16.4%
FrontierCode 1.1, Main46.4%42.4%
OSWorld 2.1, offline subset72.4%48.9%
Chartography, no tools46.4%29.1%
BenchmarkHaiku 4.5Sonnet 5.5
Terminal-Bench 4.00.0%70.6%
FrontierCode 1.1, MainNot reported52.1%, Xhigh
OSWorld 2.1, offline subset15.7%83.9%
Chartography, no tools6.4%61.6%

Keep the benchmark versions and conditions attached to the numbers. A missing result is not a zero, and a zero on one evaluation does not establish that a model cannot perform any related task.

For an existing Haiku application, the comparison justifies testing an upgrade. For work already handled by Sonnet, require evidence that Haiku can meet the same task requirements before changing the route.

Knowledge-Work and Reasoning Scores

ModelGDPval-AA v2.1AA-Briefcase v1.1
Haiku 5.516201578
Haiku 4.5735614
GPT-6 Luna14371336
Sonnet 5.518401824
ModelHLE, no toolsHLE, with tools
Haiku 5.545.9%57.4%
Haiku 4.510.2%18.7%
GPT-6 LunaNot reportedNot reported
Sonnet 5.556.9%64.5%

Do not average these figures into a single accuracy score. The knowledge-work values are not percentages, while the percentage-based evaluations cover different task sets.

Tool-assisted and unassisted reasoning also represent different conditions. Their results help identify what to investigate; they do not establish one model’s accuracy across all professional work.

Humanity’s Last Exam scores without and with tools: Haiku 5.5 45.9% and 57.4%, Haiku 4.5 10.2% and 18.7%, Sonnet 5.5 56.9% and 64.5%.

Claude Haiku 5.5 Coding Benchmarks: From Scores to Working Changes

Coding agents must investigate a problem, implement a change, verify it, and recover when an assumption fails. Producing plausible code is only part of that process.

The published results make Haiku worth evaluating for bounded development assignments. Longer engineering workflows need separate checks because mistakes can accumulate across tool calls.

Coding benchmark scores for Haiku 5.5, GPT-6 Luna, and Sonnet 5.5 on Terminal-Bench 4.0 and FrontierCode 1.1 Main.

Focused Coding Tasks Are Easier to Validate

Begin with assignments whose completion you can inspect directly.

AssignmentAcceptance check
Fix a contained defectReproduce the failure and verify the fix
Update a testConfirm it catches the intended behavior
Investigate repository behaviorCheck the supporting files and findings
Make a small implementation changeVerify requirements and regressions
Summarize code for another agentConfirm accuracy and missing context

Judge the completed change, including failed attempts and reviewer repair. A quick patch that needs substantial correction may offer less value than a slower result that passes.

For broad refactoring or an unfamiliar production defect, compare Haiku with a stronger coding model under the same requirements. Track whether each configuration maintains progress and recognizes when its approach is failing.

Creative-Coding Case: Full-Session Expense Matters

Our research included a voxel-pagoda coding comparison reporting API-equivalent expense.

ModelReported effortEstimated API-equivalent cost
Haiku 5.5Xhigh$24.96
GPT-6 LunaXhigh$1.96

These are estimated API-equivalent costs, not verified paid bills. The available record did not establish complete prompts, matched reasoning budgets, repeat counts, or a controlled quality comparison.

The figures cannot support a general claim that Haiku is more expensive. They identify a practical measurement requirement: record the whole coding session, including accumulated input, generated output, retries, and whether the task finished.

An identical effort label across providers does not establish an identical compute budget. Compare observed usage and accepted results instead.

Claude Haiku 5.5 Agent Benchmarks: Research Quality and Reliable Handoffs

Computer-use evaluations assess performance in a specified environment. A production agent must also preserve the goal, interpret tool results, collect sufficient evidence, and recognize incomplete work.

For research agents, successful tool execution does not establish a useful report. The final output must answer the brief with accurate, relevant support.

Four-Brief Research-Agent Comparison

Our review identified a research-agent evaluation covering brand visuals, forum quotations, company background, and advertising examples.

The evaluator compared Haiku with DeepSeek, specifically Haiku 5.5 and DeepSeek V4.1 Flash, using the same search and retrieval tools, including Exa, Brave, webpage retrieval, and image-related tools.

Haiku effortWins / ties / losses against DeepSeek
Low1 / 0 / 3
Medium1 / 1 / 2
High0 / 1 / 3
ConfigurationReported timeReported cost
Haiku low480 seconds$0.09
Haiku medium555 seconds$0.12
Haiku high992 seconds$0.31
DeepSeek low721 seconds$0.35

The evaluator retained DeepSeek for the workflow. Lower reported Haiku costs did not compensate for the assessed quality differences.

Reported research-agent time and cost for Haiku low, medium, and high and DeepSeek low, with time and USD shown in separate panels.

This was a small, self-reported comparison. Our research record contains available summaries rather than the complete experimental record. The full grading rubric was unavailable, and the aggregation of time and cost was insufficiently clear to calculate per-task averages.

The case supports investigating quality and cost together. It does not establish a general ranking between the two models.

What to Check in a Research Subagent

A research subagent should be evaluated on:

- Coverage of the actual brief.

- Accuracy of quotations and factual details.

- Suitability of sources for important claims.

- Traceability between evidence and conclusions.

- Disclosure of missing or unresolved information.

Access to the same tools does not guarantee the same research quality. Models may select different searches, read different pages, or stop before collecting enough evidence.

Use specific deliverables wherever possible. A request to find a financial figure and its supporting passage is easier to validate than an open-ended instruction to research an entire company.

Separate Subagent Performance from System Performance

A multi-model workflow may combine a stronger planning model, smaller execution models, tools, and recovery logic. Its overall score reflects that system.

Evaluate Haiku’s assigned contribution separately. Check whether its handoff preserves failed searches, missing evidence, and uncertainty. Otherwise, incomplete findings can become confident downstream conclusions.

A combined system result should not be presented as a standalone Haiku score.

Claude Haiku 5.5 Speed: What Is Known and What Remains Unmeasured

The reviewed evidence supports evaluating Haiku for latency-sensitive work, but it does not establish a universal completion-time advantage. Anthropic’s standard-speed positioning includes an exception for Opus models in Fast Mode. Application-level results still depend on effort, routing, tools, output length, and retries.

Measure Speed at the Right Stage

MeasurementWhat it answers
Time to first tokenHow soon does the response begin?
Output generation speedHow quickly does text arrive?
End-to-end completion timeHow long does the workflow take?
Tail latencyHow slow are less predictable requests?
Time to accepted resultHow long after retries and correction?

First-token latency matters for conversation. Successful throughput matters for batch processing. A coding agent needs to finish and verify its change.

A fast incomplete response is not a fast completed task. Measure acceptance alongside latency so failures cannot make a configuration appear artificially efficient.

Published Speed Observations Have Limited Scope

In the DataCamp invoice evaluation, the author reported faster completion and less than a quarter of Haiku 4.5’s output tokens. The report did not supply a standardized latency distribution.

Our research also reviewed a creative-output discussion describing an approximately 40-minute wait at maximum effort. A token count in that discussion was questioned as potentially anomalous, and the output images were not verified in the research record.

That observation is a reason to investigate long-running behavior, rather than a stable measurement of model speed. It should not be combined with the invoice test to calculate general throughput.

The important remaining gap is a matched-route comparison of first-token latency, generation speed, and tail latency under representative workloads.

Claude Haiku 5.5 Invoice Benchmark: Equal Accuracy, Different Costs

The invoice evaluation provides a concrete example of matching task decisions at different reported costs.

Each model received 24 invoices with purchase orders and delivery notes, then selected pay, hold, or return according to a fixed policy.

Published Invoice Results and Configurations

The author’s test report records the following results. Both models reached the sample’s accuracy ceiling.

ModelCorrect decisionsReported full-run cost
Haiku 5.524/24$0.0098
Haiku 4.524/24$0.25

Haiku 5.5 used high effort. Haiku 4.5 used a 16,000-token thinking budget, so the reasoning configurations differed.

The cases included price-tolerance boundaries, split delivery lines, mixed date formats, incorrect order numbers, and vendor notes arguing for payment.

In one example, 80 boxes were billed but only 68 had arrived. Both models correctly selected hold despite the request for full payment.

Both Haiku 5.5 and Haiku 4.5 made 24 of 24 correct invoice decisions, with reported full-run costs of $0.0098 and $0.25 under different reasoning configurations.

Why Perfect Scores Do Not Establish Equal Capability

The test distinguishes reported expense, not accuracy improvement. Haiku 4.5 already answered every case correctly.

It also does not establish equal capability ceilings. Harder cases or repeated runs could reveal differences that this sample did not expose.

For deployment, extend testing to incomplete paperwork, ambiguous policies, conflicting records, and consequential errors. Keep incorrect financial decisions separate from harmless presentation mistakes.

The invoice and research-agent cases produce different practical conclusions because their acceptance criteria differ. A model can be economical for fixed-policy decisions and still unsuitable for a source-sensitive research workflow.

Claude Haiku 5.5 Effort Settings: When Additional Reasoning Helps

Higher effort should earn its place through better accepted results. Extra waiting and usage are worthwhile when they resolve failures that matter.

In the four-brief research comparison, the high-effort configuration recorded no wins and one tie, alongside greater reported time and cost than the lower settings. That observation is task-specific, but it rejects the assumption that the highest setting must deliver the best practical outcome.

Haiku 5.5 outcomes against DeepSeek across four research briefs: low 1 win, 0 ties, 3 losses; medium 1, 1, 2; high 0, 1, 3.

Compare Effort Under Consistent Conditions

Keep the inputs, prompt, tools, and grading rules steady. Record acceptance, error type, completion time, usage, retries, and reviewer intervention.

Additional reasoning may help with an ambiguous decision. It may add little to a classification task that already passes reliably.

A longer explanation is not proof of improvement. Check whether the configuration fixes the actual failure.

Escalate Persistent Failures

Repeated errors may originate in unclear instructions, weak retrieval, or an assignment that is too broad.

Before increasing effort again, inspect the failure. Improve the evidence, clarify the rule, or divide the task into checkable parts. If the result still fails, route it to a stronger model or human review.

This approach makes effort a measured operational choice rather than a proxy for quality.

Claude Haiku 5.5 Benchmark Value: Cost per Accepted Task

Published input/output rates are $0.10/$0.50 per million tokens for prompts up to 100,000 tokens, and $0.50/$2.50 above that threshold. The threshold is a pricing boundary, not context capacity.

Haiku 5.5 input and output prices per million tokens: $0.10 and $0.50 for prompts up to 100K tokens, versus $0.50 and $2.50 above 100K.

Why Token Rates Do Not Determine Workflow Economics

The reviewed cases describe different outcomes:

- Invoice processing reported lower expense at equal measured decisions.

- Research-agent configurations were cheaper but did not satisfy the evaluator.

- Creative coding raised questions about accumulated session expense.

These results cannot be averaged. Their tasks, configurations, quality standards, and cost measurements differ.

Their shared lesson is to measure the expense required to obtain usable work.

Calculate Cost per Accepted Task

Cost per accepted task = total workflow expense ÷ accepted results.

Include unsuccessful attempts in the expense. Where relevant, record paid tools, retries, escalation, and human correction alongside model usage.

For long sessions, track growing inputs and repeated processing. For cached workflows, compare observed cache behavior through the same route rather than assuming savings from a headline rate.

This metric prevents inexpensive failures from appearing more economical than accepted results.

Claude Haiku 5.5 Third-Party Rankings: Estimated Scores and Evidence

Our research recorded an October 7, 2026 snapshot of the BenchLM model page. The composite result was marked Estimated.

Recorded fieldSnapshot value
Composite score66.32/100
Overall position28/216
Agentic position18/122
Coding position20/146

These values describe the dated research record. They are not a claim that the live ranking remains unchanged.

Use Composite Rankings to Shortlist Candidates

A composite score is not a measured percentage of real tasks completed correctly. Category positions also use different comparison pools.

The research record identified inconsistent evidence counts: five visible records in one area and seventeen referenced elsewhere. It did not provide enough information to reconstruct the composite score.

Use an estimated ranking as a screening tool, then inspect the underlying evidence. Missing categories, provider-reported scores, and independently reproduced tests require different interpretations.

Evaluating Claude Haiku 5.5 Before Migration

Choose an initial workload with representative inputs, explicit success criteria, and an output you can inspect.

WorkloadPrimary acceptance check
Fixed-rule document decisionsCorrect application of policy
Focused codingWorking change and meaningful verification
Research subagentsSource support and adequate coverage
Summaries and extractionFidelity to supplied material
Interactive assistanceCorrect completion within latency limits
Long agent sessionsRecovery, completion, and total expense

Separate Workflow Ideas from Validated Results

Our research also found reports of document-to-JSON extraction, conversational collection of custom-order requirements, prompt-evaluation pipelines, and proposed image-recognition trials.

Those scenarios lacked confirmed Haiku 5.5 measurements or complete version information. They identify possible evaluation targets rather than validated outcomes for this model.

This distinction keeps useful ideas in the comparison without turning them into unsupported performance claims.

Keep a Defensible Evaluation Record

For each case, preserve the model, provider route, date, effort, prompt version, usage, tool expense, latency, retries, and acceptance decision.

Repeat important cases and retain failures. Compare against the current workflow, including its review requirements.

Migrate when the candidate consistently meets the required standard at acceptable cost and latency. Keep an escalation path for cases that fail.

Conclusion

Claude Haiku 5.5’s published benchmarks make it a credible candidate for focused coding, computer-use workflows, and high-volume document tasks. The cases explain how to turn that promise into a decision: invoice processing matched decisions at lower reported cost, research agents showed that cheaper output can miss the quality threshold, and long coding sessions highlighted accumulated expense. Choose the configuration that reliably delivers accepted work within your time and cost limits. Representative inputs, consistent comparisons, and complete records provide a stronger basis for that choice than a headline score alone.

Frequently Asked Questions

Does Haiku 5.5 Beat GPT-6 Luna or Replace Sonnet?

Haiku leads Luna on their shared published comparison rows, while Sonnet remains ahead. These results support workload-specific evaluation. They do not establish that Haiku can replace either model across every application.

Is Haiku 5.5 Fast, or Does It Overthink and Produce Verbose Answers?

Our review of user questions identified concerns about overthinking and verbosity, but it did not establish general rates for either behavior. Compare effort settings by accepted results and full completion time rather than response onset alone.

Do the Benchmarks Establish Low Hallucination or Tool-Call Error Rates?

No. The reviewed scores do not supply general hallucination or tool-error rates for your application. Test factual support, uncertainty handling, invalid actions, recovery, and subagent handoffs directly.

Will Haiku 5.5 Reduce Long-Agent Costs After Caching and Tokenizer Changes?

Possibly, but the evidence does not establish a universal saving. Request size, tokenization, cache behavior, effort, and retries affect expense. Compare observed usage and accepted results through the same route; our research did not establish a uniform cached-cost advantage over competing models.

Should I Upgrade from Haiku 4.5 or Add Haiku Subagents?

Both are reasonable options to evaluate. The invoice case supports investigating fixed-rule work, while the research comparison shows why subagent quality needs separate checks. Confirm compatibility, acceptance rates, latency, and total expense before changing production traffic.

About the author

Fiona Thorne

Fiona Thorne

AI Model & API Researcher at LinkModel

Fiona Thorne is an AI model and API researcher at LinkModel, focusing on generative AI technologies, model capabilities, API pricing, and practical integration strategies. She explores developments across leading AI providers, drawing on official documentation, technical specifications, and comparative research to help developers and businesses evaluate AI solutions, understand their trade-offs, and make informed technology decisions.

Related Posts