Claude Haiku 5.5 API Pricing: Input, Output, Cache & Batch Costs

Why can cheap Haiku tokens produce a costly bill? Compare Claude Haiku 5.5 API pricing, the 100K price jump, cache savings, Batch rates, and worked cost examples.

Claude Haiku 5.5 API Pricing: Input, Output, Cache & Batch Costs

Claude Haiku 5.5 API pricing starts at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Longer prompts cost $0.50 for input and $2.50 for output. Cache reads start at $0.01 per million tokens, while Batch processing reduces applicable token rates by 50%.

A low starting price does not guarantee a low completed-task cost. Longer context, thinking tokens, repeated outputs, cache misses, and retries can change the bill substantially. Start with the correct pricing tier, then calculate what your workflow actually consumes. For broader savings, use our guide on how to reduce AI API costs to review your workflow’s spending.

For broader API budgeting, explore LinkModel’s available models and its pay-as-you-go terms. The platform advertises no minimum spend or monthly platform fee. Confirm model availability, endpoint features, and applicable rates before choosing an access path or assuming a Haiku 5.5 discount.

Updated October 8, 2026. This article reviews official documentation and published measurements.

linkmodel.png

Claude Haiku 5.5 API Pricing at a Glance

Haiku 5.5 uses two prompt-length tiers, with separate charges for ordinary input, output, cache writes, and cache reads. All prices below are USD per million tokens.

Billing categoryPrompts ≤100,000 tokensPrompts >100,000 tokens
Standard input$0.10$0.50
Standard output$0.50$2.50
Five-minute cache write$0.125$0.625
One-hour cache write$0.20$1.00
Cache read$0.01$0.05
Batch input$0.05$0.25
Batch output$0.25$1.25

“Per million tokens” is a billing unit, not a minimum purchase. At the shorter-prompt rate, 10,000 ordinary input tokens cost $0.001.

These rates cover model token charges. Add applicable server-side tool fees, geographic pricing modifiers, provider charges, taxes, and account-specific adjustments separately.

How Much Cheaper Is Haiku 5.5 Than Haiku 4.5?

Haiku 5.5’s standard input and output rates are 90% lower for prompts up to 100K and 50% lower for longer prompts, compared with Haiku 4.5’s published $1 input and $5 output rates.

These are per-token comparisons. Anthropic’s approximately 75% average running-cost reduction is a different measure that considers workload distribution and tokenizer changes. It should not be presented as a guaranteed reduction in every application’s bill.

Claude Haiku 4.5 and 5.5 token prices compared, showing 90% lower rates for prompts up to 100K tokens and 50% lower rates above 100K.

How Does the 100K Pricing Threshold Work?

The threshold applies to each request’s prompt length, rather than monthly usage or the latest message alone. Above 100,000 input tokens, the request uses the higher pricing tier.

This is a tier change for the request, rather than an incremental charge applied only to tokens beyond 100K. Exactly 100,000 tokens remains within the lower tier.

Claude Haiku 5.5 input and output prices increase fivefold above 100,000 prompt tokens, from $0.10 and $0.50 to $0.50 and $2.50 per million tokens.

Cached Content Still Counts Toward Prompt Length

A cached document can make a request long even when the new question is short.

For example:

  • Cached input: 100,000 tokens.
  • New ordinary input: 50 tokens.
  • Total input processed: 100,050 tokens.

Use the total prompt length to select the tier. Then price ordinary input, cache creation, and cache reads separately. Do not charge the total prompt again after calculating its components.

System instructions, tool definitions, documents, and submitted conversation history can all contribute to the prompt.

A 150K-Token Request Example

Assume 150,000 ordinary input tokens and 2,000 billable output tokens, with no caching:

ComponentCalculationCost
Input150,000 ÷ 1,000,000 × $0.50$0.075
Output2,000 ÷ 1,000,000 × $2.50$0.005
TotalInput + output$0.080

Using the lower tier would produce $0.016, understating this example by $0.064 per request, or $640 across 10,000 identical requests.

Reduce irrelevant context when it is safe to do so, then check that the shorter prompt still produces acceptable results. Savings disappear if missing context causes additional attempts.

How Do You Calculate a Haiku 5.5 Request’s Cost?

Multiply each reported token category by its applicable rate, divide by one million, and add the results.

For a request without caching:

Token cost = (ordinary input tokens × input rate + billable output tokens × output rate) ÷ 1,000,000.

For cached requests, add cache-read and cache-write charges as separate components.

A Shorter-Prompt Request Example

Assume 10,000 ordinary input tokens and 2,000 billable output tokens:

Input: 10,000 ÷ 1,000,000 × $0.10 = $0.001.
Output: 2,000 ÷ 1,000,000 × $0.50 = $0.001.
Total: $0.002 per request.

Output costs five times as much per token as ordinary input. In this example, it contributes half the bill despite representing only one-sixth of total tokens.

If billable output rises to 10,000 tokens, the request costs $0.006. Across 100,000 requests, that change adds $400.

Do Thinking Tokens Add to the Bill?

Yes. Thinking contributes to billable output usage, so visible answer length alone is insufficient for estimating cost.

Haiku 5.5 has thinking enabled by default. Thinking also consumes part of the max_tokens allowance, leaving less room for the final response.

In the worked examples, “output tokens” means total billable output, not just the text displayed to the user. Use the endpoint’s usage reporting rather than counting visible words.

For extraction or classification, define the required fields or labels and evaluate whether additional reasoning improves accepted-result quality enough to justify its cost.

Claude Haiku 5.5 Cache Pricing and Break-Even Costs

Caching saves money when successful reuse offsets the initial write premium. Under a fixed pricing tier, a five-minute cache becomes cheaper after one successful read; a one-hour cache needs two.

Reuse patternCost relative to one ordinary input passWithout cachingCheaper?
Five-minute write + one read1.35×2×Yes
One-hour write + one read2.10×2×No
One-hour write + two reads2.20×3×Yes

These comparisons cover repeated content only. New input and output remain additional charges.

Claude Haiku 5.5 cache prices per million tokens: five-minute writes cost $0.125 or $0.625, one-hour writes $0.20 or $1.00, and reads $0.01 or $0.05, depending on prompt length.

Ten Uses of a 50K Document

Assume a 50,000-token document is written once to a five-minute cache and successfully read nine times. All requests remain in the lower tier.

Repeated-document componentCost
Ten ordinary inputs without caching$0.05000
One five-minute cache write$0.00625
Nine cache reads$0.00450
Total with caching$0.01075
Saving$0.03925 — 78.5%

The 78.5% saving applies to repeated document input, rather than the whole application bill. Questions, answers, tool use, and retries are excluded.

A Request with Mixed Cache Usage

Assume 50,000 cache-read tokens, 1,000 ordinary input tokens, and 2,000 billable output tokens, with no new writes:

Cache reads: $0.0005.
Ordinary input: $0.0001.
Output: $0.0010.
Total: $0.0016.

Output contributes 62.5% of the token cost. A high cache-hit share therefore does not establish a low total cost.

What Can Prevent Cache Savings?

Check minimum length, prefix consistency, expiration, and configuration changes. Anthropic documents a 512-token minimum cacheable prompt length for Haiku 5.5 on the listed supported platforms; Bedrock has separate platform guidance.

Confirm actual reads and writes in response usage. A cache marker alone does not prove a hit.Use our guide to organize stable prefixes and identify common causes of cache misses.

Claude Haiku 5.5 Batch API Pricing for Offline Workloads

Batch processing halves applicable token rates for asynchronous work. It suits queued extraction, evaluations, and summaries that do not require immediate responses.

Processing can take up to 24 hours, and unfinished requests can expire.

Claude Haiku 5.5 standard and Batch input and output rates compared across both prompt-length tiers, showing a 50% Batch discount.

Processing 100,000 Records

Assume each record uses 1,000 ordinary input tokens and 200 billable output tokens. Every request stays in the lower tier and uses no caching.

ComponentStandardBatch
100 million input tokens$10$5
20 million output tokens$10$5
Total$20$10

The calculated saving is $10, before additional charges and retries.

Track each record’s completion status. Retry costs belong in the workload total.

Can Batch and Caching Be Combined?

Yes. Anthropic documents that caching multipliers can stack with the Batch discount. Successful reuse still depends on request timing and cache behavior.

Halving the earlier ten-use document estimate gives $0.005375, assuming one write and nine successful reads still occur. That is a conditional calculation, not a guaranteed batch outcome.

Why Can a Cheap Model Produce an Expensive Workflow?

Total spending depends on accumulated usage across all attempts, not just the advertised rate. A published pagoda-generation comparison illustrates this distinction.

Both models ran at xhigh effort through subscriptions in atomic.chat. The author reported API-equivalent estimates rather than direct API invoices.

Reported metricHaiku 5.5GPT-6 Luna
Total input268.8M tokens64.5M tokens
Cached input262.7M tokens61.1M tokens
Output4.46M tokens1.18M tokens
Estimated API-equivalent cost$24.96$1.96

Haiku’s reported estimate was approximately 12.73 times higher, with approximately 4.17 times the input and 3.78 times the output.

This single comparison does not establish a universal winner or equal output quality. Its useful lesson is that extensive cache reuse can coexist with substantial accumulated spending.For a broader model comparison, read our Claude Haiku 5.5 vs GPT-6 Luna guide alongside this single-workflow example.

Reported voxel pagoda token usage: Haiku 5.5 used 268.8 million input and 4.46 million output tokens; GPT-6 Luna used 64.5 million input and 1.18 million output tokens.

Can the Reported $24.96 Be Reconstructed?

Subtracting cached input from total input leaves 6.1 million noncached tokens.

Assuming all requests use the higher tier, treating noncached tokens as ordinary input, and excluding cache-write premiums produces:

Ordinary input: 6.1 × $0.50 = $3.05.
Cache reads: 262.7 × $0.05 = $13.135.
Output: 4.46 × $2.50 = $11.15.
Total: $27.335.

Using lower-tier rates instead produces $5.467. Neither reconstruction reproduces the reported $24.96.

Request-level tiers, cache-write details, and the estimator’s method are needed to reconcile the total. Aggregate token counts alone are insufficient.When estimating the Luna side of the comparison, use our GPT-6 Luna API pricing guide to check its billing rules separately.

Reported API-equivalent voxel pagoda cost estimates: $24.96 for Claude Haiku 5.5 and $1.96 for GPT-6 Luna, approximately a 12.73-fold difference; these are not verified API invoices.

How Do Effort and Latency Affect the Budget?

Start with medium effort, then compare settings using accepted-result cost and delivery time. Anthropic documents medium as the default, low as the cheapest and fastest level, and high, xhigh, and max as options for more demanding work.

Higher effort can generate substantially more thinking and output. Changing top-level effort between requests also invalidates the conversation-message cache.

A useful comparison is:

Cost per accepted result = total workflow spend ÷ accepted results.

For example:

$20 ÷ 8,000 accepted results = $0.0025.
$18 ÷ 6,000 accepted results = $0.0030.

The second run spends less overall but costs more per usable result. These are illustrative calculations, not measured model outcomes.

Throughput Is Different from Initial Waiting Time

Artificial Analysis’s Haiku 5.5 Low-effort snapshot reports:

185 output tokens per second.
12.79 seconds to first token.
Intelligence Index: 29, a composite score rather than percentage accuracy.
Approximately $0.02 per Intelligence Index task.

These measurements describe the benchmark’s conditions. They are not a production latency guarantee or a universal request price.

Use our AI API latency guide to plan your evaluation, then measure first useful answer time, total completion time, and retries on representative requests before making a deployment decision.

How Should You Compare Direct API and Provider Pricing?

Compare the deployed route’s complete billing rules, rather than its starting-price card.

The direct Claude model identifier is claude-haiku-5-5. OpenRouter lists anthropic/claude-haiku-5.5, with starting prices matching the lower-tier input and output rates.

Before comparing access paths, confirm:

Exact model and provider route.
Long-prompt pricing.
Cache support and charges.
Batch availability.
Geographic modifiers and additional fees.
Usage visibility and completed-task reliability.

For LinkModel, check its model catalog before assuming Haiku 5.5 availability. Its advertised “up to 30%” savings is a platform proposition, not a verified Haiku 5.5 discount.

Research Sources and Method

The pricing tables use official published rates. Worked examples, break-even comparisons, percentages, and cost reconstructions are calculations under stated assumptions.

We did not independently run the external tests, inspect private invoices, or verify Haiku 5.5 availability through LinkModel.

EvidenceSourceUse in this article
Official launchAnthropic’s Haiku 5.5 announcementVersion comparison and average savings distinction.
Official billingClaude API pricingStandard rates, tiers, and pricing modifiers.
Cache mechanicsPrompt caching documentationInput categories, minimum length, and cache behavior.
Offline processingBatch processing documentationDiscount, timing, and feature combination.
ConfigurationHaiku 5.5 prompting guideEffort defaults, output allowance, and cache invalidation.
Thinking usageClaude thinking documentationBillable reasoning usage and adaptive-thinking migration.
Independent measurementArtificial Analysis Low-effort pageThroughput, latency, index, and evaluation cost.
Published developer caseVoxel pagoda comparisonReported usage and API-equivalent estimates.
API access listingOpenRouter’s Haiku 5.5 pageStarting prices and platform identifier
Product termsLinkModelAdvertised platform terms and savings; model-specific availability remains unverified.

Conclusion: Budget for Completed Work

Claude Haiku 5.5 starts at $0.10 input and $0.50 output per million tokens, but an accurate budget requires the correct prompt tier and complete usage breakdown. Separate cache operations, include thinking and generated output, and account for retries.

Use the worked examples for an initial estimate, then reconcile actual usage with charges. Choose the effort setting and access path that meet your quality and delivery requirements at the lowest measured cost per accepted result.

Frequently Asked Questions

Is Haiku 5.5 $0.10 Input and $0.50 Output for Every Request?

No. Those standard rates apply to prompts up to 100,000 tokens. Longer prompts use $0.50 input and $2.50 output rates per million tokens.

Does the Higher Tier Apply Only to Tokens Above 100K?

No. Prompt length selects the pricing tier for the request. Calculate each input category and output at the applicable tier’s rates.

Do Cached Tokens Count Toward the Threshold?

Yes. Include cache reads, cache creation, and ordinary input when determining total prompt length.

Is a Cache Read Charged Again as Ordinary Input?

No. Calculate each reported category separately. Charging cache-read tokens again at the ordinary input rate would double-count them.

Can I Combine Batch Processing with Prompt Caching?

Yes, where supported. The discounts can stack, but expected savings require actual cache hits. Confirm them in completed-request usage.

Does Low Effort Have a Fixed Price per Request?

No.Effort influences usage, but request cost still depends on prompt length, billable output, caching, and additional attempts.

About the author

Fiona Thorne

Fiona Thorne

AI Model & API Researcher at LinkModel

Fiona Thorne is an AI model and API researcher at LinkModel, focusing on generative AI technologies, model capabilities, API pricing, and practical integration strategies. She explores developments across leading AI providers, drawing on official documentation, technical specifications, and comparative research to help developers and businesses evaluate AI solutions, understand their trade-offs, and make informed technology decisions.

Related Posts