Claude Haiku 5.5 API pricing starts at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Longer prompts cost $0.50 for input and $2.50 for output. Cache reads start at $0.01 per million tokens, while Batch processing reduces applicable token rates by 50%.
A low starting price does not guarantee a low completed-task cost. Longer context, thinking tokens, repeated outputs, cache misses, and retries can change the bill substantially. Start with the correct pricing tier, then calculate what your workflow actually consumes. For broader savings, use our guide on how to reduce AI API costs to review your workflow’s spending.
For broader API budgeting, explore LinkModel’s available models and its pay-as-you-go terms. The platform advertises no minimum spend or monthly platform fee. Confirm model availability, endpoint features, and applicable rates before choosing an access path or assuming a Haiku 5.5 discount.
Updated October 8, 2026. This article reviews official documentation and published measurements.

Claude Haiku 5.5 API Pricing at a Glance
Haiku 5.5 uses two prompt-length tiers, with separate charges for ordinary input, output, cache writes, and cache reads. All prices below are USD per million tokens.
| Billing category | Prompts ≤100,000 tokens | Prompts >100,000 tokens |
|---|---|---|
| Standard input | $0.10 | $0.50 |
| Standard output | $0.50 | $2.50 |
| Five-minute cache write | $0.125 | $0.625 |
| One-hour cache write | $0.20 | $1.00 |
| Cache read | $0.01 | $0.05 |
| Batch input | $0.05 | $0.25 |
| Batch output | $0.25 | $1.25 |
“Per million tokens” is a billing unit, not a minimum purchase. At the shorter-prompt rate, 10,000 ordinary input tokens cost $0.001.
These rates cover model token charges. Add applicable server-side tool fees, geographic pricing modifiers, provider charges, taxes, and account-specific adjustments separately.
How Much Cheaper Is Haiku 5.5 Than Haiku 4.5?
Haiku 5.5’s standard input and output rates are 90% lower for prompts up to 100K and 50% lower for longer prompts, compared with Haiku 4.5’s published $1 input and $5 output rates.
These are per-token comparisons. Anthropic’s approximately 75% average running-cost reduction is a different measure that considers workload distribution and tokenizer changes. It should not be presented as a guaranteed reduction in every application’s bill.

How Does the 100K Pricing Threshold Work?
The threshold applies to each request’s prompt length, rather than monthly usage or the latest message alone. Above 100,000 input tokens, the request uses the higher pricing tier.
This is a tier change for the request, rather than an incremental charge applied only to tokens beyond 100K. Exactly 100,000 tokens remains within the lower tier.

Cached Content Still Counts Toward Prompt Length
A cached document can make a request long even when the new question is short.
For example:
- Cached input: 100,000 tokens.
- New ordinary input: 50 tokens.
- Total input processed: 100,050 tokens.
Use the total prompt length to select the tier. Then price ordinary input, cache creation, and cache reads separately. Do not charge the total prompt again after calculating its components.
System instructions, tool definitions, documents, and submitted conversation history can all contribute to the prompt.
A 150K-Token Request Example
Assume 150,000 ordinary input tokens and 2,000 billable output tokens, with no caching:
| Component | Calculation | Cost |
|---|---|---|
| Input | 150,000 ÷ 1,000,000 × $0.50 | $0.075 |
| Output | 2,000 ÷ 1,000,000 × $2.50 | $0.005 |
| Total | Input + output | $0.080 |
Using the lower tier would produce $0.016, understating this example by $0.064 per request, or $640 across 10,000 identical requests.
Reduce irrelevant context when it is safe to do so, then check that the shorter prompt still produces acceptable results. Savings disappear if missing context causes additional attempts.
How Do You Calculate a Haiku 5.5 Request’s Cost?
Multiply each reported token category by its applicable rate, divide by one million, and add the results.
For a request without caching:
Token cost = (ordinary input tokens × input rate + billable output tokens × output rate) ÷ 1,000,000.
For cached requests, add cache-read and cache-write charges as separate components.
A Shorter-Prompt Request Example
Assume 10,000 ordinary input tokens and 2,000 billable output tokens:
Input: 10,000 ÷ 1,000,000 × $0.10 = $0.001.
Output: 2,000 ÷ 1,000,000 × $0.50 = $0.001.
Total: $0.002 per request.
Output costs five times as much per token as ordinary input. In this example, it contributes half the bill despite representing only one-sixth of total tokens.
If billable output rises to 10,000 tokens, the request costs $0.006. Across 100,000 requests, that change adds $400.
Do Thinking Tokens Add to the Bill?
Yes. Thinking contributes to billable output usage, so visible answer length alone is insufficient for estimating cost.
Haiku 5.5 has thinking enabled by default. Thinking also consumes part of the max_tokens allowance, leaving less room for the final response.
In the worked examples, “output tokens” means total billable output, not just the text displayed to the user. Use the endpoint’s usage reporting rather than counting visible words.
For extraction or classification, define the required fields or labels and evaluate whether additional reasoning improves accepted-result quality enough to justify its cost.
Claude Haiku 5.5 Cache Pricing and Break-Even Costs
Caching saves money when successful reuse offsets the initial write premium. Under a fixed pricing tier, a five-minute cache becomes cheaper after one successful read; a one-hour cache needs two.
| Reuse pattern | Cost relative to one ordinary input pass | Without caching | Cheaper? |
|---|---|---|---|
| Five-minute write + one read | 1.35× | 2× | Yes |
| One-hour write + one read | 2.10× | 2× | No |
| One-hour write + two reads | 2.20× | 3× | Yes |
These comparisons cover repeated content only. New input and output remain additional charges.

Ten Uses of a 50K Document
Assume a 50,000-token document is written once to a five-minute cache and successfully read nine times. All requests remain in the lower tier.
| Repeated-document component | Cost |
|---|---|
| Ten ordinary inputs without caching | $0.05000 |
| One five-minute cache write | $0.00625 |
| Nine cache reads | $0.00450 |
| Total with caching | $0.01075 |
| Saving | $0.03925 — 78.5% |
The 78.5% saving applies to repeated document input, rather than the whole application bill. Questions, answers, tool use, and retries are excluded.
A Request with Mixed Cache Usage
Assume 50,000 cache-read tokens, 1,000 ordinary input tokens, and 2,000 billable output tokens, with no new writes:
Cache reads: $0.0005.
Ordinary input: $0.0001.
Output: $0.0010.
Total: $0.0016.
Output contributes 62.5% of the token cost. A high cache-hit share therefore does not establish a low total cost.
What Can Prevent Cache Savings?
Check minimum length, prefix consistency, expiration, and configuration changes. Anthropic documents a 512-token minimum cacheable prompt length for Haiku 5.5 on the listed supported platforms; Bedrock has separate platform guidance.
Confirm actual reads and writes in response usage. A cache marker alone does not prove a hit.Use our guide to organize stable prefixes and identify common causes of cache misses.
Claude Haiku 5.5 Batch API Pricing for Offline Workloads
Batch processing halves applicable token rates for asynchronous work. It suits queued extraction, evaluations, and summaries that do not require immediate responses.
Processing can take up to 24 hours, and unfinished requests can expire.

Processing 100,000 Records
Assume each record uses 1,000 ordinary input tokens and 200 billable output tokens. Every request stays in the lower tier and uses no caching.
| Component | Standard | Batch |
|---|---|---|
| 100 million input tokens | $10 | $5 |
| 20 million output tokens | $10 | $5 |
| Total | $20 | $10 |
The calculated saving is $10, before additional charges and retries.
Track each record’s completion status. Retry costs belong in the workload total.
Can Batch and Caching Be Combined?
Yes. Anthropic documents that caching multipliers can stack with the Batch discount. Successful reuse still depends on request timing and cache behavior.
Halving the earlier ten-use document estimate gives $0.005375, assuming one write and nine successful reads still occur. That is a conditional calculation, not a guaranteed batch outcome.
Why Can a Cheap Model Produce an Expensive Workflow?
Total spending depends on accumulated usage across all attempts, not just the advertised rate. A published pagoda-generation comparison illustrates this distinction.
Both models ran at xhigh effort through subscriptions in atomic.chat. The author reported API-equivalent estimates rather than direct API invoices.
| Reported metric | Haiku 5.5 | GPT-6 Luna |
|---|---|---|
| Total input | 268.8M tokens | 64.5M tokens |
| Cached input | 262.7M tokens | 61.1M tokens |
| Output | 4.46M tokens | 1.18M tokens |
| Estimated API-equivalent cost | $24.96 | $1.96 |
Haiku’s reported estimate was approximately 12.73 times higher, with approximately 4.17 times the input and 3.78 times the output.
This single comparison does not establish a universal winner or equal output quality. Its useful lesson is that extensive cache reuse can coexist with substantial accumulated spending.For a broader model comparison, read our Claude Haiku 5.5 vs GPT-6 Luna guide alongside this single-workflow example.

Can the Reported $24.96 Be Reconstructed?
Subtracting cached input from total input leaves 6.1 million noncached tokens.
Assuming all requests use the higher tier, treating noncached tokens as ordinary input, and excluding cache-write premiums produces:
Ordinary input: 6.1 × $0.50 = $3.05.
Cache reads: 262.7 × $0.05 = $13.135.
Output: 4.46 × $2.50 = $11.15.
Total: $27.335.
Using lower-tier rates instead produces $5.467. Neither reconstruction reproduces the reported $24.96.
Request-level tiers, cache-write details, and the estimator’s method are needed to reconcile the total. Aggregate token counts alone are insufficient.When estimating the Luna side of the comparison, use our GPT-6 Luna API pricing guide to check its billing rules separately.

How Do Effort and Latency Affect the Budget?
Start with medium effort, then compare settings using accepted-result cost and delivery time. Anthropic documents medium as the default, low as the cheapest and fastest level, and high, xhigh, and max as options for more demanding work.
Higher effort can generate substantially more thinking and output. Changing top-level effort between requests also invalidates the conversation-message cache.
A useful comparison is:
Cost per accepted result = total workflow spend ÷ accepted results.
For example:
$20 ÷ 8,000 accepted results = $0.0025.
$18 ÷ 6,000 accepted results = $0.0030.
The second run spends less overall but costs more per usable result. These are illustrative calculations, not measured model outcomes.
Throughput Is Different from Initial Waiting Time
Artificial Analysis’s Haiku 5.5 Low-effort snapshot reports:
185 output tokens per second.
12.79 seconds to first token.
Intelligence Index: 29, a composite score rather than percentage accuracy.
Approximately $0.02 per Intelligence Index task.
These measurements describe the benchmark’s conditions. They are not a production latency guarantee or a universal request price.
Use our AI API latency guide to plan your evaluation, then measure first useful answer time, total completion time, and retries on representative requests before making a deployment decision.
How Should You Compare Direct API and Provider Pricing?
Compare the deployed route’s complete billing rules, rather than its starting-price card.
The direct Claude model identifier is claude-haiku-5-5. OpenRouter lists anthropic/claude-haiku-5.5, with starting prices matching the lower-tier input and output rates.
Before comparing access paths, confirm:
Exact model and provider route.
Long-prompt pricing.
Cache support and charges.
Batch availability.
Geographic modifiers and additional fees.
Usage visibility and completed-task reliability.
For LinkModel, check its model catalog before assuming Haiku 5.5 availability. Its advertised “up to 30%” savings is a platform proposition, not a verified Haiku 5.5 discount.
Research Sources and Method
The pricing tables use official published rates. Worked examples, break-even comparisons, percentages, and cost reconstructions are calculations under stated assumptions.
We did not independently run the external tests, inspect private invoices, or verify Haiku 5.5 availability through LinkModel.
| Evidence | Source | Use in this article |
|---|---|---|
| Official launch | Anthropic’s Haiku 5.5 announcement | Version comparison and average savings distinction. |
| Official billing | Claude API pricing | Standard rates, tiers, and pricing modifiers. |
| Cache mechanics | Prompt caching documentation | Input categories, minimum length, and cache behavior. |
| Offline processing | Batch processing documentation | Discount, timing, and feature combination. |
| Configuration | Haiku 5.5 prompting guide | Effort defaults, output allowance, and cache invalidation. |
| Thinking usage | Claude thinking documentation | Billable reasoning usage and adaptive-thinking migration. |
| Independent measurement | Artificial Analysis Low-effort page | Throughput, latency, index, and evaluation cost. |
| Published developer case | Voxel pagoda comparison | Reported usage and API-equivalent estimates. |
| API access listing | OpenRouter’s Haiku 5.5 page | Starting prices and platform identifier |
| Product terms | LinkModel | Advertised platform terms and savings; model-specific availability remains unverified. |
Conclusion: Budget for Completed Work
Claude Haiku 5.5 starts at $0.10 input and $0.50 output per million tokens, but an accurate budget requires the correct prompt tier and complete usage breakdown. Separate cache operations, include thinking and generated output, and account for retries.
Use the worked examples for an initial estimate, then reconcile actual usage with charges. Choose the effort setting and access path that meet your quality and delivery requirements at the lowest measured cost per accepted result.
Frequently Asked Questions
Is Haiku 5.5 $0.10 Input and $0.50 Output for Every Request?
No. Those standard rates apply to prompts up to 100,000 tokens. Longer prompts use $0.50 input and $2.50 output rates per million tokens.
Does the Higher Tier Apply Only to Tokens Above 100K?
No. Prompt length selects the pricing tier for the request. Calculate each input category and output at the applicable tier’s rates.
Do Cached Tokens Count Toward the Threshold?
Yes. Include cache reads, cache creation, and ordinary input when determining total prompt length.
Is a Cache Read Charged Again as Ordinary Input?
No. Calculate each reported category separately. Charging cache-read tokens again at the ordinary input rate would double-count them.
Can I Combine Batch Processing with Prompt Caching?
Yes, where supported. The discounts can stack, but expected savings require actual cache hits. Confirm them in completed-request usage.
Does Low Effort Have a Fixed Price per Request?
No.Effort influences usage, but request cost still depends on prompt length, billable output, caching, and additional attempts.
