Anthropic/claude-opus-4-8
Anthropic's flagship LLM for agentic coding, multidisciplinary reasoning, and knowledge work, with sharper judgment and stronger honesty, a 1M-token context window, and fast mode—built for code migration, deep research, and legal/financial analysis.
More from Anthropic
README
Anthropic/claude-opus-4-8
Supported Functionality
| Item | Specification |
|---|---|
| Input | Text, Image |
| Output | Text |
| Context | 1,000,000 tokens (default on Claude API / Amazon Bedrock / Vertex AI; 200k on Microsoft Foundry) |
| Max Output | 128,000 tokens |
| Vision | ✓ Supported |
| Function Calling | ✓ Supported |
Description
Claude Opus 4.8 is Anthropic's flagship large language model, released on May 28, 2026. It is the newest member of the Opus series and the company's most capable generally available model to date. Building on Opus 4.7, it delivers improvements across coding, agentic skills, reasoning, and knowledge-work benchmarks, and supports text and image input, adaptive thinking, and full tool use. Arriving just ~41 days after Opus 4.7, it marks a notably faster release cadence, with pricing unchanged from its predecessor.
The most prominent gains in this release are in honesty and judgment. Opus 4.8 is more likely to flag uncertainties and less likely to make unsupported claims; in Anthropic's evaluations it is roughly four times less likely than its predecessor to let flaws in its own code pass unremarked. Alignment assessments report new highs on prosocial traits like supporting user autonomy and acting in the user's best interest, with rates of misaligned behavior substantially lower than Opus 4.7 and similar to the best-aligned Claude Mythos Preview. A new fast mode runs at 2.5× speed and is three times cheaper than the prior fast mode.
Key Capabilities
- Agentic Coding: Scores 69.2% on SWE-Bench Pro, outperforming GPT-5.5 and Gemini 3.1 Pro, and can run codebase-scale migrations across hundreds of thousands of lines from kickoff to merge in Claude Code.
- Multidisciplinary Reasoning: Multidisciplinary reasoning with tools rises from 54.7% to 57.9%, handling deeper, multi-step problems.
- Honesty & Self-Verification: Proactively flags issues in inputs and outputs and is ~4× less likely than its predecessor to let its own code flaws pass unremarked, raising the signal-to-noise ratio on analysis.
- Computer & Browser Use: Scores 84% on Online-Mind2Web, a meaningful jump over both Opus 4.7 and GPT-5.5, staying reflective and on-task across long workloads.
- Long Context: Supports a 1M-token context window by default, sustaining long content, style, and context consistency within a single session.
- Multimodal Understanding: Accepts image input and reasons directly over unstructured content such as PDFs and diagrams at lower token cost than Opus 4.7.
- Knowledge Work & Effort Control: Leads on legal, finance, and research workloads, with adjustable effort to balance response speed, reasoning depth, and cost.
Technical Strengths
| Feature | Benefit |
|---|---|
| Higher honesty | Fewer overconfident false-progress claims, making long agentic outputs more trustworthy and auditable |
| Sharper judgment | Asks the right questions, catches its own mistakes, and pushes back on unsound plans, cutting rework |
| Fast mode | 2.5× speed and three times cheaper than the prior fast mode, balancing latency and budget |
| Effort control | High / extra / max settings let users trade quality against cost per task difficulty |
| Million-token context | Holds massive codebases and documents in one session for codebase-scale migration and long-doc analysis |
| Dynamic workflows (Claude Code) | Plans work, runs hundreds of parallel subagents, and verifies before reporting—handling very large-scale problems |
Capability Ratings
| Dimension | Rating | Notes |
|---|---|---|
| Reasoning | Top-tier | Multidisciplinary reasoning with tools up to 57.9%, strong on deep multi-step problems |
| Coding | Top-tier | 69.2% on SWE-Bench Pro, ahead of major rivals; excels at codebase-scale migration |
| Creative Writing | Excellent | Holds voice, style, and technical execution consistently across long sessions |
| Multimodal | Strong | Image input and reasoning over PDFs and diagrams, though output is text-only |
| Response Speed | Moderate | Defaults to high effort; reaches 2.5× in fast mode |
| Context Window | Huge | 1M tokens by default, among the largest available |
Use Cases
- Large-scale code migration: Run migrations across hundreds of thousands of lines from kickoff to merge in Claude Code, using the existing test suite as the bar.
- Autonomous agentic workflows: Use tools cleanly and follow instructions unattended for long stretches, ideal for autonomous engineering and super-agent scenarios.
- Deep research & analysis: Produce higher-density analysis while proactively flagging issues in inputs and outputs, improving efficiency and trust.
- Legal & financial professional work: Deliver greater consistency, citation precision, and reasoning quality in legal-agent and financial-document workflows.
- Computer & browser automation: Execute end-to-end web-operation agent tasks with best-in-class computer-use performance.
- Long-document & multimodal processing: Retrieve and reason over dense filings, diagrams, and PDFs using the million-token context and image understanding.
- Enterprise knowledge work: Provide reliable, effort-adjustable output across data Q&A, translation, and slide generation.
Pricing
| Token Type | LinkAI Price | Official Price |
|---|---|---|
| input | $4.250000 / 1M tokens | $5.000000 / 1M tokens |
| output | $21.250000 / 1M tokens | $25.000000 / 1M tokens |
| cache_read | $0.425000 / 1M tokens | $0.500000 / 1M tokens |
| cache_write_5m | $5.312500 / 1M tokens | $6.250000 / 1M tokens |
| cache_write_1h | $8.500000 / 1M tokens | $10.000000 / 1M tokens |