Anthropic/claude-opus-4-6
Anthropic's frontier flagship model built for deep reasoning, long-horizon agentic coding, and enterprise knowledge work — tops Terminal-Bench 2.0 and HLE with 1M context and adaptive thinking for the hardest tasks.
More from Anthropic
README
Anthropic/claude-opus-4-6
Supported Functionality
| Item | Specification |
|---|---|
| Input | Text, Image |
| Output | Text |
| Context | 200,000 tokens (1,000,000 in beta) |
| Max Output | 128,000 tokens |
| Vision | ✓ Supported |
| Function Calling | ✓ Supported (Agent Teams, Context Compaction, MCP) |
Description
Claude Opus 4.6 is Anthropic's frontier flagship model, released on February 5, 2026 and available via the Claude API as claude-opus-4-6. It is the first Opus-class model to ship a 1M-token context window (beta) and the first Claude model with Adaptive Thinking plus four effort levels (low / medium / high / max), producing up to 128K output tokens. Pricing is $5 / $25 per million input/output tokens (with $10 / $37.50 for inputs over 200K).
Opus 4.6's defining breakthrough is turning frontier reasoning into sustained, reliable long-horizon execution. It tops Terminal-Bench 2.0 (65.4%), Humanity's Last Exam (53.0% with tools), ARC-AGI-2 (69.17%), ARC-AGI-1 (94.00%), and BrowseComp (84.0%); it finishes Vending-Bench 2's year-long simulation with a final balance of $8,017 (well above Gemini 3 Pro's prior SOTA); and on MRCR v2 8-needle 1M retrieval it scores 76% versus Sonnet 4.5's 18.5%. The release also introduces Agent Teams in Claude Code (parallel multi-agent collaboration), Context Compaction on the API, a major Claude in Excel upgrade, and a Claude in PowerPoint research preview — a full lift for enterprise workflows.
Key Capabilities
- Top-Tier Agentic Coding: #1 on Terminal-Bench 2.0 at 65.4%; SWE-Bench Verified 80.8%; handles cross-file refactors and complex bug diagnosis autonomously.
- Frontier Deep Reasoning: HLE 53.0% (with tools), ARC-AGI-2 69.17%, ARC-AGI-1 94.00% — leads on expert-level multi-disciplinary and novel-problem reasoning.
- Long-Context Retrieval: MRCR v2 8-needle 1M at 76% (Sonnet 4.5: 18.5%) — dramatically reduced context rot, supporting deep reasoning across full repos and large document sets.
- Adaptive Thinking + Effort Control: The model auto-tunes reasoning depth; developers further adjust via low / medium / high / max for precise quality/latency/cost tradeoffs.
- Agent Teams in Claude Code: Multiple Claude instances collaborate in parallel on a single task — ideal for code review and read-heavy large-codebase work.
- Enterprise Knowledge Work: GDPval-AA +190 Elo over Opus 4.5 and +144 over GPT-5.2; BigLaw Bench legal reasoning 90.2%; BrowseComp agentic search 84.0%.
- Life Sciences & Cybersecurity: ~2× improvement over Opus 4.5 on computational/structural biology, organic chemistry, and phylogenetics; industry-leading on CyberGym vulnerability discovery.
Technical Strengths
| Feature | Benefit |
|---|---|
| First 1M Context in the Opus Tier | Frontier-grade reasoning over entire codebases, hundreds of contracts, or full research corpora — no RAG plumbing required |
| Adaptive Thinking + Effort Levels | Model auto-tunes reasoning depth and developers get four explicit dials for quality/latency/cost trade-offs |
| Context Compaction (API) | Automatic context summarization lets very long agentic runs continue beyond raw token limits |
| Agent Teams Orchestration | Multiple Claude agents run in parallel inside Claude Code, accelerating large-scale and read-heavy engineering |
| Sustained Coherence | Vending-Bench 2 final balance of $8,017 demonstrates strategic planning and self-correction across thousands of decisions |
| Anthropic Safety & Alignment | Automated audits show low rates of deception, sycophancy, and delusion reinforcement — capability gains without safety regressions |
Capability Ratings
| Dimension | Rating | Notes |
|---|---|---|
| Reasoning | Top-tier | Leads HLE, ARC-AGI-2, and BigLaw Bench — undisputed frontier expert reasoning |
| Coding | Top-tier | #1 on Terminal-Bench 2.0; SWE-Bench Verified 80.8%; strongest at complex cross-file refactors |
| Creative Writing | Top-tier | Anthropic's signature voice with strong pacing, restraint, and cultural depth |
| Multimodal | Strong | Text + image input; no native video/audio (later improved in Opus 4.7) |
| Response Speed | Moderate | Flagship priced for depth over latency; Low effort mode trades reasoning for speed |
| Context Window | Huge | 200K default and 1M beta; MRCR long-context retrieval is best-in-class among contemporary models |
Use Cases
- Complex Codebase Refactoring & Diagnosis: #1 Terminal-Bench plus 1M context enables whole-repo analysis, multi-file refactors, and tough concurrency bug fixes.
- Multi-Agent Engineering Workflows: Claude Code Agent Teams let multiple Opus instances review, write, and test in parallel for large engineering projects.
- Enterprise Research & Deep Retrieval: 84.0% on BrowseComp plus long context make it ideal for multi-step research, cross-document synthesis, and intelligence work.
- Legal / Finance / Life Sciences: BigLaw Bench 90.2%, Real-World Finance SOTA, and near-2× life sciences gains support high-stakes professional agents.
- Long-Horizon Strategic Planning: Vending-Bench 2 results show sustained coherence over thousands of decisions — useful for business simulation and operational planning.
- Office Productivity: Upgraded Claude in Excel and Claude in PowerPoint research preview extend Opus into spreadsheets, documents, and slide decks.
- Highest-Difficulty Research Assistant: SOTA on HLE, GPQA, and ARC-AGI-2 makes it a reliable partner on frontier math, physics, biology, and CS problems.
Pricing
| Token Type | LinkAI Price | Official Price |
|---|---|---|
| input | $3.750000 / 1M tokens | $5.000000 / 1M tokens |
| output | $18.750000 / 1M tokens | $25.000000 / 1M tokens |
| cache_read | $0.375000 / 1M tokens | $0.500000 / 1M tokens |
| cache_write_5m | $4.687500 / 1M tokens | $6.250000 / 1M tokens |
| cache_write_1h | $7.500000 / 1M tokens | $10.000000 / 1M tokens |