Anthropic

Anthropic/claude-opus-4-6

From $0.375 / 1M tokens

Anthropic's frontier flagship model built for deep reasoning, long-horizon agentic coding, and enterprise knowledge work — tops Terminal-Bench 2.0 and HLE with 1M context and adaptive thinking for the hardest tasks.

Chat

More from Anthropic

README

Anthropic/claude-opus-4-6

Supported Functionality

ItemSpecification
InputText, Image
OutputText
Context200,000 tokens (1,000,000 in beta)
Max Output128,000 tokens
Vision✓ Supported
Function Calling✓ Supported (Agent Teams, Context Compaction, MCP)

Description

Claude Opus 4.6 is Anthropic's frontier flagship model, released on February 5, 2026 and available via the Claude API as claude-opus-4-6. It is the first Opus-class model to ship a 1M-token context window (beta) and the first Claude model with Adaptive Thinking plus four effort levels (low / medium / high / max), producing up to 128K output tokens. Pricing is $5 / $25 per million input/output tokens (with $10 / $37.50 for inputs over 200K).

Opus 4.6's defining breakthrough is turning frontier reasoning into sustained, reliable long-horizon execution. It tops Terminal-Bench 2.0 (65.4%), Humanity's Last Exam (53.0% with tools), ARC-AGI-2 (69.17%), ARC-AGI-1 (94.00%), and BrowseComp (84.0%); it finishes Vending-Bench 2's year-long simulation with a final balance of $8,017 (well above Gemini 3 Pro's prior SOTA); and on MRCR v2 8-needle 1M retrieval it scores 76% versus Sonnet 4.5's 18.5%. The release also introduces Agent Teams in Claude Code (parallel multi-agent collaboration), Context Compaction on the API, a major Claude in Excel upgrade, and a Claude in PowerPoint research preview — a full lift for enterprise workflows.

Key Capabilities

  • Top-Tier Agentic Coding: #1 on Terminal-Bench 2.0 at 65.4%; SWE-Bench Verified 80.8%; handles cross-file refactors and complex bug diagnosis autonomously.
  • Frontier Deep Reasoning: HLE 53.0% (with tools), ARC-AGI-2 69.17%, ARC-AGI-1 94.00% — leads on expert-level multi-disciplinary and novel-problem reasoning.
  • Long-Context Retrieval: MRCR v2 8-needle 1M at 76% (Sonnet 4.5: 18.5%) — dramatically reduced context rot, supporting deep reasoning across full repos and large document sets.
  • Adaptive Thinking + Effort Control: The model auto-tunes reasoning depth; developers further adjust via low / medium / high / max for precise quality/latency/cost tradeoffs.
  • Agent Teams in Claude Code: Multiple Claude instances collaborate in parallel on a single task — ideal for code review and read-heavy large-codebase work.
  • Enterprise Knowledge Work: GDPval-AA +190 Elo over Opus 4.5 and +144 over GPT-5.2; BigLaw Bench legal reasoning 90.2%; BrowseComp agentic search 84.0%.
  • Life Sciences & Cybersecurity: ~2× improvement over Opus 4.5 on computational/structural biology, organic chemistry, and phylogenetics; industry-leading on CyberGym vulnerability discovery.

Technical Strengths

FeatureBenefit
First 1M Context in the Opus TierFrontier-grade reasoning over entire codebases, hundreds of contracts, or full research corpora — no RAG plumbing required
Adaptive Thinking + Effort LevelsModel auto-tunes reasoning depth and developers get four explicit dials for quality/latency/cost trade-offs
Context Compaction (API)Automatic context summarization lets very long agentic runs continue beyond raw token limits
Agent Teams OrchestrationMultiple Claude agents run in parallel inside Claude Code, accelerating large-scale and read-heavy engineering
Sustained CoherenceVending-Bench 2 final balance of $8,017 demonstrates strategic planning and self-correction across thousands of decisions
Anthropic Safety & AlignmentAutomated audits show low rates of deception, sycophancy, and delusion reinforcement — capability gains without safety regressions

Capability Ratings

DimensionRatingNotes
ReasoningTop-tierLeads HLE, ARC-AGI-2, and BigLaw Bench — undisputed frontier expert reasoning
CodingTop-tier#1 on Terminal-Bench 2.0; SWE-Bench Verified 80.8%; strongest at complex cross-file refactors
Creative WritingTop-tierAnthropic's signature voice with strong pacing, restraint, and cultural depth
MultimodalStrongText + image input; no native video/audio (later improved in Opus 4.7)
Response SpeedModerateFlagship priced for depth over latency; Low effort mode trades reasoning for speed
Context WindowHuge200K default and 1M beta; MRCR long-context retrieval is best-in-class among contemporary models

Use Cases

  • Complex Codebase Refactoring & Diagnosis: #1 Terminal-Bench plus 1M context enables whole-repo analysis, multi-file refactors, and tough concurrency bug fixes.
  • Multi-Agent Engineering Workflows: Claude Code Agent Teams let multiple Opus instances review, write, and test in parallel for large engineering projects.
  • Enterprise Research & Deep Retrieval: 84.0% on BrowseComp plus long context make it ideal for multi-step research, cross-document synthesis, and intelligence work.
  • Legal / Finance / Life Sciences: BigLaw Bench 90.2%, Real-World Finance SOTA, and near-2× life sciences gains support high-stakes professional agents.
  • Long-Horizon Strategic Planning: Vending-Bench 2 results show sustained coherence over thousands of decisions — useful for business simulation and operational planning.
  • Office Productivity: Upgraded Claude in Excel and Claude in PowerPoint research preview extend Opus into spreadsheets, documents, and slide decks.
  • Highest-Difficulty Research Assistant: SOTA on HLE, GPQA, and ARC-AGI-2 makes it a reliable partner on frontier math, physics, biology, and CS problems.

Pricing

Token TypeLinkAI PriceOfficial Price
input$3.750000 / 1M tokens$5.000000 / 1M tokens
output$18.750000 / 1M tokens$25.000000 / 1M tokens
cache_read$0.375000 / 1M tokens$0.500000 / 1M tokens
cache_write_5m$4.687500 / 1M tokens$6.250000 / 1M tokens
cache_write_1h$7.500000 / 1M tokens$10.000000 / 1M tokens