Anthropic

Anthropic/claude-opus-4-5-20251101

From $0.375 / 1M tokens

Anthropic's 4th-gen flagship setting state-of-the-art on SWE-Bench Verified and ARC-AGI-2, introducing effort control and context compaction, with industry-leading prompt-injection robustness — built for coding, reasoning, and enterprise agents.

Chat

More from Anthropic

README

Anthropic/claude-opus-4-5-20251101

Supported Functionality

ItemSpecification
InputText, Image
OutputText
Context200,000 tokens
Max Output64,000 tokens
Vision✓ Supported
Function Calling✓ Supported (computer use, context compaction, MCP)

Description

Claude Opus 4.5 is Anthropic's 4th-generation flagship model, released on November 24, 2025 as claude-opus-4-5-20251101. It is a major generational upgrade over Opus 4.1 and lands at $5 / $25 per million input/output tokens — a 67% price reduction from the prior Opus flagship — with up to 95% additional savings via prompt caching and Batch API. It introduces the effort parameter (low / medium / high) for fine-grained reasoning control and supports autonomous coding sessions of up to 30 minutes.

Opus 4.5's defining breakthrough is achieving harder results with fewer tokens: SWE-Bench Verified 80.9% beats GPT-5.1 (76.3%) and Gemini 3 Pro (76.2%); ARC-AGI-2 37.6% more than doubles GPT-5.1 and exceeds Gemini 3 Pro by ~6 points; on the Artificial Analysis Intelligence Index it scores 70, second only to Gemini 3 Pro. At medium effort, it matches Sonnet 4.5's best SWE-Bench score while using 76% fewer output tokens. The release also introduces Context Compaction, a Claude for Chrome browser extension, and Claude for Excel, and sets industry-best prompt-injection robustness at just 4.7% attack success rate (vs Gemini 3 Pro 12.5% and GPT-5.1 21.9%) — making it the best-aligned frontier model on the market at launch.

Key Capabilities

  • Top-Tier Agentic Coding: #1 on SWE-Bench Verified at 80.9%; outscored every human candidate on Anthropic's internal performance-engineer take-home test.
  • Abstract Reasoning Leap: ARC-AGI-2 37.6% — more than 2× GPT-5.1 and ~6 points above Gemini 3 Pro; Terminal-Bench Hard 44% leads the industry.
  • Token Efficiency Revolution: Matches Sonnet 4.5's top SWE-Bench score at medium effort while using 76% fewer output tokens; customers report 50–75% drops in tool-call and build/lint errors.
  • Long-Horizon Autonomy: Sustains 30-minute autonomous coding sessions; Vending-Bench final balance of $4,967 (+23% over Sonnet 4.5); strong multi-turn logical coherence.
  • Best-in-Class Computer Use: OSWorld 66.26%; Claude for Chrome extension opened to all Claude Max subscribers at launch.
  • Frontier Reasoning Gains: +11 points over Sonnet 4.5 on Humanity's Last Exam, +16 on LiveCodeBench, +12 on τ²-Bench Telecom; MMLU-Pro 90% ties Gemini 3 Pro.
  • Industry-Leading Prompt-Injection Robustness: Combined attack success rate of just 4.7% — substantially lower than Gemini 3 Pro (12.5%) and GPT-5.1 (21.9%).

Technical Strengths

FeatureBenefit
Effort Parameter (low / medium / high)Developers tune reasoning depth per call to trade quality, latency, and cost on the fly
Context CompactionLong-running agentic tasks auto-summarize earlier context, removing the hard token-limit wall
Extreme Token EfficiencySame or better quality with roughly half to a quarter the tokens of peers, dramatically improving cost and throughput
30-Minute Autonomous Coding SessionsSustains goal coherence across long coding tasks with fewer dead-ends, retries, and false-success claims
Industry-Leading Safety Alignment4.7% prompt-injection attack success rate; lowest concerning-behavior rate among frontier models
67% Price ReductionBrings flagship-tier capabilities into reach of smaller teams and high-volume production for the first time

Capability Ratings

DimensionRatingNotes
ReasoningTop-tierAA Intelligence Index 70 (second only to Gemini 3 Pro); leads on ARC-AGI-2 and Terminal-Bench Hard
CodingTop-tier#1 on SWE-Bench Verified at 80.9%; outperformed all human candidates on Anthropic's internal engineering exam
Creative WritingTop-tierAnthropic's signature voice with strong pacing, restraint, and cultural depth
MultimodalStrongMMMU 80.72%; text + image input, no native video/audio
Response SpeedModerateFlagship priced for depth; Low effort balances speed with much lower token usage than peers
Context WindowLarge200K tokens (1M context arrived later in Opus 4.6)

Use Cases

  • Complex Codebase Refactoring & Engineering: #1 on SWE-Bench plus 30-minute autonomous sessions — ideal for long-running migrations, refactors, and modernization.
  • AI Coding Assistants & Claude Code: Already integrated into GitHub Copilot, Cursor, Lovable, and similar tools; particularly strong at planning and code migration.
  • Enterprise Automation & Desktop Agents: Best-in-class computer use plus Claude for Chrome and Excel enable browser and spreadsheet automation.
  • Abstract Reasoning & Research: Large ARC-AGI-2 lead makes it well-suited to ICPC-grade algorithmic problems, physics research, and novel problem-solving.
  • Security-Sensitive Agent Deployments: Industry-leading prompt-injection robustness plus ASL-3 safety standard make it the choice for finance, healthcare, and legal automations.
  • Multi-Step Research & Synthesis: Strong gains on AA-LCR, HLE, and τ²-Bench versus Sonnet 4.5 support deep research and intelligence-synthesis workflows.
  • High-Volume API & Cost-Sensitive Production: 67% price cut plus token efficiency and 95% cache/batch savings put flagship quality within reach of large-scale production.

Pricing

Token TypeLinkAI PriceOfficial Price
input$3.755000 / 1M tokens$5.000000 / 1M tokens
output$18.750000 / 1M tokens$25.000000 / 1M tokens
cache_read$0.375000 / 1M tokens$0.500000 / 1M tokens
cache_write_5m$4.687500 / 1M tokens$6.250000 / 1M tokens
cache_write_1h$7.500000 / 1M tokens$10.000000 / 1M tokens