Anthropic/claude-opus-4-5-20251101
Anthropic's 4th-gen flagship setting state-of-the-art on SWE-Bench Verified and ARC-AGI-2, introducing effort control and context compaction, with industry-leading prompt-injection robustness — built for coding, reasoning, and enterprise agents.
More from Anthropic
README
Anthropic/claude-opus-4-5-20251101
Supported Functionality
| Item | Specification |
|---|---|
| Input | Text, Image |
| Output | Text |
| Context | 200,000 tokens |
| Max Output | 64,000 tokens |
| Vision | ✓ Supported |
| Function Calling | ✓ Supported (computer use, context compaction, MCP) |
Description
Claude Opus 4.5 is Anthropic's 4th-generation flagship model, released on November 24, 2025 as claude-opus-4-5-20251101. It is a major generational upgrade over Opus 4.1 and lands at $5 / $25 per million input/output tokens — a 67% price reduction from the prior Opus flagship — with up to 95% additional savings via prompt caching and Batch API. It introduces the effort parameter (low / medium / high) for fine-grained reasoning control and supports autonomous coding sessions of up to 30 minutes.
Opus 4.5's defining breakthrough is achieving harder results with fewer tokens: SWE-Bench Verified 80.9% beats GPT-5.1 (76.3%) and Gemini 3 Pro (76.2%); ARC-AGI-2 37.6% more than doubles GPT-5.1 and exceeds Gemini 3 Pro by ~6 points; on the Artificial Analysis Intelligence Index it scores 70, second only to Gemini 3 Pro. At medium effort, it matches Sonnet 4.5's best SWE-Bench score while using 76% fewer output tokens. The release also introduces Context Compaction, a Claude for Chrome browser extension, and Claude for Excel, and sets industry-best prompt-injection robustness at just 4.7% attack success rate (vs Gemini 3 Pro 12.5% and GPT-5.1 21.9%) — making it the best-aligned frontier model on the market at launch.
Key Capabilities
- Top-Tier Agentic Coding: #1 on SWE-Bench Verified at 80.9%; outscored every human candidate on Anthropic's internal performance-engineer take-home test.
- Abstract Reasoning Leap: ARC-AGI-2 37.6% — more than 2× GPT-5.1 and ~6 points above Gemini 3 Pro; Terminal-Bench Hard 44% leads the industry.
- Token Efficiency Revolution: Matches Sonnet 4.5's top SWE-Bench score at medium effort while using 76% fewer output tokens; customers report 50–75% drops in tool-call and build/lint errors.
- Long-Horizon Autonomy: Sustains 30-minute autonomous coding sessions; Vending-Bench final balance of $4,967 (+23% over Sonnet 4.5); strong multi-turn logical coherence.
- Best-in-Class Computer Use: OSWorld 66.26%; Claude for Chrome extension opened to all Claude Max subscribers at launch.
- Frontier Reasoning Gains: +11 points over Sonnet 4.5 on Humanity's Last Exam, +16 on LiveCodeBench, +12 on τ²-Bench Telecom; MMLU-Pro 90% ties Gemini 3 Pro.
- Industry-Leading Prompt-Injection Robustness: Combined attack success rate of just 4.7% — substantially lower than Gemini 3 Pro (12.5%) and GPT-5.1 (21.9%).
Technical Strengths
| Feature | Benefit |
|---|---|
| Effort Parameter (low / medium / high) | Developers tune reasoning depth per call to trade quality, latency, and cost on the fly |
| Context Compaction | Long-running agentic tasks auto-summarize earlier context, removing the hard token-limit wall |
| Extreme Token Efficiency | Same or better quality with roughly half to a quarter the tokens of peers, dramatically improving cost and throughput |
| 30-Minute Autonomous Coding Sessions | Sustains goal coherence across long coding tasks with fewer dead-ends, retries, and false-success claims |
| Industry-Leading Safety Alignment | 4.7% prompt-injection attack success rate; lowest concerning-behavior rate among frontier models |
| 67% Price Reduction | Brings flagship-tier capabilities into reach of smaller teams and high-volume production for the first time |
Capability Ratings
| Dimension | Rating | Notes |
|---|---|---|
| Reasoning | Top-tier | AA Intelligence Index 70 (second only to Gemini 3 Pro); leads on ARC-AGI-2 and Terminal-Bench Hard |
| Coding | Top-tier | #1 on SWE-Bench Verified at 80.9%; outperformed all human candidates on Anthropic's internal engineering exam |
| Creative Writing | Top-tier | Anthropic's signature voice with strong pacing, restraint, and cultural depth |
| Multimodal | Strong | MMMU 80.72%; text + image input, no native video/audio |
| Response Speed | Moderate | Flagship priced for depth; Low effort balances speed with much lower token usage than peers |
| Context Window | Large | 200K tokens (1M context arrived later in Opus 4.6) |
Use Cases
- Complex Codebase Refactoring & Engineering: #1 on SWE-Bench plus 30-minute autonomous sessions — ideal for long-running migrations, refactors, and modernization.
- AI Coding Assistants & Claude Code: Already integrated into GitHub Copilot, Cursor, Lovable, and similar tools; particularly strong at planning and code migration.
- Enterprise Automation & Desktop Agents: Best-in-class computer use plus Claude for Chrome and Excel enable browser and spreadsheet automation.
- Abstract Reasoning & Research: Large ARC-AGI-2 lead makes it well-suited to ICPC-grade algorithmic problems, physics research, and novel problem-solving.
- Security-Sensitive Agent Deployments: Industry-leading prompt-injection robustness plus ASL-3 safety standard make it the choice for finance, healthcare, and legal automations.
- Multi-Step Research & Synthesis: Strong gains on AA-LCR, HLE, and τ²-Bench versus Sonnet 4.5 support deep research and intelligence-synthesis workflows.
- High-Volume API & Cost-Sensitive Production: 67% price cut plus token efficiency and 95% cache/batch savings put flagship quality within reach of large-scale production.
Pricing
| Token Type | LinkAI Price | Official Price |
|---|---|---|
| input | $3.755000 / 1M tokens | $5.000000 / 1M tokens |
| output | $18.750000 / 1M tokens | $25.000000 / 1M tokens |
| cache_read | $0.375000 / 1M tokens | $0.500000 / 1M tokens |
| cache_write_5m | $4.687500 / 1M tokens | $6.250000 / 1M tokens |
| cache_write_1h | $7.500000 / 1M tokens | $10.000000 / 1M tokens |