Anthropic/claude-sonnet-4-6
Anthropic's high-value mid-tier model approaching Opus-level performance on coding, computer use, and agent orchestration, with 1M context and adaptive thinking — built for scaled coding, desktop automation, and enterprise agents.
More from Anthropic
README
Anthropic/claude-sonnet-4-6
Supported Functionality
| Item | Specification |
|---|---|
| Input | Text, Image |
| Output | Text |
| Context | 200,000 tokens (1,000,000 in beta) |
| Max Output | 64,000 tokens |
| Vision | ✓ Supported |
| Function Calling | ✓ Supported (computer use, code execution, memory, MCP) |
Description
Claude Sonnet 4.6 is Anthropic's mid-tier flagship released on February 17, 2026, available via the Claude API as claude-sonnet-4-6 and now the default model on claude.ai and Claude Cowork for Free and Pro users. Pricing holds at $3 / $15 per million input/output tokens — identical to Sonnet 4.5 — with up to 90% savings via prompt caching and 50% via batch processing.
Sonnet 4.6's defining breakthrough is delivering near-Opus performance at Sonnet pricing: 79.6% on SWE-Bench Verified (within 1.2 points of Opus 4.6), 72.5% on OSWorld-Verified (essentially tied with Opus 4.6's 72.7%), and a 4.3× leap on ARC-AGI-2 (13.6% → 58.3%) in a single generation. In Claude Code, 70% of users preferred Sonnet 4.6 over 4.5 and 59% preferred it over the previous flagship Opus 4.5; it actually surpasses Opus 4.6 on real-world office tasks like GDPval-AA (1633 Elo) and Finance Agent (63.3%), while showing substantially improved resistance to prompt-injection attacks.
Key Capabilities
- Real-World Software Engineering: 79.6% on SWE-Bench Verified, 75.9% on SWE-Bench Multilingual; matches Opus 4.5 on long-horizon coding with fewer tokens.
- Computer Use: 72.5% on OSWorld-Verified — autonomously operates Chrome, LibreOffice, VS Code, and other GUI-only legacy software (94% accuracy on insurance vertical tasks).
- Adaptive Thinking: Four effort levels and interleaved reasoning between tool calls allow developers to tune the latency/quality trade-off precisely.
- 1M Long Context + Auto Compaction: Beta 1M-token window holds entire codebases or hundreds of contracts; Context Compaction auto-summarizes older turns to prevent hard truncation.
- Office & Financial Agents: 1633 Elo on GDPval-AA (above Opus 4.6) and 63.3% on Finance Agent — the strongest production model on real-world knowledge work today.
- Abstract Reasoning Leap: ARC-AGI-2 jumped 4.3× from 13.6% to 58.3% in one generation; ARC-AGI-1 reaches 86.5%.
- Mature Developer Ecosystem: Code execution, memory, programmatic tool calling, tool search, and code-filtered web search are all generally available, with native MCP support.
Technical Strengths
| Feature | Benefit |
|---|---|
| Opus-Level Performance at Sonnet Pricing | Most coding, office, and computer-use tasks achieve near-flagship quality without paying 5× Opus pricing |
| 1M Context + Compaction | Whole-repo analysis and very long sessions become practical; old turns compact instead of hard-truncating |
| Frontier Computer Use | OSWorld-Verified rose from 14.9% to 72.5% in 16 months — one of the few models reliably operating real GUIs |
| Adaptive + Interleaved Thinking | Model decides when and how deeply to reason inside agent loops, improving long-horizon reliability |
| Hardened Prompt-Injection Resistance | Major improvement over Sonnet 4.5 in computer-use and web-agent settings; on par with Opus 4.6 |
| Anthropic Safety & Alignment | System card characterizes the model as warm, honest, prosocial with strong safety behaviors and no major misalignment concerns |
Capability Ratings
| Dimension | Rating | Notes |
|---|---|---|
| Reasoning | Excellent | ARC-AGI-2 58.3%, GPQA Diamond 74.1%; trails Opus 4.6 on the hardest scientific reasoning |
| Coding | Top-tier | SWE-Bench Verified 79.6% within 1.2 of Opus 4.6; long-horizon coding matches Opus 4.5 |
| Creative Writing | Excellent | Warm, restrained, well-paced prose — Anthropic's signature style at scale |
| Multimodal | Strong | Text + image input; no native video/audio, optimized for UI screenshots and documents |
| Response Speed | Fast | Mid-tier latency, faster than Opus; thinking-off mode delivers near-instant responses |
| Context Window | Huge | 200K default, 1M in beta; auto-compaction extends effective context further |
Use Cases
- AI Coding Assistants & Claude Code: 70% of users prefer 4.6 over 4.5 — a strong default for Cursor, Claude Code, GitHub Copilot, and similar tools.
- Enterprise Desktop Automation: 72.5% OSWorld and 94% on insurance workflows make it viable for ERP, government portals, and legacy GUI software lacking APIs.
- Large Codebase & Contract Analysis: 1M context plus compaction handles entire repositories, hundreds of contracts, or dozens of research papers in a single session.
- Office & Finance Agents: Leads GDPval-AA and Finance Agent; ideal for automated report writing, financial analysis, and Excel + MCP data workflows.
- Long-Horizon Multi-Step Agents: Adaptive and interleaved thinking deliver strong results on Terminal, browser, and simulated-business tasks like Vending-Bench.
- Front-End UI & Data Reports: Customers report excellent design taste — generates production-quality pages and dashboards with minimal hand-holding.
- Enterprise-Grade Deployment: Available via Claude API, Amazon Bedrock, Google Vertex AI, and Microsoft Foundry with enterprise security and compliance controls.
Pricing
| Token Type | LinkAI Price | Official Price |
|---|---|---|
| input | $2.250000 / 1M tokens | $3.000000 / 1M tokens |
| output | $11.250000 / 1M tokens | $15.000000 / 1M tokens |
| cache_read | $0.225000 / 1M tokens | $0.300000 / 1M tokens |
| cache_write_5m | $2.812500 / 1M tokens | $3.750000 / 1M tokens |
| cache_write_1h | $4.500000 / 1M tokens | $6.000000 / 1M tokens |