Claude Sonnet 5: Model Introduction & Practical Guide
Claude Sonnet 5 is Anthropic's most agentic mid-tier model - and the one most people actually use, since it became the default for Claude Free and Pro users at launch. It plans, calls browser and terminal tools, and runs autonomously, with performance approaching Opus 4.8 in agentic coding, multidisciplinary reasoning, and computer use - at Sonnet-tier pricing.
Here's the short version: Sonnet 5 is the model that made "mid-tier" stop meaning "compromise." SWE-bench Pro lands at 63.2%, Terminal-Bench 2.1 at 80.4%, OSWorld-Verified at 81.2% (near Opus 4.8), Humanity's Last Exam at 43.2% without tools and 57.4% with them, and GDPval-AA v2 at 1618. A five-level effort control (low → max) tunes the cost/quality trade per request, and Anthropic's alignment work reduced hallucination rates, sycophancy, and prompt-injection susceptibility versus Sonnet 4.6. The trade-offs: the updated tokenizer produces 1.0-1.35x more tokens for the same text (changing cost math for cached prompt libraries), the absolute frontier still belongs to Opus 5 and Fable 5, and the browser-search capability, while improved, trails the leaders.
This guide covers model overview, core features, technical specifications, capability comparison, core advantages, recommended use cases, example prompts, and selection recommendations.
Quick Facts
| Attribute | Value |
|---|---|
| Model name | Claude Sonnet 5 |
| Developer | Anthropic |
| Category | Agentic mid-tier model |
| Positioning | Near-Opus-4.8 agentic capability at Sonnet pricing |
| Key benchmarks | SWE-bench Pro 63.2% · Terminal-Bench 2.1 80.4% · OSWorld-Verified 81.2% · HLE 43.2%/57.4% (no tools/tools) · GDPval-AA v2 1618 |
| Effort levels | low / med / high / xhigh / max |
| Default status | Claude Free and Pro default model |
| Tokenizer | Updated - produces ~1.0-1.35x tokens for the same text |
| Caching | 5-minute and 1-hour cache write options |
| Access | Claude web/app, API (claude-sonnet-5), Claude Code, enterprise consoles |
Table of Contents
- Model Overview
- Core Features
- Technical Specifications
- Capability Comparison
- Core Advantages
- Recommended Use Cases
- Example Prompts
- Selection Recommendations
- FAQ
- Sources & Further Reading
1. Model Overview
Anthropic positioned Sonnet 5 as the strongest agentic model in the Sonnet line: it formulates plans, invokes browser and terminal tools, and executes autonomously, with benchmark results approaching Opus 4.8 in exactly the domains that matter for agent work. That capability profile - combined with Sonnet-tier economics - is why it became the default for Free and Pro users immediately.
The technical foundation is adaptive reasoning: rather than a fixed token budget, the model judges task complexity and allocates reasoning accordingly. On top of that sits a five-level effort parameter (low/med/high/xhigh/max) for explicit control over the cost/quality trade. Tool use is native and deep: browser operation for autonomous search and synthesis, terminal invocation for server administration and scripting, and graphical computer control (OSWorld-Verified 81.2%, close to Opus 4.8's level).
Two operational details deserve attention. First, the updated tokenizer produces roughly 1.0-1.35x more tokens for identical text - which matters for teams whose cost models assume historical token counts. Second, caching is flexible (5-minute and 1-hour write options), so prompt libraries with stable prefixes can offset some of that tokenizer drift. Safety improvements are also notable: lower hallucination rates, less sycophancy, and stronger prompt-injection resistance versus Sonnet 4.6 - the failure modes that matter most in autonomous agent deployments.
2. Core Features
Agentic coding. Complex software engineering with autonomous writing and debugging (SWE-bench Pro 63.2%).
Terminal operations. Command execution for server administration and scripting (Terminal-Bench 2.1: 80.4%).
Browser search. Autonomous web search and information synthesis, significantly improved over Sonnet 4.6.
Computer use. Graphical interface operation for complex cross-application tasks (OSWorld-Verified 81.2%).
Multidisciplinary reasoning. HLE 43.2% (no tools) / 57.4% (tools); GDPval-AA v2 1618 for knowledge work.
Five-level effort control. Low through max - explicit per-request cost/quality tuning on top of adaptive reasoning.
Improved safety alignment. Lower hallucination, sycophancy, and prompt-injection risk than Sonnet 4.6.
3. Technical Specifications
| Specification | Detail |
|---|---|
| Reasoning | Adaptive (complexity-based) + 5-level effort parameter |
| Effort levels | low / med / high / xhigh / max |
| Tokenizer | Updated; ~1.0-1.35x tokens vs prior for same text |
| Caching | 5-minute and 1-hour cache write options |
| Vision | High-resolution image understanding |
| Tools | Browser, terminal, computer use - natively integrated |
| Access | Web/app (default Free/Pro), API, Claude Code, enterprise |
Published benchmarks:
| Benchmark | Sonnet 5 |
|---|---|
| SWE-bench Pro | 63.2% |
| Terminal-Bench 2.1 | 80.4% |
| OSWorld-Verified | 81.2% |
| HLE (no tools / tools) | 43.2% / 57.4% |
| GDPval-AA v2 | 1618 |
4. Capability Comparison
| Dimension | Claude Sonnet 5 | Claude Opus 5 | Gemini 2.5 Pro (reference) |
|---|---|---|---|
| Tier | Mid (agentic) | Frontier | Flagship |
| SWE-bench Pro | 63.2% | Higher (Frontier-Bench leader) | ~63-65% (estimates) |
| Terminal-Bench | 80.4% (native depth) | Higher | Tool-call based |
| Computer use | 81.2% (verified) | 70.6% (OSWorld 2.0 variant) | Limited |
| Effort control | 5 levels | 5 levels | Thinking budget |
| Default status | Free/Pro default | Max default | — |
Positioning read. Sonnet 5's value is the proximity to Opus-class agentic performance at the tier below - and its status as the default means most Claude users experience it as "Claude." The caveats are the tokenizer change (watch token budgets on legacy prompts) and the remaining frontier gap on the hardest reasoning and long-horizon work.
5. Core Advantages
- Agentic capability at mid-tier pricing. Near-Opus-4.8 results without the flagship bill.
- Self-directed execution. Plans, calls tools, checks its own output, and follows through to completion.
- Strong terminal and computer use. 80.4% and 81.2% respectively - deployment-ready automation.
- Five-level effort control. Explicit cost/quality tuning per request.
- Improved safety profile. Lower hallucination, sycophancy, and injection risk - critical for autonomy.
- Caching flexibility. 5-minute and 1-hour writes for stable-prefix cost management.
6. Recommended Use Cases
- Software engineering agents: autonomous coding, debugging, and review.
- Server and script automation: terminal-centric operations.
- Browser research: autonomous search and synthesis workflows.
- Computer-use automation: cross-application GUI tasks.
- Knowledge work: reports, analysis, and structured deliverables.
- Default organizational assistant: the tier everyone can use by default.
7. Example Prompts
1. Autonomous coding task
2. Terminal operations
3. Browser research
4. Computer-use workflow
5. Knowledge deliverable
8. Selection Recommendations
Choose Claude Sonnet 5 if:
- You want agentic capability near Opus-class at Sonnet economics.
- Your work spans coding, terminal, browser, and computer use.
- You need per-request effort control for cost management.
- You want the improved safety profile for autonomous deployments.
- You want the default Claude experience (Free/Pro).
Upgrade to Opus 5 if:
- Your tasks demand the frontier (hardest reasoning, longest autonomy).
Consider alternatives if:
- You need open weights.
- Your token budget models depend on historical tokenizer ratios - re-baseline before migrating.
Sources & Further Reading
- Claude Sonnet 5 - AI Toolset (Chinese overview)
- Claude Sonnet 5 announcement - Anthropic
- Claude models - Anthropic
Benchmarks are as published by Anthropic and third-party summaries. Tokenizer changes affect cost calculations - validate against current pricing before budgeting migrations.



