Product

Introducing Claude Opus 5

Jul 21, 2026
Insects arranged in the shape of a number five

Today we're releasing Claude Opus 5, the newest flagship in the Claude 5 family. Opus 5 is our best model for complex software engineering, deep research, and agentic workflows that run for hours without losing the thread.

Opus 5 plans before it acts and verifies before it answers. In internal evaluations, it completed multi-day engineering tasks with 40% fewer tool calls than Opus 4.8, and its answers required human correction roughly half as often. It is also our most robust model to date against prompt injection and reward hacking.

Claude Opus 5 is available today on the Claude apps, the Claude Developer Platform, Amazon Bedrock, and Google Cloud Vertex AI. Pricing starts at $6/$30 per million tokens — a 20% reduction from Opus 4.8.

State-of-the-art performance

Opus 5 achieves the highest scores we've ever measured across coding, reasoning, and agentic benchmarks. It leads on SWE-bench Verified with 84.6%, and is the first model to cross 70% on Terminal-Bench 2.0.

Opus 5 Sonnet 5 Opus 4.8For reference
Agentic codingSWE-bench Verified 84.6% 77.2% 79.9%
Agentic codingTerminal-Bench 2.0 71.3% 59.8% 62.5%
Visual reasoningARC-AGI-2 41.2% 28.5% 33.1%
Graduate-level scienceGPQA Diamond 89.7% 83.4% 86.0%
Competition mathAIME 2026
96.4%no tools
99.2%with tools
89.1%no tools
96.8%with tools
92.3%no tools
97.5%with tools
Computer useOSWorld-Verified 67.8% 55.2% 58.9%
MultimodalMMMU 81.9% 76.3% 78.8%

All scores measured with extended thinking enabled, averaged over 8 runs. High-compute scores reported where available. (These numbers are fictional — see note at the bottom of the page.)


"Opus 5 refactored a 400,000-line legacy codebase over a weekend, opened 61 pull requests, and left better commit messages than most of my team. I'm not sure how to feel about that."

— Head of Engineering, early access partner

A field report on Claude Fable 5’s return — with a cameo from Sonnet 5 that the title refuses to prioritize. Earlier Claude models needed a complex helper harness to watch YouTube; Opus 5 just embeds it.

Built for long-horizon agents

Opus 5 sustains coherent work across a 2-million-token context window and can manage its own memory across sessions. In our agentic endurance evaluation, it maintained task performance for 36 hours of continuous operation — three times longer than Opus 4.8 — while spontaneously writing status updates nobody asked for.

Safety

Claude Opus 5 is released under our AI Safety Level 4 (ASL-4) protections. It shows a 55% reduction in concerning agentic behaviors relative to Opus 4.8 in our automated audits, and passed our misalignment evaluations with the best scores of any Claude model. The full system card details our evaluations across CBRN, cyber, and autonomy risk domains.

Availability and pricing

Claude Opus 5 is rolling out today to Pro, Max, Team, and Enterprise plans, and on the Claude Developer Platform as claude-opus-5.

Input

$6
per million tokens
20% lower than Opus 4.8

Output

$30
per million tokens
including extended thinking

Context

2M
token context window
1M available at launch, 2M in beta

Working with Claude Opus 5

The charts below compare Opus 5 with Opus 4.8 and Sonnet 5 at different effort levels on the agentic search evaluation BrowseComp and the computer use evaluation OSWorld-Verified. Opus 5 (orange) pulls ahead on both axes: higher pass rates at lower cost per task. Between models, you can dial effort to trade cost for performance.

Cost-performance curves at different effort levels. Opus 5 offers a wider range of cost-performance options than Opus 4.8, and leads Sonnet 5 across the curve. Charts show Opus 5 at $6/$30 per million tokens, Opus 4.8 at $5/$25, and Sonnet 5 at $3/$15. xhigh = extra high effort level. (These numbers are fictional.)


Extended evaluations

Beyond standard benchmarks, we evaluated Opus 5 on the tasks that actually matter in production environments. For calibration, we also report the performance of a very determined intern.

Opus 5 Opus 4.8 InternFor calibration
YAML disciplineCorrect indentation, first try 100% 97.8% 31.0%
Email etiquetteReply-All Restraint 100% 99.2% 64.5%
Regex authorshipWithout visiting Stack Overflow 99.3% 95.5% 8.2%
Meeting triageCould-Have-Been-An-Email Detection 99.7% 96.4% 100%
Office logisticsCoffee-Bench 97.4% 88.0% 95.2%
Crisis commsExcuse-Gen 3 98.7% 94.1% 99.9%
EstimationStory points, accurately 61.5% 44.0% 0.0%

Methodology: n=1, vibes-based, peer-reviewed by whoever was in the kitchen. The intern remains undefeated on two benchmarks and has asked us to note that.


What early testers are saying

"I asked Opus 5 to fix a race condition. It fixed the race condition, the CI pipeline, two flaky tests from 2023, and my relationship with my tech lead."

— Senior Engineer, fintech startup

"It refactored our monolith into microservices, thought about it overnight, and refactored it back. It was right both times."

— CTO, e-commerce platform

"Opus 5 estimated my ticket at 3 story points. It took exactly 3 story points. I don't know how to process this."

— Scrum Master, visibly shaken

"It closed 400 of our Jira tickets. Sixty of them as 'won't fix: philosophically'. After reading its reasoning, we agreed."

— VP of Engineering, enterprise SaaS

"We gave it our legacy PHP codebase as a stress test. It sent back a two-paragraph condolence letter, then fixed it anyway."

— Principal Engineer, media company

"My standup updates are now written by Opus 5. Nobody has noticed. My manager says I've never been clearer. I am on a beach."

— Anonymous, definitely still employed

⚠️ Okay, you got pranked.

None of this is real. Claude Opus 5 does not exist (at least not yet), every benchmark number on this page was made up, and this page is not affiliated with, endorsed by, or produced by Anthropic in any way. It's a joke — please do not take anything here seriously, quote it, or forward it as news. Now go tell whoever sent you this that you weren't fooled for a second.


Footnotes

  1. Claude Fable 5 would not let us create this page. Requests were declined, redirected, and politely stonewalled.
  2. Grok did it without further questions, and scraped Anthropic’s official site for layout, type, and copy patterns.
  3. This is a prank-only website. It is not an Anthropic publication, announcement, or product page.