Claude Opus 4.7: Anthropic Reclaims the Coding Crown While Deliberately Nerfing Cyber Capabilities
All blog articles
A 64.3% score on SWE-bench Pro, up from 53.4%. Anthropic just dropped Claude Opus 4.7 today, and the numbers tell a clear story. The new flagship model beats both OpenAI's GPT-5.4 and Google's Gemini 3.1 Pro across several key benchmarks.
Claude Opus 4.7 leaps forward in coding and vision
Opus 4.7 outperforms its predecessor in agentic coding, multidisciplinary reasoning, and scaled tool use. The model now verifies its own outputs before reporting back, which means you can hand off difficult tasks with less babysitting.
Vision gets a massive upgrade too. Maximum resolution jumps from 1.15 megapixels to roughly 3.75 megapixels. Pixel coordinates now map 1:1 with actual screenshots, a huge win for autonomous computer use.
Cyber capabilities deliberately throttled
Here's where it gets really interesting. Anthropic intentionally reduced Opus 4.7's cybersecurity capabilities during training. Built-in safeguards automatically detect and block high-risk cyber requests. This caution flows directly from Project Glasswing and the restricted Claude Mythos model, which remains available only to select partners like Apple, Google, and Microsoft.
Security professionals can apply for access through a new Cyber Verification Program. Think of it as a preview of a future where the most powerful AI features sit behind professional credentials.
Same price per token, but your bill might go up
Pricing holds steady at $5 per million input tokens and $25 per million output tokens. However, a new tokenizer can map the same text to up to 35% more tokens. So yes, your per-request costs could climb noticeably.
A new "xhigh" effort level slots between high and max for finer control. Claude Code also gains an /ultrareview command for thorough code reviews. Opus 4.7 is already rolling out on GitHub Copilot, plus the Claude API, Amazon Bedrock, Vertex AI, and Microsoft Foundry.
Anthropic leads, but the gap is razor-thin
On comparable benchmarks, Opus 4.7 only edges out GPT-5.4 by a slim margin. OpenAI still wins on agentic search (89.3% vs. 79.3%). And the real beast remains Mythos Preview at 77.8% on SWE-bench Pro. The question for you developers: how long before Anthropic unlocks Mythos for everyone?