Claude Opus 4.8 is released: code defects are reduced by 75%, Agentic Coding fully surpasses GPT-5.5

On May 29, Anthropic officially released its latest flagship model Claude Opus 4.8.

The most eye-catching data comes from the field of programming: in the Agentic Coding benchmark, Opus 4.8 achieved 69.2% The high score is not only significantly ahead of the 64.3% of the previous generation Opus 4.7, but also behind OpenAI's GPT-5.5 (58.6%) and Google Gemini 3.1 Pro (54.2%).Reduced by about 75%——For developers, this means less debugging time and higher quality of delivery.

Not only stronger, but also more "honest" AI

If the performance improvement is an expected improvement, then Opus 4.8's breakthrough in "honesty" is even more worthy of attention.Opus 4.8 is more likely to proactively point out uncertainties in its analysis than to give answers that seem fluent but are actually wrong.

The testing team of Bridgewater Associates, the world's largest hedge fund, gave a pithy summary: "The biggest difference of Opus 4.8 is that it actively flags potential problems in input data and output analysis - these are usually ignored by other models and left to users to discover by themselves." In scenarios such as financial analysis, where the error tolerance rate is extremely low, this ability to "dare to admit uncertainty" is more valuable than pure fluency.

Anthropic's evaluation also shows that Opus 4.8's performance in deception rate and coordination abuse is "significantly lower" than the previous model, even tying the model that was previously praised as "the best aligned"Claude Mythos Preview level.

Dynamic Workflows: Multi-agent collaboration preview

Along with Opus 4.8, there is also a software called Dynamic Workflows Research preview function. Claude Code is available in Team, Enterprise, and Max plans.

For teams that need to handle long-term, multi-step tasks, Dynamic Workflows means upgrading from "commanding an AI to work" to "commanding a group of AIs to work collaboratively."

Price remains the same, speed doubles

In terms of pricing strategy, Anthropic gave a surprise this time: the standard API price of Opus 4.8 is exactly the same as that of Opus 4.7 - $5 per million tokens for input and $25 per million tokens for output. Fast Mode Running at 2.5 times the output speed per minute, the price is only $10/$50 (per million tokens), which is about two-thirds cheaper than the fast inference of the previous generation model.

The available range is also available in one step: in addition to claude.ai and Claude API, Opus 4.8 has been launched simultaneously on Amazon Bedrock, Google Vertex AI, Microsoft Foundry, and was immediately adopted GitHub Copilot and Cursor integrated.

What does this mean for users?

If you are a developer, what can Opus 4.8 do for you?Write fewer bugs, debug less, and dare to hand over complex tasks to AI.

If you don’t write code, but need AI to help you analyze data, write reports, or perform complex reasoning, Opus 4.8’s “active reminder of uncertainty” feature is also of great significance—you will no longer be misled by an answer that seems to be sound but actually contains errors.

Summarize

Claude Opus 4.8 is a "substantial upgrade without a price increase" - stronger programming capabilities, more honest answers, faster reasoning, and future-oriented multi-agent collaboration capabilities.

To discover more useful AI tools, welcome to visit AI Dash.

🔗 Share: Twitter Weibo Copy link

📬 Like this article?

Weekly selected AI tool reviews + practical tutorials, delivered directly to you.

Subscribe to the weekly AI picks →

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top