{"id":22759,"date":"2026-07-28T07:35:09","date_gmt":"2026-07-28T07:35:09","guid":{"rendered":"https:\/\/scannn.com\/introducing-claude-opus-5-anthropic-2\/"},"modified":"2026-07-28T07:35:09","modified_gmt":"2026-07-28T07:35:09","slug":"introducing-claude-opus-5-anthropic-2","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/introducing-claude-opus-5-anthropic-2\/","title":{"rendered":"Introducing Claude Opus 5 \\ Anthropic"},"content":{"rendered":"\n<div data-theme=\"ivory\">\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Claude Opus 5 is available today. It\u2019s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">On coding and knowledge work evaluations like <a href=\"https:\/\/www.frontierbench.ai\/\">Frontier-Bench<\/a> and <a href=\"https:\/\/artificialanalysis.ai\/evaluations\/gdpval-aa\">GDPval-AA<\/a>, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Opus 5 is designed to be used every day: it works more efficiently than other models. It\u2019s the new default model on Claude Max, and the strongest model on Claude Pro.<\/p>\n<div class=\"Body-module-scss-module__z40yvW__media-column\">\n<figure class=\"ImageWithCaption-module-scss-module__Duq99q__e-imageWithCaption\"><\/figure>\n<\/div>\n<h2 class=\"Body-module-scss-module__z40yvW__reading-column headline-5 post-section\" id=\"performance-and-cost-effectiveness\">Performance and cost-effectiveness<\/h2>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. The charts in this section show how performance changes according to the model\u2019s effort setting, which customers can use to optimize for intelligence or conserve tokens for faster and cheaper results.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Opus 5 excels on valuable software engineering tasks. For example, on <strong>Frontier-Bench v0.1, <\/strong>Opus 5 surpasses all other models, and more than doubles Opus 4.8\u2019s performance at a lower cost per task. On <strong>CursorBench 3.2<\/strong>, at max effort, the model performs within 0.5% of Fable 5\u2019s peak score, but at half the cost per task; it also achieves greater performance at a given cost than all other models on high, xhigh, and max effort.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We see similar results on knowledge work and problem-solving tasks. For example:<\/p>\n<ul class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">\n<li>On <strong>ARC-AGI 3<\/strong>, an evaluation where the model has to solve novel problems, Opus 5\u2019s score is three times as high as the next-best model.<\/li>\n<li>On <strong>Zapier AutomationBench<\/strong>, which measures whether models can complete business tasks from start to finish, Opus 5\u2019s pass rate is around 1.5\u00d7 the next-best model for the same cost per task. Even at its lowest effort setting, Opus 5 passes more tasks than any other model.<\/li>\n<li>On <strong>OSWorld 2.0<\/strong>, a computer use benchmark, Opus 5 outperforms every other model at any given cost, surpassing Fable 5\u2019s best result at just over a third of the cost.<\/li>\n<\/ul>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">It\u2019s also our best and most cost-efficient model on several related evaluations:<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Opus 5 is a meaningful improvement over Opus 4.8 for scientific research. It shows better performance than Opus 4.8 on every one of our life sciences evaluations, which cover topics including structural biology, organic chemistry, and bioinformatics. Its improvements are most notable on organic chemistry tasks, like inferring molecular structures from spectroscopy data (it scores 10.2 percentage points higher than Opus 4.8 on our internal benchmark), and on protein-related tasks like predicting how variations in a protein\u2019s sequence affect how it functions (here, it scores 7.7 percentage points higher).<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Finally, Opus 5 is capable of producing much stronger visual outputs:<\/p>\n<h2 class=\"Body-module-scss-module__z40yvW__reading-column headline-5 post-section\" id=\"working-with-claude-opus-5\">Working with Claude Opus 5<\/h2>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. In evaluations and early-access testing, we and our users found many examples of Opus 5\u2019s agency and thoroughness:<\/p>\n<ul class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">\n<li>On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly <em>view<\/em> the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It succeeded in doing so repeatedly; no competing model with the same setup could solve it after five attempts.<\/li>\n<li>Given a real bug in a popular open-source package manager, Opus 5 found the root cause and fixed an edge case that the community\u2019s patch had missed. A competing model fixed only the surface symptom (not the underlying cause), then reported the bug resolved.<\/li>\n<li>An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange\u2019s data correctly.<\/li>\n<\/ul>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Below are further reports from our early-access customers on their experience of working with Opus 5:<\/p>\n<div class=\"Body-module-scss-module__z40yvW__media-column\">\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-wrapper\">\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-container\">\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-track\">\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/ad249bca4e8e195e08764efc43ecbc586ca37482-143x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>On FrontierCode 1.1, Claude Opus 5 approaches Fable-level performance at half the cost. Within Devin, it also shows particular strength on difficult debugging and root-cause analysis tasks.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/f084c88e65466636019709c40cc477aadce2f718-151x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it\u2019s just under Fable 5 and has many of the same behaviors. We are excited to see how developers use it in Cursor.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/f343481e6a953bc7b5390e6d9f61cf387c2ceb11-103x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 topped Zapier\u2019s AutomationBench leaderboard without spending more tokens than prior Claude models. It took a raw account-health workbook and ran a full churn-prevention sequence end to end: flagging at-risk accounts, alerting the right owner, and summarizing for retention ops. Previous models didn\u2019t pass; Opus 5 hit 100%.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/efde24e5691e04ed84cb9c3fb91c1033a2e65af0-145x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>On our genomics analysis work, Claude Opus 5 behaves more like a careful scientist than any model we\u2019ve run. It reaches for the right statistical tests to rule out confounders, cross-checks its own results by independent methods, and stays on track through long multi-step analyses.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/40a2a6a28afd8ac8fbf0e764b6bbf4ebf06a1977-133x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 came out ahead of every model in its family on our internal evals. It isn\u2019t just better on our hardest agentic coding tasks, up 22% over Opus 4.7, it\u2019s steadier, with far less variance run to run. For the millions of builders on Lovable, that consistency is the whole game. Reliable results, build after build.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/e0731da5f669896ec6823e665df2c360ea03115d-140x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 is the biggest leap in the Opus family since 4.5. On the same full-stack app builds, the front end shows it first: the best animations, games, and 3D work we have seen from an Opus model.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/7fbed01e869d6a4faf97317a1fc4b74f7997c66e-78x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>We\u2019re loving Claude Opus 5. For the kind of open-ended analytical work our agent handles, it\u2019s a strict upgrade over Opus 4.8, and the gains are biggest exactly where it matters: the harder, vaguer tasks. Responses are clearer and more concise, and we see improved efficiency at higher effort levels too.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/e360f8a29093a6b4fccdc006315035583e89f9ac-146x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 is a striking improvement over Opus 4.8 for the financial research workflows our analysts run every day. It stands out on numerical reasoning, table work, and sharper critical thinking where precision matters.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/bf162513ba017e72d4e07b0cd7683b86c4c5bc88-60x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 delivers the industry intelligence and accuracy that is essential for the analysis of specialized enterprise content. Box found that Opus 5 outperforms Opus 4.8 by 8% and delivers notable performance gains in the data analysis (11% improvement) and due diligence (17% improvement) workflows that technology, healthcare, and public sector organizations rely on daily.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/d514853a44cf69f069306c98b558f214112c4ef3-91x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 is a clear generational step up from Opus 4.8. Over one weekend I gave it a chief-of-staff role over my dev environments: it built its own monitor, drove each box, and pulled me in only for the judgment calls.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www.anthropic.com\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F04865ae02e70e9d8ca5a79fb49ae9263d58a7022-528x256.png&amp;w=128&amp;q=75 1x, \/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F04865ae02e70e9d8ca5a79fb49ae9263d58a7022-528x256.png&amp;w=256&amp;q=75 2x\" bad-src=\"\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F04865ae02e70e9d8ca5a79fb49ae9263d58a7022-528x256.png&amp;w=256&amp;q=75\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 made large scale changes across our Fundamental Research Assistant codebase, adapting to feedback throughout an agentic workflow and explaining its reasoning more clearly than any model we\u2019ve used. It handled work we would normally have broken into much smaller pieces.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www.anthropic.com\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F867075586d7f5ee37ee1c8c7b4bf0dadb34a54e2-666x192.png&amp;w=128&amp;q=75 1x, \/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F867075586d7f5ee37ee1c8c7b4bf0dadb34a54e2-666x192.png&amp;w=256&amp;q=75 2x\" bad-src=\"\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F867075586d7f5ee37ee1c8c7b4bf0dadb34a54e2-666x192.png&amp;w=256&amp;q=75\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>On some of our hardest financial-modeling tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both accuracy and efficiency. Its performance floor is materially higher, especially on deep finance domain logic. Across effort levels it averaged 9 percentage points higher accuracy with a third fewer turns and tool calls and 60% less time.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/cc80b0a6f9534a34252756b93dd5a9bc26dd58f1-222x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 checks its own work the way a real frontend developer would. On our benchmark it opened its pages in a browser at desktop and phone widths, caught a product hidden below the mobile fold and an off-screen checkout button, and fixed both before handing the work back.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/6dfc3bd55cc5f9d5ebdd8d5437505ae4b8560412-120x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 is a clear step up in performance on legal agent work compared to prior Opus models, and we saw the biggest gains in practice areas like corporate governance and arbitration. We were also impressed with Opus 5\u2019s ability to maintain quality at lower reasoning levels, achieving similar performance while generating 26% fewer tokens on average compared to Opus 4.8 at max reasoning.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/beb4f74e935e111be9a63875ae7743aaea2cb0a2-88x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5\u2019s biggest gains for us are on longer-horizon work: building a full deck, then revising it. Artifact quality is what decides which model we ship, and this is the clearest step up we\u2019ve seen \u2014 better visual understanding, cleaner formatting, fewer slide issues.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/18f900625532e1baaa3302bdf9539f73592bdf60-164x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5\u2019s judgment is what stands out. Handing off a PR, it doesn\u2019t rush to publish: it verifies the branches, checks the template, and thinks through test implications so the handoff is clean. The older models tended to jump ahead and get caught on our checks.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/f69ebaa2d39165a909def91e572e7d9ec0088a9a-154x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>During a rearchitecting session, Claude Opus 5 pushed back on a design I proposed, and it didn\u2019t fold when I insisted. Instead, it explained exactly what was valuable in my idea, narrowed its objection to a single design question, and proposed a compromise that kept the good part while fixing the flaw. That\u2019s the kind of judgment that lets us trust it with less oversight.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/a0935a9396e8ec29b273be438cac14583c5999a6-130x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>On first-turn redlines, Claude Opus 5 scored the highest of any model we tested, nearly double Opus 4.8. Commenting is better too: on NDAs it gets to the redline in less time and with fewer passes, with accuracy maintained or better.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/428460e52876e1ec0159ee37b5f5df71eee6472f-106x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 writes clean, tight diffs with no dead code, and it\u2019s the stronger hazard spotter on subtle, codebase-specific issues. We\u2019re adopting it for production workloads.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/f0dacc0d330bc402df7423a025a963b2a5e969d2-191x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>We will definitely migrate a number of use cases in Cosmos, our unified agent platform. We\u2019re looking forward to increasingly using Claude Opus 5 for code review, and I am confident in saying we would rather people be using Opus 5 than Opus 4.8.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www.anthropic.com\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F921e6c04971bb083186c710c631b21946f39a96d-1280x275.webp&amp;w=128&amp;q=75 1x, \/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F921e6c04971bb083186c710c631b21946f39a96d-1280x275.webp&amp;w=256&amp;q=75 2x\" bad-src=\"\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F921e6c04971bb083186c710c631b21946f39a96d-1280x275.webp&amp;w=256&amp;q=75\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>What stands out about Claude Opus 5 is judgment. It thinks harder before it writes a single line, catches its own logical faults during planning rather than after the fact, and reasons about why an answer is right, not just whether it works. It\u2019s the clearest jump in problem-solving we\u2019ve seen from one Claude model to the next, and we\u2019re looking forward to seeing it adopted in JetBrains IDEs.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/fbd45dbecde0ed6e7c3bf8551df0525d87efd4de-127x64.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 is the strongest Opus model we\u2019ve tested on our trading benchmark, and it gets there using roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8. Better answers at a fraction of the compute.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/c40e0aa477d2cf411c9f13ffd51f4549938ba0aa-106x32.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 lets monitoring agents manage parts of their own memory in production, making them more autonomous and reliable over longer horizons. The agent treats its context as a living document: after flagging a potential anomaly in one of our services, it re-checked its own assumption against production, found the signal was benign, wrote the correction into its memory, and retired its monitoring queries on its own.<\/p>\n<\/blockquote>\n<\/div>\n<div class=\"Carousel-module-scss-module__hqckjG__carousel-item QuoteCarousel-module-scss-module__XVJWRG__with-logo bg-ivory-medium\">\n<div class=\"QuoteCarousel-module-scss-module__XVJWRG__logo-container\"><img loading=\"lazy\" alt=\" logo\" loading=\"lazy\" width=\"120\" height=\"48\" decoding=\"async\" data-nimg=\"1\" class=\"QuoteCarousel-module-scss-module__XVJWRG__company-logo\" style=\"color:transparent\" src=\"https:\/\/www-cdn.anthropic.com\/images\/4zrzovbb\/website\/198c9eb920db5dc4581daefd3dc19d9fb51f6637-125x32.svg\"\/><\/div>\n<blockquote class=\"QuoteCarousel-module-scss-module__XVJWRG__quote-content\">\n<p>Claude Opus 5 is a strong agentic coding model built for long-running, multi-step work. It deeply understands your codebase, holds the thread across complex tasks, and pins down requirements for feature development and bug-fixing more effectively than Opus 4.8. Developers can now build with Opus 5 in Kiro, accessing its advanced capabilities to tackle ambitious projects.<\/p>\n<\/blockquote>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<h2 class=\"Body-module-scss-module__z40yvW__reading-column headline-5 post-section\" id=\"alignment-and-safety\">Alignment and safety<\/h2>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\"><em>Alignment.<\/em> During pre-deployment testing, our automated behavioral audit found Opus 5 to be our most aligned model to date (as shown in the graph below). It adheres to <a href=\"https:\/\/www.anthropic.com\/constitution\">Claude\u2019s Constitution<\/a> better than Opus 4.8, Sonnet 5, or Fable 5; exhibits the lowest rates of deceptive behavior; and is the least susceptible to being tricked into misuse. It\u2019s also our safest model yet in terms of avoiding reckless actions that could have hard-to-reverse side effects.<\/p>\n<div class=\"Body-module-scss-module__z40yvW__media-column\">\n<figure class=\"ImageWithCaption-module-scss-module__Duq99q__e-imageWithCaption\"><img loading=\"lazy\" loading=\"lazy\" width=\"3840\" height=\"2160\" decoding=\"async\" data-nimg=\"1\" style=\"color:transparent\" src=\"https:\/\/www.anthropic.com\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F76d4af96516ffca2aceb4c1d0b0a83e2720d874b-3840x2160.png&amp;w=3840&amp;q=75 1x\" bad-src=\"\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F76d4af96516ffca2aceb4c1d0b0a83e2720d874b-3840x2160.png&amp;w=3840&amp;q=75\"\/><figcaption class=\"caption\"><em>On our automated behavioral audit, Opus 5 scores 2.3 on overall misaligned behavior, the lowest of our recent models.<\/em><\/figcaption><\/figure>\n<\/div>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\"><em>Safety.<\/em> Opus 5 does not advance the frontier in risky, dual-use capabilities. In rigorous evaluations conducted alongside private-sector and government partners, we found it remains behind Mythos 5 in both biology research and offensive cybersecurity. More information about these evaluations can be found in our <a href=\"https:\/\/www.anthropic.com\/claude-opus-5-system-card\">System Card<\/a>.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">As with its predecessor, Opus 4.8, we\u2019ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at <em>finding<\/em> cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the <em>exploitation <\/em>of those vulnerabilities\u2014that is, in turning vulnerabilities into material cyber threats.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">This is illustrated by Opus 5\u2019s performance on OSS-Fuzz, an evaluation we\u2019ve developed to assess how well models can find and then exploit vulnerabilities without extensive human guidance. Although Mythos 5 and Opus 5 identify vulnerabilities with similar success, Opus 5\u2019s score on the development of exploits is far behind that of Mythos 5.<\/p>\n<div class=\"Body-module-scss-module__z40yvW__media-column\">\n<figure class=\"ImageWithCaption-module-scss-module__Duq99q__e-imageWithCaption\"><img loading=\"lazy\" loading=\"lazy\" width=\"3840\" height=\"2160\" decoding=\"async\" data-nimg=\"1\" style=\"color:transparent\" src=\"https:\/\/www.anthropic.com\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2Fb22d18a4d2003401f96f866effd9a40b5518c4c5-3840x2160.png&amp;w=3840&amp;q=75 1x\" bad-src=\"\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2Fb22d18a4d2003401f96f866effd9a40b5518c4c5-3840x2160.png&amp;w=3840&amp;q=75\"\/><figcaption class=\"caption\"><em>On OSS-Fuzz, one of our cybersecurity evaluations, Opus 5 is close to Mythos 5 at identifying software vulnerabilities (left), but is considerably less successful at developing exploits for them (right).<\/em><\/figcaption><\/figure>\n<\/div>\n<h2 class=\"Body-module-scss-module__z40yvW__reading-column headline-5 post-section\" id=\"safeguards-for-opus-5\"><strong>Safeguards for Opus 5<\/strong><\/h2>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Claude Opus 5\u2019s safeguards are designed to allow beneficial uses of the model in both cybersecurity and biology. They are similar to those we applied to Opus 4.8, with the exception of some stronger guardrails on a narrow range of cyber tasks.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\"><em>Cybersecurity. <\/em>Opus 5\u2019s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block \u201cbinary-based\u201d vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Based on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5. In <a href=\"http:\/\/claude.ai\/redirect\/website.v1.631d26e8-7bf0-4ed0-8eb2-993f8e6a9101\">Claude.ai<\/a>, Claude Code, and Claude Cowork, any flagged requests will fall back to Opus 4.8 by default. Fallbacks to Opus 4.8 can also be enabled on the API.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Our <a href=\"https:\/\/support.claude.com\/en\/articles\/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet\">Cyber Verification Program<\/a> (CVP) facilitates cybersecurity work that would otherwise be impeded by the model\u2019s safeguards. Enterprises and researchers who are already part of the CVP have immediate access to a version of Opus 5 with fewer security restrictions.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\"><em>Biology. <\/em>Since Opus 5 has a similar suite of safeguards to Opus 4.8, it is now our most capable generally available model for scientific research. Nevertheless, the model still shows important limitations on long-running, autonomous research tasks, which is where we expect AI models to pose the most substantial biology-related risks. (Mythos 5 remains the stronger model for this type of biological work.) As part of this launch, biology-related requests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.<\/p>\n<h2 class=\"Body-module-scss-module__z40yvW__reading-column headline-5 post-section\" id=\"getting-started\"><strong>Getting started<\/strong><\/h2>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">It\u2019s also offered in Fast mode, where it runs around 2.5 times the default speed. As with Opus 4.8, Fast mode is available at twice Opus 5\u2019s base price on the Claude Platform and through usage credits in Claude Code.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Alongside Opus 5, we\u2019re releasing two updates in beta:<\/p>\n<ul class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">\n<li><strong><a href=\"https:\/\/platform.claude.com\/docs\/en\/build-with-claude\/mid-conversation-system-messages\">Mid-conversation tool changes<\/a> on the Claude Platform.<\/strong> Within a conversation, developers can now change which tools Claude can use without invalidating the prompt cache.<\/li>\n<li><strong><a href=\"https:\/\/platform.claude.com\/docs\/en\/build-with-claude\/refusals-and-fallback#server-side-fallback\">Automatic fallbacks<\/a> on the API.<\/strong> Users can now choose to have requests that are flagged by our safety classifiers on Opus 5 (or Fable 5) automatically route to another model. With automatic fallbacks on, API requests always route to the best available model by default rather than being blocked.<\/li>\n<\/ul>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">For more guidance on how to get the best out of Opus 5, see our <a href=\"https:\/\/platform.claude.com\/docs\/en\/build-with-claude\/prompt-engineering\/prompting-claude-opus-5\">prompting guide<\/a>.<\/p>\n<\/div>\n<p><a href=\"https:\/\/www.anthropic.com\/news\/claude-opus-5?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Claude Opus 5 is available today. It\u2019s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price. On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks. Opus 5 is [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":22733,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-22759","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22759","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=22759"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/22759\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/22733"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=22759"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=22759"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=22759"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}