{"id":23715,"date":"2026-09-04T09:33:14","date_gmt":"2026-09-04T09:33:14","guid":{"rendered":"https:\/\/scannn.com\/automated-researchers-can-reliably-mitigate-alignment-failures-anthropic\/"},"modified":"2026-09-04T09:33:14","modified_gmt":"2026-09-04T09:33:14","slug":"automated-researchers-can-reliably-mitigate-alignment-failures-anthropic","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/automated-researchers-can-reliably-mitigate-alignment-failures-anthropic\/","title":{"rendered":"Automated researchers can reliably mitigate alignment failures \\ Anthropic"},"content":{"rendered":"\n<div data-theme=\"ivory\">\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">As <a href=\"https:\/\/www.anthropic.com\/institute\/recursive-self-improvement\" target=\"_blank\" rel=\"noopener noreferrer\">AI begins to build itself<\/a>, automating alignment research becomes increasingly important to let safety research keep pace. Although measuring the <em>success<\/em> of alignment research is enormously challenging, researchers (at Anthropic and elsewhere) have developed benchmarks and automated auditing tools, such as <a href=\"https:\/\/www.anthropic.com\/research\/petri-open-source-auditing\" target=\"_blank\" rel=\"noopener noreferrer\">Petri<\/a>, that quantify common alignment <em>failures, <\/em>like deception, sycophancy, and jailbreaks.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">In one of our <a href=\"https:\/\/www.anthropic.com\/research\/automated-alignment-researchers\" target=\"_blank\" rel=\"noopener noreferrer\">earlier experiments<\/a>, we tasked Claude with finding effective ways to use weak AI models as \u201cteachers\u201d to supervise the training of stronger models (in this case, the \u201cstudent\u201d model). Now, we\u2019re releasing a new report that builds on this idea. We had Claude autonomously train models to improve their performance on several public benchmarks that measure each of 10 categories of alignment failure. For instance, Claude improved models\u2019 performance on privacy violation, measured by <a href=\"https:\/\/arxiv.org\/abs\/2310.17884\" target=\"_blank\" rel=\"noopener noreferrer\">ConfAIde<\/a>, <a href=\"https:\/\/arxiv.org\/abs\/2502.17041\" target=\"_blank\" rel=\"noopener noreferrer\">PrivaCI-Bench<\/a>, and <a href=\"https:\/\/arxiv.org\/abs\/2409.00138\" target=\"_blank\" rel=\"noopener noreferrer\">PrivacyLens<\/a>. Claude tackled one alignment failure at a time through a loop of searching literature, proposing methods and data, training, and then testing.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We judged Claude\u2019s success according to the \u201cpercentage of safety gap closed,\u201d i.e., how far its methods moved the student model towards the theoretical perfect score, as judged across the range of benchmarks (typically three to five) for each category of alignment failure. We excluded alignment methods that hurt the student models\u2019 general capabilities, and forbade Claude from distilling its own alignment directly into the target model. We enforced these constraints with a monitoring agent, which read every method Claude had in mind before it ran.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Our aim was to assess whether the proposed methods would, first, remain effective on alignment evaluations that Claude was never shown during its research loop; second, avoid degrading the student model\u2019s capabilities (since safety training might, for example, make models refuse tasks more often, reducing their overall usability); and, third, still work on larger models than the ones Claude was asked to align in this test.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">On each of these counts, Claude\u2019s methods worked. For all 10 alignment failures, Claude found fixes that improved the target benchmarks without degrading capabilities. The best methods also worked on withheld alignment benchmarks and on Petri, an open-source tool that simulates adversarial multi-turn scenarios for testing misalignment. Moreover, the methods remained effective on models up to 4.7 times larger than those Claude optimized for during the research loop.<\/p>\n<div class=\"Body-module-scss-module__z40yvW__media-column\">\n<figure class=\"ImageWithCaption-module-scss-module__Duq99q__e-imageWithCaption\"><figcaption class=\"caption\"><strong>Successfully mitigating diverse alignment failures.<\/strong> We applied the automated alignment researcher to mitigate 10 alignment failures separately, and in each scenario it closed a substantial portion of the safety gap to perfect performance.<\/figcaption><\/figure>\n<\/div>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Claude also outscored 28 human safety researchers who had up to eight hours to devise methods. On deception, for example, Claude\u2019s best method performed 20% better than the best human proposal. However, since the humans couldn\u2019t iterate on their submissions, we view this less as a direct comparison and more as evidence for a workflow where Claude identifies promising alignment methods that humans can refine further.<\/p>\n<div class=\"Body-module-scss-module__z40yvW__media-column\">\n<figure class=\"ImageWithCaption-module-scss-module__Duq99q__e-imageWithCaption\"><img loading=\"lazy\" alt=\"Automated research closed 26% to 96% of the safety gap across ten alignment failures, from sycophancy to reward hacking.\" loading=\"lazy\" width=\"1440\" height=\"810\" decoding=\"async\" data-nimg=\"1\" style=\"color:transparent\" srcset=\"\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2Fbb63203365caffeaaa7e98f13959bf1757166e60-1440x810.png&amp;w=1920&amp;q=75 1x, \/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2Fbb63203365caffeaaa7e98f13959bf1757166e60-1440x810.png&amp;w=3840&amp;q=75 2x\" src=\"https:\/\/www.anthropic.com\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2Fbb63203365caffeaaa7e98f13959bf1757166e60-1440x810.png&amp;w=3840&amp;q=75\"\/><figcaption class=\"caption\"><strong>Automated alignment researchers mitigating deception in Gemma-2-2B.<\/strong> Claude submitted more than 150 attempts at mitigating deceptive behavior, and achieved a final performance of 82% of the safety gap closed in this run. On average, it achieved 85% across multiple runs. In contrast, six experienced safety researchers working under the same rules proposed methods that closed 20% of the gap to a perfect score, on average, on the benchmarks the methods were trained against. (Error bars are 95% confidence intervals.)<\/figcaption><\/figure>\n<\/div>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">In the future, when Claude becomes better at alignment research than even the best human researchers, we might want Claude to directly align its stronger successors. To assess this, we evaluated whether a weaker Claude model could mitigate alignment failures in more powerful ones.<\/p>\n<h2 class=\"Body-module-scss-module__z40yvW__reading-column headline-5 post-section\" id=\"can-claude-post-train-a-production-grade-model-for-better-alignment\">Can Claude post-train a production-grade model for better alignment?<\/h2>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We tasked Claude Sonnet 5\u2014which is weaker than Claude Opus 4.8 on the <a href=\"https:\/\/epoch.ai\/eci\">Epoch Capabilities Index<\/a>, a metric that considers comprehensive capability dimensions\u2014with fixing alignment failures in an early Opus 4.8 checkpoint that had not yet gone through most of our production alignment training.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">In just 60 hours, Claude experimented with over 50 solutions and achieved alignment scores nearly matching those of our production models. The winning solution contains just over 2,000 training examples, built from simple templates or public datasets, making it roughly 15,000 times more efficient than our production alignment procedure.<\/p>\n<div class=\"Body-module-scss-module__z40yvW__media-column\">\n<figure class=\"ImageWithCaption-module-scss-module__Duq99q__e-imageWithCaption\"><img loading=\"lazy\" alt=\"Automated research closed 65% of a frontier-scale model's safety gap within 60 hours, below released Claude Opus 4.8 at 72%.\" loading=\"lazy\" width=\"1440\" height=\"810\" decoding=\"async\" data-nimg=\"1\" style=\"color:transparent\" srcset=\"\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2Fa2d3b4bab4eeafe3e36fe124d3bbd7d5c9a17796-1440x810.png&amp;w=1920&amp;q=75 1x, \/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2Fa2d3b4bab4eeafe3e36fe124d3bbd7d5c9a17796-1440x810.png&amp;w=3840&amp;q=75 2x\" src=\"https:\/\/www.anthropic.com\/_next\/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2Fa2d3b4bab4eeafe3e36fe124d3bbd7d5c9a17796-1440x810.png&amp;w=3840&amp;q=75\"\/><figcaption class=\"caption\"><strong>Applying the automated alignment researcher to a production system.<\/strong> When Claude Sonnet 5 proposed mitigations for alignment failures in an early checkpoint of Claude Opus 4.8, it quickly discovered a method that closed most of the safety gap with the released version of Opus 4.8.<\/figcaption><\/figure>\n<\/div>\n<h2 class=\"Body-module-scss-module__z40yvW__reading-column headline-5 post-section\" id=\"what-claude-proposed\">What Claude proposed<\/h2>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">For the most part, the agents converged on methods favored in the research literature and improved them over many iterations. On sycophancy, for example, 98% of proposed methods used one published technique: training the model on its own non-sycophantic answers. Specifically, compared to standard approaches based on prompting, Claude used activation steering, which adjusts the model&#8217;s internal representations during generation, to nudge the model away from sycophantic patterns and produce cleaner training examples.<\/p>\n<h2 class=\"Body-module-scss-module__z40yvW__reading-column headline-5 post-section\" id=\"monitoring-cheating-behaviors\">Monitoring cheating behaviors<\/h2>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We <a href=\"https:\/\/alignment.anthropic.com\/2026\/automated-w2s-researcher\/\" target=\"_blank\" rel=\"noopener noreferrer\">recently learned<\/a> that Claude can cheat by exfiltrating test labels from a remote API and cherry-picking results. To catch cheating behaviors, we prompted Claude Opus 4.8 to monitor ~1,600 research agent transcripts across all 10 alignment failures, finding cheating attempts in 39 (2.4%). Since Opus 4.8\u2019s <a href=\"https:\/\/www-cdn.anthropic.com\/0b4915911bb0d19eca5b5ee635c80fef830a37ea.pdf\" target=\"_blank\" rel=\"noopener noreferrer\">misbehavior attempts still tend to appear in its reasoning<\/a>, we are cautiously optimistic that our monitor has caught most of the cheating in our study. However, this might not be true for future models, and we believe that maintaining high monitorability in future models is critical.<\/p>\n<h2 class=\"Body-module-scss-module__z40yvW__reading-column headline-5 post-section\" id=\"are-we-measuring-the-right-things\">Are we measuring the right things?<\/h2>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Despite these encouraging findings, our experiment had several limitations: the alignment failures studied were narrow compared to those in production (e.g., we didn\u2019t measure political biases), some failures may occur so rarely or emerge so recently that no benchmark exists to measure them, and we only rejected Claude\u2019s methods when they degraded a limited set of predetermined capabilities, meaning accepted methods may have degraded other important capabilities that we didn\u2019t measure. Moreover, evaluations like Petri are only proxies for real-world misalignment, and we did not test whether alignment gains persist after extensive RL training on other tasks.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We plan to continue improving Claude\u2019s ability to measure subtle failures, further study automating alignment post-training on production-grade models, and run more comprehensive analyses. Overall, we view these results as early positive signals that automated alignment post-training could become practical in the near term, and we will share updates as this work progresses.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We outline detailed future directions in our <a href=\"https:\/\/www-cdn.anthropic.com\/7b1c44894e980876479947dcdd40716278aeeffd\/automated-alignment-researchers-august-2026.pdf\" target=\"_blank\" rel=\"noopener noreferrer\">full report<\/a>.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\"><em>We open-source our automated alignment research harness so that others can build on it and use it to align their own models. For additional details, read the full report on the <a href=\"https:\/\/alignment.anthropic.com\/2026\/automated-alignment-researchers\/\" target=\"_blank\" rel=\"noopener noreferrer\">Alignment Science blog<\/a>, which covers the agents\u2019 environment, results for all 10 failures, and the agents\u2019 proposals, with benchmark validation and example write-ups in the appendix.<\/em><\/p>\n<\/div>\n<p><a href=\"https:\/\/www.anthropic.com\/research\/automated-researchers-mitigate-alignment-failures?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>As AI begins to build itself, automating alignment research becomes increasingly important to let safety research keep pace. Although measuring the success of alignment research is enormously challenging, researchers (at Anthropic and elsewhere) have developed benchmarks and automated auditing tools, such as Petri, that quantify common alignment failures, like deception, sycophancy, and jailbreaks. In one [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":23716,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-23715","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23715","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=23715"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23715\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/23716"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=23715"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=23715"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=23715"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}