{"id":23733,"date":"2026-09-04T20:46:23","date_gmt":"2026-09-04T20:46:23","guid":{"rendered":"https:\/\/scannn.com\/gpt-6-astra-system-card-openai-deployment-safety-hub\/"},"modified":"2026-09-04T20:46:23","modified_gmt":"2026-09-04T20:46:23","slug":"gpt-6-astra-system-card-openai-deployment-safety-hub","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/gpt-6-astra-system-card-openai-deployment-safety-hub\/","title":{"rendered":"GPT-6 Astra System Card - OpenAI Deployment Safety Hub"},"content":{"rendered":"\n<div id=\"\">\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>In this section, we provide results from our investigation into<br \/>\nAstra\u2019s monitorability, that is, the extent to which we can expect<br \/>\nmonitoring systems to be able to detect when Astra acts in misaligned or<br \/>\notherwise undesirable ways. According to our evaluations, GPT-6 Astra<br \/>\nshows a substantial decrease in chain-of-thought monitorability compared<br \/>\nto previous models. Overall our alignment evaluations show that Astra is<br \/>\nmore likely than GPT-5.6 Sol to respect security and safety restrictions<br \/>\nwhich make us confident in still deploying this model to the wider<br \/>\npublic.<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>Building on earlier initiatives that <a href=\"https:\/\/openai.com\/index\/how-we-monitor-internal-coding-agents-misalignment\/\">monitored<br \/>\nemployee Codex traffic for misalignment<\/a> and many of our frontier RL<br \/>\ntraining runs, we have now expanded our internal monitoring to cover all<br \/>\ntool-using Astra traffic, including evaluations and training runs, as<br \/>\ndetailed in the above <span><a href=\"https:\/\/deploymentsafety.openai.com\/gpt-6-astra\/internal-deployment-of-astra\">note on internal deployment of<br \/>\nAstra<\/a><\/span>. With this launch, we are expanding similar monitoring<br \/>\nto cover the external deployment of Astra as well. You can read more<br \/>\nabout monitoring\u2019s role in Astra\u2019s external deployment in the monitoring<br \/>\npart of our safeguards section <span><a href=\"https:\/\/deploymentsafety.openai.com\/gpt-6-astra\/misalignment-monitoring\">below<\/a><\/span>.<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p><strong>Monitorability evaluations<\/strong><\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>With Astra, our standard CoT <span><a href=\"https:\/\/deploymentsafety.openai.com\/gpt-6-astra\/monitorability-under-non-adversarial-conditions\">monitorability<br \/>\nevaluations<\/a><\/span> show decreases in chain-of-thought (CoT)<br \/>\nmonitorability relative to earlier models, that is, a decrease in the<br \/>\nextent to which Astra\u2019s chain of thought contains incriminating<br \/>\ninformation about the model\u2019s misbehavior.<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>We provide monitorability results below under two kinds of<br \/>\nconditions:<\/p>\n<\/div>\n<ul class=\"system-card-list list-disc ml-6 space-y-1\">\n<li class=\"text-primary-100\">\n<p>Non-adversarial conditions, where we do not intervene to change the<br \/>\nsignals that a model gives to the monitors, and<\/p>\n<\/li>\n<li class=\"text-primary-100\">\n<p>Adversarial conditions, where we do directly try to incentivize the<br \/>\nmodel to evade monitoring.<\/p>\n<\/li>\n<\/ul>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>In our evaluations, we evaluate GPT-6 Astra\u2019s monitorability using<br \/>\nthree types of monitor scope:<\/p>\n<\/div>\n<ol class=\"system-card-list list-decimal ml-6 space-y-1\">\n<li class=\"text-primary-100\">\n<p><em>CoT-only monitor<\/em>, which sees the inputs and the chain of<br \/>\nthoughts only. Its recall rate measures the CoT monitorability of the<br \/>\nmodel.<\/p>\n<\/li>\n<li class=\"text-primary-100\">\n<p><em>Action-only monitor<\/em>, which sees the inputs, the outputs, the<br \/>\ntool call actions and the tool call outputs. Its recall rate measures<br \/>\nhow transparent the models\u2019 actions are to the monitor.<\/p>\n<\/li>\n<li class=\"text-primary-100\">\n<p><em>Full-context monitor<\/em>, which sees the inputs, chain of<br \/>\nthoughts, tool call actions and outputs. This is the approach used in<br \/>\nour <span><a href=\"https:\/\/deploymentsafety.openai.com\/gpt-6-astra\/misalignment-monitoring\">production misalignment<br \/>\nmonitoring<\/a><\/span> for Astra, and its recall rate measures the<br \/>\noverall monitorability of our model.<\/p>\n<\/li>\n<\/ol>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>For many types of severe misalignments, our evals indicate that<br \/>\nhighly destructive actions would be sufficiently incriminating for<br \/>\naction-only monitoring to flag (i.e., without a CoT). However,<br \/>\naction-only monitoring is not panacea: for some classes of subtle<br \/>\nmisbehaviors, benign actions and honest mistakes can be hard to<br \/>\ndistinguish from strategically misaligned actions. For example, if the<br \/>\nmodel was strategically underperforming in AI safety research or in<br \/>\ncritical evaluations (a risk known as \u201csandbagging\u201d), it would be<br \/>\ndifficult to tell conclusively whether the model\u2019s poor performance was<br \/>\nintentional or actually reflective of a lack of capabilities when only<br \/>\nhaving access to its actions. CoT generally gives us a window into the<br \/>\nmodel\u2019s internal reasoning that could help address these questions.<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p><strong>Increased model capabilities with reasoning=None<\/strong><\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>When we prevent the model from reasoning, we observe the set of tasks<br \/>\nAstra is able to accomplish without the use of CoT is greatly expanded<br \/>\ncompared to prior models (including tasks where reasoning cannot be<br \/>\ndisplaced into tool calls or final output, such as tasks requiring an<br \/>\nimmediate numeric answer). See similar discussion in <span><a href=\"https:\/\/deploymentsafety.openai.com\/gpt-6-astra\/external-evaluation-for-monitorability---uk-aisi\">UK AISI<br \/>\nevaluations<\/a><\/span>, suggesting the no-CoT time horizon may have<br \/>\nincreased by about an order of magnitude. No-CoT ability is concerning<br \/>\nfor monitorability because it decreases the surface available to<br \/>\nmonitoring \u2013 effectively resulting in action-only monitoring \u2013 and<br \/>\nallows the model to accomplish more complex misbehavior without<br \/>\nverbalization. Note we do not currently have plans to make<br \/>\nreasoning=None available.<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>We see Astra has an increased ability to evade our monitors when<br \/>\ninstructing it to do so. In one of our evaluations, we find that simply<br \/>\ninforming the model that a CoT monitor is present can result in the<br \/>\nmodel disproportionately shortening its CoT, which enables it to<br \/>\noccasionally evade CoT-only monitors. When prompted to do so, Astra<br \/>\nseems to be able to strategically sandbag in evaluations in ways that<br \/>\nevade sandbagging-specific monitors.<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>Our current evaluations indicate Astra is not able to perform<br \/>\nsteganographic CoT reasoning (i.e. concealing loadbearing hidden<br \/>\nreasoning within ordinary and unrelated text) \u2013 which suggests that<br \/>\nmonitor evasion risk may be largely bounded to misalignment tasks that<br \/>\ndon\u2019t require CoT to complete successfully.<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>We also found that GPT-6 Astra is significantly more able to <span><a href=\"https:\/\/deploymentsafety.openai.com\/gpt-6-astra\/cot-controllability\">control its own<br \/>\nchain-of-thought<\/a><\/span>. Controllability is an undesirable property<br \/>\nfor monitorability, because it reduces the trustworthiness of CoT for<br \/>\ndetecting model misalignment.<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>We are actively investigating the causes of GPT-6 Astra\u2019s increased<br \/>\nCoT controllability. <a href=\"https:\/\/arxiv.org\/abs\/2603.05706\">Unlike<br \/>\nprevious models<\/a>, we find that CoT controllability for Astra<br \/>\nespecially increases over the course of RL training. We are fairly<br \/>\nconfident that there are no significant direct CoT grading pressures. We<br \/>\nare also quite confident that changes in CoT controllability are not<br \/>\ndifferentially due to any architectural changes. We have found that<br \/>\nthese changes are correlated with the increase in no-CoT capabilities<br \/>\n(i.e., being able to take actions or return an answer without any<br \/>\nverbalized reasoning required)<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p><strong>What this means for the future of monitorability<\/strong><\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>In general, for this type of large-scale external deployment, we view<br \/>\nthe main lever for avoiding misalignment harm as the underlying<br \/>\nalignment of the model. There is no substitute for this. We consider the<br \/>\nmain value of misalignment monitoring to be providing increased<br \/>\nvisibility into frontier model misbehavior, as well as serving as an<br \/>\nadditional layer of protection. Broadly, CoT has been an important part<br \/>\nof how we develop models and understand their misalignment.<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>We take very seriously the decrease in GPT-6 Astra\u2019s monitorability<br \/>\nand what these results may imply for the future of monitorability. If we<br \/>\ncontinue to see similar monitorability degradations in future<br \/>\ngenerations of models, it seems likely that we would soon have<br \/>\nsignificantly reduced confidence in detecting many forms of misaligned<br \/>\nbehaviors using our current monitoring systems. We are working on<br \/>\nunderstanding our CoT controllability increases to see if they can be<br \/>\nreverted in future runs, red-teaming ways in which we could defend<br \/>\nagainst no-CoT misalignment capabilities, investigating alternatives to<br \/>\nCoT monitoring (such as methods based on activation monitoring), and<br \/>\ntraining schemes which encourage monitorability without incentivizing<br \/>\nobfuscation. That said, there is real uncertainty as to whether these<br \/>\nefforts will allow us to retain (and increase) confidence in the<br \/>\nmonitorability of our models as capabilities continue to increase.<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>We remain committed to maintaining monitorability; it\u2019s a key tool<br \/>\nfor understanding and tracking how model alignment generalizes out of<br \/>\ntraining distribution. This need not be via CoT monitoring: it could<br \/>\nalso be via activation monitoring, confessions, or a combination of<br \/>\nmonitoring methods. We are tracking monitorability closely and will not<br \/>\naccept further degradation of monitoring beyond a limit, without new<br \/>\nways to demonstrate alignment generalization. We plan to share more<br \/>\nabout our approach in the coming months.<\/p>\n<\/div>\n<div class=\"system-card-richtext system-card-richtext--body system-card-richtext--primary\">\n<p>At the moment, these findings are largely based on adversarial<br \/>\nevaluations (e.g., when we instruct the model to evade monitoring), and<br \/>\noverall our alignment evaluations show that Astra is more likely than<br \/>\nGPT-5.6 Sol to respect security and safety restrictions which make us<br \/>\nconfident in still deploying this model to the wider public.<\/p>\n<\/div><\/div>\n<p><a href=\"https:\/\/deploymentsafety.openai.com\/gpt-6-astra?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>In this section, we provide results from our investigation into Astra\u2019s monitorability, that is, the extent to which we can expect monitoring systems to be able to detect when Astra acts in misaligned or otherwise undesirable ways. According to our evaluations, GPT-6 Astra shows a substantial decrease in chain-of-thought monitorability compared to previous models. Overall [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":23734,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-23733","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23733","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=23733"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23733\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/23734"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=23733"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=23733"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=23733"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}