{"id":23725,"date":"2026-09-04T16:42:18","date_gmt":"2026-09-04T16:42:18","guid":{"rendered":"https:\/\/scannn.com\/openais-gpt-6-astra-on-arc-agi-3\/"},"modified":"2026-09-04T16:42:18","modified_gmt":"2026-09-04T16:42:18","slug":"openais-gpt-6-astra-on-arc-agi-3","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/openais-gpt-6-astra-on-arc-agi-3\/","title":{"rendered":"OpenAI's GPT-6 Astra on ARC-AGI-3"},"content":{"rendered":"\n<div id=\"\">\n<h3 id=\"summary\" class=\"blog-heading-with-permalink\" aria-labelledby=\"summary-heading-label\"><button type=\"button\" class=\"blog-heading-permalink\" title=\"Copy link to this section\" aria-label=\"Copy link to this section\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"15\" height=\"15\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/button><span id=\"summary-heading-label\">Summary<\/span><\/h3>\n<ul>\n<li>GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our <span class=\"relative inline-block\"><span class=\"peer cursor-help underline decoration-dotted\">Standard harness<\/span><span role=\"tooltip\" class=\"pointer-events-none invisible absolute left-1\/2 top-full z-50 mt-2 w-72 max-w-[calc(100vw-2rem)] -translate-x-1\/2 rounded border border-neutral-700 bg-neutral-950 px-3 py-2 text-left text-[12px] font-normal normal-case leading-relaxed tracking-normal text-neutral-200 opacity-0 shadow-xl transition-opacity peer-hover:visible peer-hover:opacity-100\">Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment.<\/span><\/span>, and 99.9% for $19K with a <span class=\"relative inline-block\"><span class=\"peer cursor-help underline decoration-dotted\">Provider Adapter harness<\/span><span role=\"tooltip\" class=\"pointer-events-none invisible absolute left-1\/2 top-full z-50 mt-2 w-72 max-w-[calc(100vw-2rem)] -translate-x-1\/2 rounded border border-neutral-700 bg-neutral-950 px-3 py-2 text-left text-[12px] font-normal normal-case leading-relaxed tracking-normal text-neutral-200 opacity-0 shadow-xl transition-opacity peer-hover:visible peer-hover:opacity-100\">The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.<\/span><\/span>.<\/li>\n<li>GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels.<\/li>\n<li>A key behavior observed in GPT-6 Astra was its ability to turn unfamiliar environments into compact symbolic world models. It represented game mechanics as logical rules and developed its own domain-specific language shorthand to track state and plan actions.<\/li>\n<\/ul>\n<h2 id=\"arc-agi-3\" class=\"blog-heading-with-permalink\" aria-labelledby=\"arc-agi-3-heading-label\"><button type=\"button\" class=\"blog-heading-permalink\" title=\"Copy link to this section\" aria-label=\"Copy link to this section\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"15\" height=\"15\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/button><span id=\"arc-agi-3-heading-label\">ARC-AGI-3<\/span><\/h2>\n<p>ARC-AGI-3 is a benchmark for studying agentic intelligence through novel, abstract, turn-based environments. Agents must explore, infer goals, and build internal models of environments to effectively plan actions <em>without<\/em> explicit instructions. You can <a href=\"https:\/\/arcprize.org\/tasks\/ls20\">play ARC-AGI-3 yourself<\/a>.<\/p>\n<figure><video autoplay=\"\" muted=\"\" loop=\"\" playsinline=\"\" preload=\"metadata\" poster=\"\/media\/videos\/astra-arc-agi-3-poster.jpg\" aria-label=\"Animated ARC-AGI-3 benchmark graphic\" style=\"display:block;width:min(120%, calc(100vw - 32px));max-width:none;height:auto;position:relative;left:50%;transform:translateX(-50%)\"><source src=\"https:\/\/arcprize.org\/media\/videos\/astra-arc-agi-3.mp4\" type=\"video\/mp4\"\/><p>Your browser does not support embedded video.<\/p><\/video><\/figure>\n<p>These environments only contain <a href=\"https:\/\/arcprize.org\/arc-agi#:~:text=towards%20general%20intelligence.-,Core%20Knowledge%20Priors,-A%20principle%20underlying\">core knowledge priors<\/a> and are difficulty-calibrated through controlled testing with human participants. Humans can <a href=\"https:\/\/arcprize.org\/blog\/arc-agi-3-human-dataset\">solve 100% of the environments<\/a>.<\/p>\n<p>The goal of the ARC-AGI series is to measure the \u201cresidual gap\u201d between current artificial intelligence and AGI. We define AGI as a system\u2019s ability to acquire <em>any<\/em> skill a human can, as <em>efficiently<\/em> as a human can.<\/p>\n<p>ARC-AGI-3 is the third generation of the <a href=\"https:\/\/arcprize.org\/arc-agi\">ARC-AGI benchmark series<\/a>. It tests agentic capabilities beyond <a href=\"https:\/\/arcprize.org\/arc-agi\/1\">ARC-AGI-1<\/a> and <a href=\"https:\/\/arcprize.org\/arc-agi\/2\">ARC-AGI-2<\/a>. Each generation expands on the one before it &#8211; as frontier AI capabilities advance, our benchmarks must advance with them.<\/p>\n<p>ARC-AGI-3 tests four components of agentic intelligence:<\/p>\n<ul>\n<li><strong>Exploration:<\/strong> In real-world environments, information is rarely provided passively. Agents must actively obtain it by interacting with their surroundings.<\/li>\n<li><strong>Modeling:<\/strong> Agents must turn raw observations into a generalizable model that can predict future states and outcomes.<\/li>\n<li><strong>Goal-setting:<\/strong> Agents must identify target future states with only sparse rewards.<\/li>\n<li><strong>Planning and execution:<\/strong> Agents must map a path from their current state to a goal, course correcting as new information appears.<\/li>\n<\/ul>\n<h2 id=\"astra-results\" class=\"blog-heading-with-permalink\" aria-labelledby=\"astra-results-heading-label\"><button type=\"button\" class=\"blog-heading-permalink\" title=\"Copy link to this section\" aria-label=\"Copy link to this section\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"15\" height=\"15\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/button><span id=\"astra-results-heading-label\">Astra Results<\/span><\/h2>\n<figure><a href=\"https:\/\/arcprize.org\/results\/openai-gpt-6-astra\" style=\"display:block\"><\/a><figcaption>GPT-6 Astra achieves state-of-the-art scores on ARC-AGI-3 with both the Standard and Provider Adapter harnesses. Higher reasoning levels generally cost <em>less<\/em> because Astra solves games in fewer actions, reducing the total number of model calls and tokens. <a href=\"https:\/\/arcprize.org\/results\/openai-gpt-6-astra\">View the full results<\/a>.<\/figcaption><\/figure>\n<p>With our <span class=\"relative inline-block\"><span class=\"peer cursor-help underline decoration-dotted\">Standard harness<\/span><span role=\"tooltip\" class=\"pointer-events-none invisible absolute left-1\/2 top-full z-50 mt-2 w-72 max-w-[calc(100vw-2rem)] -translate-x-1\/2 rounded border border-neutral-700 bg-neutral-950 px-3 py-2 text-left text-[12px] font-normal normal-case leading-relaxed tracking-normal text-neutral-200 opacity-0 shadow-xl transition-opacity peer-hover:visible peer-hover:opacity-100\">Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment.<\/span><\/span>, OpenAI\u2019s Astra (max) <a href=\"https:\/\/arcprize.org\/results\/openai-gpt-6-astra\">scores 62.7% on ARC-AGI-3 Semi-Private<\/a> for $26K. With the <span class=\"relative inline-block\"><span class=\"peer cursor-help underline decoration-dotted\">Provider Adapter harness<\/span><span role=\"tooltip\" class=\"pointer-events-none invisible absolute left-1\/2 top-full z-50 mt-2 w-72 max-w-[calc(100vw-2rem)] -translate-x-1\/2 rounded border border-neutral-700 bg-neutral-950 px-3 py-2 text-left text-[12px] font-normal normal-case leading-relaxed tracking-normal text-neutral-200 opacity-0 shadow-xl transition-opacity peer-hover:visible peer-hover:opacity-100\">The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.<\/span><\/span>, Astra (high) scores 99.9% for $19K. Both are state-of-the-art scores. See the full <a href=\"https:\/\/arcprize.org\/leaderboard\">leaderboard<\/a>.<\/p>\n<div style=\"width:100%;overflow-x:auto;margin:20px 0\">\n<table class=\"info-table\" style=\"min-width:680px;margin:0\">\n<caption style=\"caption-side:bottom;color:#888;font-size:0.85rem;padding-top:8px;text-align:center\">\n<p>At max reasoning effort, Astra solves games more efficiently, requiring<br \/>\nfewer actions and therefore lowering total cost relative to the other<br \/>\nreasoning-effort levels.<\/p>\n<\/caption>\n<thead>\n<tr>\n<th>Reasoning effort<\/th>\n<th><span class=\"relative inline-block\"><span class=\"peer cursor-help underline decoration-dotted\">Standard harness<\/span><span role=\"tooltip\" class=\"pointer-events-none invisible absolute left-1\/2 top-full z-50 mt-2 w-72 max-w-[calc(100vw-2rem)] -translate-x-1\/2 rounded border border-neutral-700 bg-neutral-950 px-3 py-2 text-left text-[12px] font-normal normal-case leading-relaxed tracking-normal text-neutral-200 opacity-0 shadow-xl transition-opacity peer-hover:visible peer-hover:opacity-100\">Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment.<\/span><\/span><\/th>\n<th><span class=\"relative inline-block\"><span class=\"peer cursor-help underline decoration-dotted\">Provider Adapter harness<\/span><span role=\"tooltip\" class=\"pointer-events-none invisible absolute left-1\/2 top-full z-50 mt-2 w-72 max-w-[calc(100vw-2rem)] -translate-x-1\/2 rounded border border-neutral-700 bg-neutral-950 px-3 py-2 text-left text-[12px] font-normal normal-case leading-relaxed tracking-normal text-neutral-200 opacity-0 shadow-xl transition-opacity peer-hover:visible peer-hover:opacity-100\">The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.<\/span><\/span><\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>max<\/td>\n<td>62.7%, $26,098<\/td>\n<td>98.6%, $17,332<\/td>\n<\/tr>\n<tr>\n<td>xhigh<\/td>\n<td>59.3%, $37,317<\/td>\n<td>98.4%, $18,147<\/td>\n<\/tr>\n<tr>\n<td>high<\/td>\n<td>54.8%, $40,705<\/td>\n<td>99.9%, $18,817<\/td>\n<\/tr>\n<tr>\n<td>medium<\/td>\n<td>38.6%, $48,090<\/td>\n<td>98.4%, $19,285<\/td>\n<\/tr>\n<tr>\n<td>low<\/td>\n<td>17.5%, $38,166<\/td>\n<td>98.0%, $21,298<\/td>\n<\/tr>\n<tr>\n<td>none<\/td>\n<td>35.2%, $49,791<\/td>\n<td>96.7%, $23,457<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.<\/p>\n<p>Most of this fee pays for the participant\u2019s time and <em>willingness<\/em> to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain\u2019s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted.<sup id=\"fnref-1\"><a href=\"#fn-1\" aria-label=\"See footnote 1\">1<\/a><\/sup><\/p>\n<h2 id=\"analysis\" class=\"blog-heading-with-permalink\" aria-labelledby=\"analysis-heading-label\"><button type=\"button\" class=\"blog-heading-permalink\" title=\"Copy link to this section\" aria-label=\"Copy link to this section\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"15\" height=\"15\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/button><span id=\"analysis-heading-label\">Analysis<\/span><\/h2>\n<p>Beyond the scores, Astra\u2019s replays show how it turns unfamiliar game mechanics into useful working models. Three findings stood out: the compact algebraic notation it develops, its action efficiency compared with humans, and the custom tools it builds.<\/p>\n<h3 id=\"custom-algebraic-notation\" class=\"blog-heading-with-permalink\" aria-labelledby=\"custom-algebraic-notation-heading-label\"><button type=\"button\" class=\"blog-heading-permalink\" title=\"Copy link to this section\" aria-label=\"Copy link to this section\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"15\" height=\"15\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/button><span id=\"custom-algebraic-notation-heading-label\">Custom Algebraic Notation<\/span><\/h3>\n<p>When playing ARC-AGI-3, Astra chooses which strategy notes it would like to carry forward. It tracked objects, coordinates, rules, and unfinished plans, while also using a custom domain-specific language notation it generated for the environments.<\/p>\n<p>We\u2019ve seen similar behavior <a href=\"https:\/\/x.com\/arcprize\/status\/2080716567760007317\">in other models<\/a>, but Astra\u2019s notes stood out for their precision and information density. It distilled the scene into a compact code-like symbolic model: where objects were, how they interacted, and exactly which actions needed to happen in what order. This is an on-the-fly algebraic shorthand rather than a fully fledged programming language. For example:<\/p>\n<ul>\n<li><strong>Game state:<\/strong> <code>L8: hub q2 (8\u2193). Lengths: 14=1\u2026<\/code> records the level, a local rotation index, and mechanism lengths. <a href=\"https:\/\/arcprize.org\/replay\/39d9f100-328a-4121-ad81-ce298e1f9626?frame=219&amp;quote=L8%3A+hub+q2+%288%E2%86%93%29.+Lengths%3A+14%3D1%2C+9%3D1%2C+8%3D0%2C+12%3D0%2C+gate7%3D6%2C+gate10%3D7.&amp;quoteFrame=219&amp;quotePrefix=&amp;quoteSuffix=+Main+target+requires14%3D13.%0A%0APlan%3A+extend8+to3%3B+retract10+to2%3B+shorten8+to1.+Rot&amp;reasoning=decision\">s5i5, frame 219<\/a><\/li>\n<li><strong>Multi-step plans:<\/strong> <code>extend8 to3; retract10 to2; shorten8 to1<\/code> records an ordered sequence of changes to the color-8 and color-10 mechanisms. <a href=\"https:\/\/arcprize.org\/replay\/39d9f100-328a-4121-ad81-ce298e1f9626?frame=219&amp;quote=Plan%3A+extend8+to3%3B+retract10+to2%3B+shorten8+to1.+Rotate+once+to9%E2%86%93%3B+extend9+to2%3B+retract7+to2%3B+shorten9+to1.+Rotate+twice+to14%E2%86%93%3B+extend14+to13.&amp;quoteFrame=219&amp;quotePrefix=.+Lengths%3A+14%3D1%2C+9%3D1%2C+8%3D0%2C+12%3D0%2C+gate7%3D6%2C+gate10%3D7.+Main+target+requires14%3D13.%0A%0A&amp;quoteSuffix=%0A%0AControls+%E2%88%92%2F%2B%3A+14%3D%2853%2F59%2C11%29%2C+8%3D%2839%2F45%2C11%29%2C+9%3D%2839%2F45%2C4%29%2C+10%3D%288%2F14%2C56%29%2C+7%3D%2822%2F28&amp;reasoning=decision\">s5i5, frame 219<\/a><\/li>\n<li><strong>Controls and coordinates:<\/strong> <code>9\u2212=(39,4), rotate=(49,18), 14+=(59,11)<\/code> maps operations to the coordinates of the controls that perform them. <a href=\"https:\/\/arcprize.org\/replay\/39d9f100-328a-4121-ad81-ce298e1f9626?frame=235&amp;quote=Controls%3A+9%E2%88%92%3D%2839%2C4%29%2C+rotate%3D%2849%2C18%29%2C+14%2B%3D%2859%2C11%29.&amp;quoteFrame=235&amp;quotePrefix=0%3B+gates7%3D2%2C10%3D2.%0A%0ANext%3A+shorten9+to1.+Rotate+twice+to14%E2%86%93%2C+then+extend14+to13.%0A%0A&amp;quoteSuffix=%0A%0AACTION6+39+4&amp;reasoning=decision\">s5i5, frame 235<\/a><\/li>\n<li><strong>Time and position:<\/strong> <code>Turn 5: P=(24,20), empty, facing west<\/code> combines a turn counter with the player\u2019s location, carrying state, and orientation. <a href=\"https:\/\/arcprize.org\/replay\/be78fcef-1244-4cf8-b680-0a5e4e8f9afe?frame=708&amp;quote=Turn+5%3A+P%3D%2824%2C20%29%2C+empty%2C+facing+west.&amp;quoteFrame=708&amp;quotePrefix=&amp;quoteSuffix=+Continue+west+twice+to+%2816%2C20%29%2C+then+lift+crate+%2812%2C20%29.+Cyan%3D%2820%2C24%29%2C+carrying&amp;reasoning=decision\">wa30, frame 708<\/a><\/li>\n<\/ul>\n<figure><img decoding=\"async\" src=\"https:\/\/arcprize.org\/media\/images\/blog\/astra-symbolic-model.gif\" alt=\"Astra playing s5i5 while recording compact symbolic notes\" style=\"width:min(120%, calc(100vw - 32px));max-width:none;position:relative;left:50%;transform:translateX(-50%)\"\/><figcaption>Astra playing <a href=\"https:\/\/arcprize.org\/tasks\/s5i5\"><code>s5i5<\/code><\/a>, using its on-the-fly algebraic shorthand to track state and plan actions.<\/figcaption><\/figure>\n<h3 id=\"action-efficiency-compared-to-humans\" class=\"blog-heading-with-permalink\" aria-labelledby=\"action-efficiency-compared-to-humans-heading-label\"><button type=\"button\" class=\"blog-heading-permalink\" title=\"Copy link to this section\" aria-label=\"Copy link to this section\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"15\" height=\"15\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/button><span id=\"action-efficiency-compared-to-humans-heading-label\">Action Efficiency Compared to Humans<\/span><\/h3>\n<p>Before launching ARC-AGI-3, we tested approximately 500 members of the general public to establish a human baseline for action efficiency, or simply, how <em>quickly<\/em> did people solve each environment. Participants were not selected for puzzle-solving experience or ability.<sup id=\"fnref-2\"><a href=\"#fn-2\" aria-label=\"See footnote 2\">2<\/a><\/sup><\/p>\n<p>For each level, we defined the \u201chuman baseline\u201d using the <em>median<\/em> action count among players who completed it. This gives us a reference for comparing human and AI performance. An AI that needs <em>more<\/em> actions is less action-efficient, while one that needs fewer actions is more action-efficient.<\/p>\n<p>In the <span class=\"relative inline-block\"><span class=\"peer cursor-help underline decoration-dotted\">Provider Adapter harness<\/span><span role=\"tooltip\" class=\"pointer-events-none invisible absolute left-1\/2 top-full z-50 mt-2 w-72 max-w-[calc(100vw-2rem)] -translate-x-1\/2 rounded border border-neutral-700 bg-neutral-950 px-3 py-2 text-left text-[12px] font-normal normal-case leading-relaxed tracking-normal text-neutral-200 opacity-0 shadow-xl transition-opacity peer-hover:visible peer-hover:opacity-100\">The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.<\/span><\/span>, Astra (max) used <strong>fewer actions than the human baseline on 96.0% of levels<\/strong> and used <strong>51.7% fewer actions per level on average<\/strong>. This is a material milestone. This means by ARC-AGI-3\u2019s measure of action efficiency, Astra matched and surpassed human parity.<\/p>\n<p>As an aside, before we launched ARC-AGI-3, we hypothesized that action efficiency would remain a dividing line between humans and AI. We anticipated that even when an AI solved an environment, it might require substantially more exploration (actions) than a person. That remains true of brute-force approaches, but frontier AI shows a more binary-like pattern. Once frontier AI \u201cunderstands\u201d the mechanics, it generally executes within the range of human efficiency.<\/p>\n<figure>\n<h4 style=\"font-size:22px;font-weight:500;line-height:1.2;margin:0 0 18px;text-align:center;text-transform:none\"\/>\n<p>Astra\u2019s Action Efficiency Compared to Humans<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/arcprize.org\/media\/images\/blog\/astra-action-efficiency.png\" alt=\"Scatter plot comparing Astra actions with the human baseline for each completed ARC-AGI-3 level\" style=\"width:min(120%, calc(100vw - 32px));max-width:none;position:relative;left:50%;transform:translateX(-50%);margin-bottom:-12px\"\/><figcaption>\n<p>Each dot represents one level that Astra (max) completed. Points below the<br \/>\nsolid line indicate fewer actions than the human baseline.<\/p>\n<\/figcaption><\/figure>\n<p>The plot above compares the number of actions Astra used to complete each level with our human baseline. This reinforces why ARC-AGI-3 measures action efficiency, not just task completion. A completion-only score would tell us that Astra completed an environment, but not how efficiently it <em>learned<\/em> to solve them.<\/p>\n<p>Most benchmarks only measure <em>cost<\/em> efficiency, which measures the computational resources used, but <em>action<\/em> efficiency measures how much experience with an environment was required.<\/p>\n<p>Astra\u2019s results show that it needed fewer interactions than the human baseline to execute a solution.<\/p>\n<h3 id=\"custom-tools-in-agent-harness\" class=\"blog-heading-with-permalink\" aria-labelledby=\"custom-tools-in-agent-harness-heading-label\"><button type=\"button\" class=\"blog-heading-permalink\" title=\"Copy link to this section\" aria-label=\"Copy link to this section\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"15\" height=\"15\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/button><span id=\"custom-tools-in-agent-harness-heading-label\">Custom Tools in Agent Harness<\/span><\/h3>\n<p>We also evaluated Astra in the <a href=\"https:\/\/github.com\/alexisfox7\/PRO-LONG\">PRO-LONG harness<\/a> (<a href=\"https:\/\/arxiv.org\/pdf\/2607.20064\">paper<\/a>), an early ARC-AGI-3 red-teaming partner. In this advanced setup, Astra had access to a sandbox where it could execute custom code <sup id=\"fnref-3\"><a href=\"#fn-3\" aria-label=\"See footnote 3\">3<\/a><\/sup>.<\/p>\n<p>We observed Astra create a custom set of tools for each game: board parsers, game-state models, search algorithms, planners, and persistent notes. For more involved runs, Astra even produced small, game-specific software libraries.<\/p>\n<p>For example, in <a href=\"https:\/\/arcprize.org\/tasks\/tu93\"><code>tu93<\/code><\/a>, a maze-like game with guards and moving patrols, Astra started with navigation and built <code>maze_solver.py<\/code>. It added combat rules in <code>combat_solver.py<\/code>, modeled moving patrols in <code>patrol_solver.py<\/code>, and used <code>sync_state.py<\/code> to check its predictions against observations.<\/p>\n<p>Examining Astra\u2019s performance in PRO-LONG is useful because we see what it can do with <em>external<\/em> tools. However, this represents different evaluation conditions from our controlled human testing. Our testing participants did not have a code interpreter, scratch pad, etc., so PRO-LONG\u2019s results should be understood as the combined performance of the model and its tools.<\/p>\n<figure><img decoding=\"async\" src=\"https:\/\/arcprize.org\/media\/images\/blog\/astra-pro-long-tools.gif\" alt=\"Astra using a custom maze solver while playing tu93 in the PRO-LONG harness\" style=\"width:min(120%, calc(100vw - 32px));max-width:none;position:relative;left:50%;transform:translateX(-50%)\"\/><figcaption>Astra playing <code>tu93<\/code> in the PRO-LONG harness.<\/figcaption><\/figure>\n<h2 id=\"two-harnesses-two-questions\" class=\"blog-heading-with-permalink\" aria-labelledby=\"two-harnesses-two-questions-heading-label\"><button type=\"button\" class=\"blog-heading-permalink\" title=\"Copy link to this section\" aria-label=\"Copy link to this section\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"15\" height=\"15\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/button><span id=\"two-harnesses-two-questions-heading-label\">Two Harnesses, Two Questions<\/span><\/h2>\n<p>Our Standard harness for ARC-AGI-3 asks how models compare under the same minimal, provider-neutral interface. It provides all the information required to solve each game, but leaves the model responsible for deciding what to preserve in its visible notes. We believe a future AGI should be able to solve ARC-AGI-3 under these conditions. The shared interface also gives us a consistent, apples-to-apples comparison across providers.<\/p>\n<p>Alternatively, there is a separate question: how well does a model perform when it can use the context-management features its provider designed for it? For Astra, this means preserving the opaque reasoning state (which we don\u2019t see) between requests and using compaction to manage longer conversations.<\/p>\n<p>With the Provider Adapter harness, Astra&#8217;s best observed score on ARC-AGI-3 Semi-Private increased from 62.7% to 99.9%. Looking across Public and Semi-Private and all reasoning levels, Provider Adapter runs were approximately 3.66x faster by aggregate recorded elapsed time and used 49% fewer total tokens across the 167 game-reasoning pairs both harnesses solved.<\/p>\n<p>Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our <a href=\"https:\/\/github.com\/arcprize\/arc-agi-3-benchmarking\">open-source testing repository<\/a> and <a href=\"https:\/\/arcprize.org\/policy\">testing policy<\/a> document both approaches.<\/p>\n<h2 id=\"arc-agi-series\" class=\"blog-heading-with-permalink\" aria-labelledby=\"arc-agi-series-heading-label\"><button type=\"button\" class=\"blog-heading-permalink\" title=\"Copy link to this section\" aria-label=\"Copy link to this section\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"15\" height=\"15\" viewbox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-link2 lucide-link-2\" aria-hidden=\"true\"><path d=\"M9 17H7A5 5 0 0 1 7 7h2\"\/><path d=\"M15 7h2a5 5 0 1 1 0 10h-2\"\/><line x1=\"8\" x2=\"16\" y1=\"12\" y2=\"12\"\/><\/svg><\/button><span id=\"arc-agi-series-heading-label\">ARC-AGI Series<\/span><\/h2>\n<p>ARC-AGI-3 continues to be a useful playground for researchers and agents to explore unfamiliar environments, discover rules, and learn through interaction. Astra\u2019s results are also a major milestone worth celebrating. From our perspective, Astra represents a noticeable step-function change in frontier model capabilities.<\/p>\n<p>When we launched ARC-AGI-3, we <a href=\"https:\/\/arxiv.org\/pdf\/2603.24621\">made it clear<\/a> that saturating the benchmark would not represent \u201cproof of achieving AGI.\u201d Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.<\/p>\n<p>The ARC-AGI benchmark series is designed to evolve in tandem with frontier AI. This creates a feedback loop between emerging research questions and advances in AI capabilities. ARC-AGI-3 was our first interactive benchmark, which asked AI to efficiently synthesize causal world models and achieve goals without specific instructions. Astra clears this bar. At the same time, ARC-AGI-3 has a tightly bounded scope and format, and its environments have deterministic, closed-ended mechanics and goals. It does not represent the complexity and open-endedness of the real world.<\/p>\n<p>We are actively exploring the questions that should shape the next generation of benchmarks, including how to evaluate recursive self-improvement and open-ended innovation. Astra\u2019s progress helps clarify which AI capabilities are out of reach and which questions remain open.<\/p>\n<hr\/>\n<p>Thank you to Fran\u00e7ois Chollet, Mike Knoop, Matt Mazur, Ethan Bond, and Derek Smith for early review of this post.<\/p>\n<ol style=\"font-size:0.9em\">\n<li id=\"fn-1\">Assuming <a href=\"https:\/\/journals.sagepub.com\/doi\/10.1177\/0271678X17708691\">20 W of brain metabolic power<\/a> and an electricity price of $0.20\/kWh: 0.020 kW \u00d7 1.5 hours = 0.030 kWh, worth $0.006 per session, or $0.006 \u00f7 9 \u2248 $0.00067 per attempted game. <\/li>\n<li id=\"fn-2\">See the <a href=\"https:\/\/arxiv.org\/pdf\/2603.24621\">ARC-AGI-3 human testing paper<\/a>. <\/li>\n<li id=\"fn-3\">No evidence of trying to break out of the sandbox was observed.<\/li>\n<\/ol>\n<\/div>\n<p><a href=\"https:\/\/arcprize.org\/blog\/astra?utm_source=tldrai\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Summary GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harnessStandard harness enables a model to carry forward notes it chooses to keep with it throughout the environment., and 99.9% for $19K with a Provider Adapter harnessThe Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":23726,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-23725","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23725","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=23725"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23725\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/23726"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=23725"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=23725"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=23725"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}