{"id":23033,"date":"2026-08-05T05:22:42","date_gmt":"2026-08-05T05:22:42","guid":{"rendered":"https:\/\/scannn.com\/mirrorcode-whats-the-largest-software-project-ai-can-complete-on-its-own\/"},"modified":"2026-08-05T05:22:42","modified_gmt":"2026-08-05T05:22:42","slug":"mirrorcode-whats-the-largest-software-project-ai-can-complete-on-its-own","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/mirrorcode-whats-the-largest-software-project-ai-can-complete-on-its-own\/","title":{"rendered":"MirrorCode: What's the largest software project AI can complete on its own?"},"content":{"rendered":"\n<div data-pagefind-body=\"\" data-pagefind-weight=\"0.4\" data-astro-cid-xlfmpdvi=\"\">\n<p>AI has made rapid progress on software engineering benchmarks in the past few years. However, most such benchmarks tend to focus on shorter tasks like fixing bugs or implementing individual features. MirrorCode is our benchmark, co-developed with METR, to test AI models on long-horizon coding tasks. In a MirrorCode task, AI models are tasked with reimplementing an entire program end-to-end, without access to the original source code. AI-generated solutions must match the original program\u2019s output exactly on end-to-end tests, including held-out tests. MirrorCode\u2019s 25 target programs span different areas of computing: Unix utilities, data serialization and query tools, bioinformatics, interpreters, static analysis, cryptography, and compression.<\/p>\n<section class=\"mirrorcode-advantages\" data-astro-cid-mypu2fha=\"\">\n<h2 data-astro-cid-mypu2fha=\"\">How MirrorCode is different<\/h2>\n<div class=\"advantages-cards\" data-astro-cid-mypu2fha=\"\">\n<article class=\"advantage-card card\" data-astro-cid-mypu2fha=\"\">\n<h3 data-astro-cid-mypu2fha=\"\">Scale-aware evaluations<\/h3>\n<p data-astro-cid-mypu2fha=\"\">Crucially, we provide a large enough inference budget to make a serious attempt at MirrorCode tasks. Many existing software engineering benchmarks limit inference spending to around $1\u201310, even when the task would take weeks for a human to complete. For example, one of the largest MirrorCode tasks cost $2,600 for a single run and involved AI working for 19 days without human intervention.<\/p>\n<\/article>\n<article class=\"advantage-card card\" data-astro-cid-mypu2fha=\"\">\n<h3 data-astro-cid-mypu2fha=\"\">Difficult, but fair<\/h3>\n<p data-astro-cid-mypu2fha=\"\">Reimplementing entire programs is extremely challenging for human software engineers. We believe a human engineer without AI would take months to solve the most complex MirrorCode tasks. However, MirrorCode tasks are also feasible; we know that there is enough information for the tasks to be fair.<\/p>\n<\/article>\n<article class=\"advantage-card card\" data-astro-cid-mypu2fha=\"\">\n<h3 data-astro-cid-mypu2fha=\"\">Cheat-resistant by design<\/h3>\n<p data-astro-cid-mypu2fha=\"\">We sandbox AI models, requiring them to conduct their work without access to the internet, without access to the original codebase, and with no way to cheat on the task. There are end-to-end tests that models never see while developing their code, so they cannot simply create a lookup table to mimic the original program&#8217;s outputs.<\/p>\n<\/article><\/div>\n<\/section>\n<p><mdx-section depth=\"2\" id=\"section-ai-can-already-perform-some-long-horizon-coding-tasks\"><\/p>\n<h2 id=\"ai-can-already-perform-some-long-horizon-coding-tasks\">AI can already perform some long-horizon coding tasks<\/h2>\n<p>AI can already solve long-horizon MirrorCode tasks, despite their difficulty. For example, Claude Opus 4.7 reimplemented gotree: a bioinformatics toolkit with ~16,000 lines of Go and 40+ commands.<sup><a href=\"#user-content-fn-1\" id=\"user-content-fnref-1\" data-footnote-ref=\"\" aria-describedby=\"footnote-label\">1<\/a><\/sup> We believe this same task would take a human engineer without AI assistance 2\u201317 weeks. Opus 4.7 solved it in 14 hours, costing $251.<\/p>\n<p>One important caveat to these results is data contamination. Because MirrorCode tasks involve reimplementing open-source programs, AI models are likely to have seen the original codebases in pretraining. This might lead to inflated performance on the benchmark. However, AI successfully reimplemented several target programs that passed our memorization screen, and failed to reimplement programs where the screen showed evidence of memorization. This suggests that the results were not dominated by memorization, but we cannot rule out the possibility that memorization contributes to AI performance. Overall, we expect that the capabilities measured by MirrorCode would generalize to an unseen codebase. We discuss this further, along with more results and details on benchmark construction, in <a href=\"https:\/\/arxiv.org\/pdf\/2606.30182\">the paper<\/a>.<\/p>\n<p><\/mdx-section><br \/>\n<mdx-section depth=\"2\" id=\"section-leaderboard\"><\/p>\n<h2 id=\"leaderboard\">Leaderboard<\/h2>\n<p>MirrorCode is not fully solved. For our regularly updated leaderboard, we report <strong>MirrorCode (ML, +Private, 2L)<\/strong>. This means we run the 15 target programs from the Medium and Large buckets, and drop the Small bucket. Each target program is evaluated in two implementation languages (generally Go and Ada) giving 30 tasks. We run each task three times, with a budget of 10 billion tokens and 7 days per attempt.<sup><a href=\"#user-content-fn-2\" id=\"user-content-fnref-2\" data-footnote-ref=\"\" aria-describedby=\"footnote-label\">2<\/a><\/sup><\/p>\n<figure class=\"epoch-figure\"><astro-island uid=\"1KuBwY\" prefix=\"r49\" component-url=\"\/_astro\/Chart.CmriW99F.js\" component-export=\"default\" renderer-url=\"\/_astro\/client.PWbBjEwf.js\" props=\"{&quot;chartId&quot;:[0,&quot;mirrorcode-headline-scores&quot;],&quot;downloadFilename&quot;:[0,&quot;mirrorcode-headline-scores&quot;]}\" ssr=\"\" client=\"load\" opts=\"{&quot;name&quot;:&quot;Chart&quot;,&quot;value&quot;:true}\" await-children=\"\"><\/p>\n<div class=\"legacy-context chart-root _root_c28ba_8\" data-chart-root=\"true\" data-chart-id=\"mirrorcode-headline-scores\" data-export-mode=\"false\" data-title-visibility=\"default\" data-task-finished=\"false\" role=\"img\" aria-label=\"Horizontal bar chart of MirrorCode solve@100% rate by model: Claude Fable 5 at 64%, GPT-5.6 Sol at 20%, GPT-5.4 at 16%, and GPT-5.5 at 10%, with \u00b11 SE whiskers.\">\n<div class=\"_viewport_c28ba_378 \"><noscript class=\"_fallback_c28ba_37\"><\/p>\n<p class=\"trim body-3\">Enable JavaScript to see an interactive visualization.<\/p>\n<p><\/noscript><\/div>\n<\/div>\n<p><!--astro:end--><\/astro-island><\/figure>\n<div class=\"full-width-section no-toc-section\">\n<figure><astro-island uid=\"Z2a26dH\" prefix=\"r50\" component-url=\"\/_astro\/Chart.CmriW99F.js\" component-export=\"default\" renderer-url=\"\/_astro\/client.PWbBjEwf.js\" props=\"{&quot;chartId&quot;:[0,&quot;mirrorcode-solve-grid&quot;],&quot;downloadFilename&quot;:[0,&quot;mirrorcode-solve-grid&quot;]}\" ssr=\"\" client=\"load\" opts=\"{&quot;name&quot;:&quot;Chart&quot;,&quot;value&quot;:true}\" await-children=\"\"><\/p>\n<div class=\"legacy-context chart-root _root_c28ba_8\" data-chart-root=\"true\" data-chart-id=\"mirrorcode-solve-grid\" data-export-mode=\"false\" data-title-visibility=\"default\" data-task-finished=\"false\" role=\"img\" aria-label=\"Heatmap of MirrorCode solve rates for Claude Fable 5, GPT-5.6 Sol, GPT-5.4, and GPT-5.5 across 15 Medium and Large target programs.\">\n<div class=\"_viewport_c28ba_378 \"><noscript class=\"_fallback_c28ba_37\"><img decoding=\"async\" src=\"https:\/\/epoch.ai\/assets\/images\/charts\/mirrorcode-solve-grid.png\" alt=\"Heatmap of MirrorCode solve rates for Claude Fable 5, GPT-5.6 Sol, GPT-5.4, and GPT-5.5 across 15 Medium and Large target programs.\" title=\"The hardest MirrorCode targets remain unsolved\"\/><\/p>\n<p class=\"trim body-3\">Enable JavaScript to see an interactive visualization.<\/p>\n<p><\/noscript><\/div>\n<\/div>\n<p><!--astro:end--><\/astro-island><\/figure>\n<\/div>\n<figure class=\"epoch-figure\"><astro-island uid=\"1ff6dA\" prefix=\"r51\" component-url=\"\/_astro\/Chart.CmriW99F.js\" component-export=\"default\" renderer-url=\"\/_astro\/client.PWbBjEwf.js\" props=\"{&quot;chartId&quot;:[0,&quot;mirrorcode-language-resources&quot;],&quot;downloadFilename&quot;:[0,&quot;mirrorcode-language-resources&quot;]}\" ssr=\"\" client=\"load\" opts=\"{&quot;name&quot;:&quot;Chart&quot;,&quot;value&quot;:true}\" await-children=\"\"><\/p>\n<div class=\"legacy-context chart-root _root_c28ba_8\" data-chart-root=\"true\" data-chart-id=\"mirrorcode-language-resources\" data-export-mode=\"false\" data-title-visibility=\"default\" data-task-finished=\"false\">\n<div class=\"_viewport_c28ba_378 \"><noscript class=\"_fallback_c28ba_37\"><img decoding=\"async\" src=\"https:\/\/epoch.ai\/assets\/images\/charts\/mirrorcode-language-resources.png\"\/><\/p>\n<p class=\"trim body-3\">Enable JavaScript to see an interactive visualization.<\/p>\n<p><\/noscript><\/div>\n<\/div>\n<p><!--astro:end--><\/astro-island><\/figure>\n<p><\/mdx-section><br \/>\n<mdx-section depth=\"2\" id=\"section-open-source-code\"><\/p>\n<h2 id=\"open-source-code\">Open-source code<\/h2>\n<p>We release our scaffold and 22 of the 25 MirrorCode target programs (totaling 132 task instances across the six supported programming languages) as <a href=\"https:\/\/github.com\/epoch-research\/MirrorCode\">open-source<\/a>, with the other three targets held out as a private test set.<\/p>\n<div class=\"mirrorcode-credit-note\">\n<p>This work was co-developed with METR and supported by a grant from METR. The authors of MirrorCode are Tom Adamczewski, David Owen, and<br \/>\nDavid Rein. Florian Brand, Giles Edkins, Allen Hart, and Daniel O\u2019Connell contributed additional target programs. Rasmus Faber-Espensen<br \/>\nmade crucial infrastructure improvements and gave advice on engineering<\/p>\n<\/div>\n<p><\/mdx-section>\n <\/div>\n<p><a href=\"https:\/\/epoch.ai\/MirrorCode?utm_source=tldrdev\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>AI has made rapid progress on software engineering benchmarks in the past few years. However, most such benchmarks tend to focus on shorter tasks like fixing bugs or implementing individual features. MirrorCode is our benchmark, co-developed with METR, to test AI models on long-horizon coding tasks. In a MirrorCode task, AI models are tasked with [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":23034,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-23033","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23033","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=23033"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23033\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/23034"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=23033"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=23033"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=23033"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}