{"id":23898,"date":"2026-09-11T13:10:14","date_gmt":"2026-09-11T13:10:14","guid":{"rendered":"https:\/\/scannn.com\/loops-graphs-harnesses-getting-quality-out-of-a-software-factory\/"},"modified":"2026-09-11T13:10:14","modified_gmt":"2026-09-11T13:10:14","slug":"loops-graphs-harnesses-getting-quality-out-of-a-software-factory","status":"publish","type":"post","link":"https:\/\/scannn.com\/lv\/loops-graphs-harnesses-getting-quality-out-of-a-software-factory\/","title":{"rendered":"Loops, graphs & harnesses \u2013 getting quality out of a software factory"},"content":{"rendered":"\n<div>\n<p>It&#8217;s a mess, right? If you have not been thoroughly disappointed in AI capabilities, you have not tried enough. And the expectations grow. We&#8217;ve been through the ladder of prompt engineering, context engineering, harness engineering. Then Steinberg is tweeting about <a href=\"https:\/\/x.com\/steipete\/status\/2063697162748260627?ref=ivokund.com\">Loops<\/a> one month, <a href=\"https:\/\/x.com\/steipete\/status\/2078277297791189132?ref=ivokund.com\">Graphs<\/a> the next. <\/p>\n<p>Should your <a href=\"https:\/\/addyosmani.com\/blog\/software-factories\/?ref=ivokund.com\">software factory<\/a> be more autonomous, more dark? Have you sacrificed quality for speed under pressure to deliver? Or is your job now dealing with what happens when <em>other people<\/em> trade quality for speed? Is AI output mostly just rubbish? Is it all a huge mess? Yes, maybe, maybe not.<\/p>\n<p>I have decided not to give up. And probably you can&#8217;t give up either, maybe because your manager wants you to use AI and ship faster. Maybe you have Kool-AId in your stomach (like me) and are determined to make the most of this new world. <\/p>\n<p>And hey! We&#8217;re not even at the peak of inflated expectations! Even if you wanted to, it would be lame to give up that early, right?<\/p>\n<figure class=\"kg-card kg-image-card\"><\/figure>\n<p>As it&#8217;s mostly only rant online about how you <em>can&#8217;t <\/em>do stuff properly with AI, I decided to write about how you <em>can. <\/em>Or at least how to improve your odds. <\/p>\n<p>This is the second part of my &#8220;Lies and deception of the yes-man&#8221; series. It talks about working with someone who is essentially an <a href=\"https:\/\/link.springer.com\/article\/10.1007\/s13347-025-00856-x?ref=ivokund.com\">AI psychopath<\/a>. I kid you not, check this out, reminds you of someone?<\/p>\n<figure class=\"kg-card kg-image-card\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-20.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"1788\" height=\"736\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-20.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-20.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/image-20.png 1600w, https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-20.png 1788w\" sizes=\"auto, (min-width: 720px) 720px\"\/><\/figure>\n<p><a href=\"https:\/\/www.ivokund.com\/lies-and-deception-of-the-yes-man-how-to-ensure-quality-in-agentic-workflows-part-1\/\">The previous post<\/a> explained that there are two important parts to it, and went deep into the first one \u2013 <strong>Alignment<\/strong>. Alignment is needed to make AI decisions better \u2013 more aligned with what you expect. <\/p>\n<p>The second part is here. It&#8217;s the control layer, whether you call it the harness, building loops, graph engineering or by some other name (I bet it will be called &#8220;the program&#8221; soon). This is what drives the prompts, the loops, and guides work through the graph. It used to be humans, but it&#8217;s now agents connected by messages and actions in your system. <\/p>\n<p>I&#8217;ll start with an <strong>example<\/strong>. This is one of my real factory runs from last week. It has 7 tasks and 90k LOC changes, agents working for 15h, with 2 human actions \u2013 one starting it, the other reviewing before final deploy. As it was a rather big change, I asked to review before deploying. 91 unique agents were involved. Most runs aren&#8217;t that big in terms of LOC, but are usually more parallel. <\/p>\n<figure class=\"kg-card kg-image-card kg-width-full\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-28.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"2000\" height=\"616\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-28.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-28.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/image-28.png 1600w, https:\/\/www.ivokund.com\/content\/images\/size\/w2400\/2026\/09\/image-28.png 2400w\"\/><\/figure>\n<p>Surprisingly, it turned out well. Actually, not surprisingly, because it usually does. But not always. This is the reality, success is not guaranteed, you must be ready for disappointments. You must know where the next one comes from and be ready to turn it into a factory improvement. Sometimes this is not possible, often it is hard, but mostly it&#8217;s possible. I&#8217;ve been building this factory for the last 6 months, and made all kinds of bad choices. So perhaps you find something useful here. <\/p>\n<p>So, I&#8217;ll start with how the factory works, go into some principles that guide the design, and then explain two of the workhorse loops that I use to get stuff done.<\/p>\n<p>By the way, if you are confused about all the different things people mean by <em>Loops<\/em>, then <a href=\"https:\/\/www.langchain.com\/blog\/the-art-of-loop-engineering?ref=ivokund.com\">this article<\/a> explains four types of loops very well. <\/p>\n<h2 id=\"the-main-loop\">The Main Loop<\/h2>\n<p>On the highest level, this is how it works \u2013 two loops, one for planning (creating specs), the other for execution (creating code), with a few human chokepoints for catching drift and problems. Separating <em>creating<\/em> work from <em>doing<\/em> work has been the single biggest efficiency gain for me, unlocking parallelism for both. <\/p>\n<figure class=\"kg-card kg-image-card kg-card-hascaption\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-22.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"2000\" height=\"562\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-22.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-22.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/image-22.png 1600w, https:\/\/www.ivokund.com\/content\/images\/size\/w2400\/2026\/09\/image-22.png 2400w\" sizes=\"auto, (min-width: 720px) 720px\"\/><figcaption><span style=\"white-space: pre-wrap;\">The most important part here is the feedback from the end to the beginning \u2013 improving the outer loop \u2013 the process, the machine, the factory. When I started off, a third of my time went there. Now it&#8217;s a couple of improvements every 1.8 days on average. <\/span><\/figcaption><\/figure>\n<p>I don&#8217;t run a distributed system where work happens and ships night and day. I want to be present in key moments and in key roles \u2013 be the bottleneck just enough to stop major chaos from happening and retain understanding of architecture and patterns. <\/p>\n<h2 id=\"the-main-principles\">The Main Principles<\/h2>\n<h3 id=\"principle-1-%E2%80%93-be-explicit-about-what-you-want-to-bottleneck\">Principle 1 \u2013 be explicit about what you want to bottleneck<\/h3>\n<p>Agents go very fast without humans. They also make more mistakes the longer they go on their own. So humans can be used to slow down the process where needed, and make sure the chaos is nipped in the bud.<\/p>\n<p>Usually, the limit is not how much work can you input to the machine. It&#8217;s how much output can you verify and still remain sane. <\/p>\n<p>This can be done in several ways. One time-consuming way is to review every line of code as the agent writes it. Another way is to direct your attention to a narrower choke-point, like when it&#8217;s done working and presents a PR. It&#8217;s better because you just slow the agent down every time it pushes a PR. You can move to higher and higher level there, but it&#8217;s important to be explicit about how your workflow expects human involvement. <\/p>\n<p>In my case, I <strong>always<\/strong> want to do the following myself:<\/p>\n<ol>\n<li>Plan product features<\/li>\n<li>Decide what to do when the plan changes during review (e.g. a critique agent found a wrong base assumption and the scope changed during planning)<\/li>\n<li>Approve improvements to the main outer loop (how the workflow works)<\/li>\n<li>Review infrastructure and data model changes in detail<\/li>\n<\/ol>\n<p><strong>Optionally<\/strong>, I want to sometimes be there for:<\/p>\n<ol>\n<li>Reviewing UX work once it&#8217;s done and before it&#8217;s shipped<\/li>\n<li>Review dangerous data heals in production before they&#8217;re run<\/li>\n<\/ol>\n<p>I absolutely <strong>never<\/strong> ever want to do the following:<\/p>\n<ol>\n<li>Review all the code agents produce, if my harness can do it<\/li>\n<li>Test the work of an agent, if the harness can do it<\/li>\n<li>Fix or improve the agent&#8217;s output directly, if I can do it through outer loop improvement instead<\/li>\n<li>Approve <em>safe<\/em> actions in production (safe deploys, migrations etc)<\/li>\n<\/ol>\n<p>I have seen a lot of frustration with agents because engineers do not believe that they can reduce this last list by outer loop improvements. I&#8217;ll go into this later. <\/p>\n<h3 id=\"principle-2-%E2%80%93-i-must-not-write-any-lines-of-code\">Principle 2 \u2013 I must not write any lines of code<\/h3>\n<p>It is so easy to just fix the crap that the agent produced. Or maybe just do it all yourself? A popular opinion is that LLMs write bad code, but in my experience this is not a problem with the model. For popular languages like TypeScript, the model can write <strong>any<\/strong> kind of code. The code that you do not like, someone else will love. It&#8217;s a mind with a billion split personalities, so more often than not, it&#8217;s not a problem of skill, it&#8217;s a problem of <strong>alignment<\/strong> \u2013 problem of defining what you expect. <\/p>\n<p>So this principle is simple. <strong>When I don&#8217;t like what comes out, I fix the machine, not its output, then run it again. <\/strong>Often this is one added line in <code>agents.md<\/code>, sometimes it&#8217;s a linter rule (more on that below). <\/p>\n<p>Here&#8217;s an important observation I&#8217;ve made: you don&#8217;t need to enumerate <em>all <\/em>the ways you don&#8217;t want code to be written. After surprisingly few guidelines (let&#8217;s say 10-20 principles you hold dearest) the model picks up the general <em>style<\/em> of your preferences, and starts producing code that is quite close to your expectations. <\/p>\n<p>Spend your precious human time on long-term improvement, not testing agents&#8217; work or fixing their code. Do work that compounds. <\/p>\n<h3 id=\"principle-3-%E2%80%93-shift-validation-left\">Principle 3 \u2013 Shift validation left<\/h3>\n<p>This one is obvious, but the power of it comes from actually applying it in practice. Left is close to ideation and planning, right is close to production and customers. <\/p>\n<figure class=\"kg-card kg-image-card\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-12.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"1522\" height=\"312\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-12.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-12.png 1000w, https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-12.png 1522w\" sizes=\"auto, (min-width: 720px) 720px\"\/><\/figure>\n<p>The more left you find a problem, the less work you have to undo and the less time you have wasted. Discovering that a button does the wrong thing in a UX test on the built app is a lot worse than finding the problem in plan validation. <\/p>\n<p>This is why most of my machinery focuses on plan validation. On average, my planning loop takes 1M tokens per issue, whereas the build loop takes 0.5M tokens per issue. <a href=\"https:\/\/www.codecentric.de\/en\/knowledge-hub\/blog\/shift-left-and-right-how-ai-integraion-ais-moving-beyond-coding?ref=ivokund.com\">Here&#8217;s<\/a> a good article on shifting validation both left and right. <\/p>\n<h3 id=\"principle-4-%E2%80%93-enforce-over-document\">Principle 4 \u2013 Enforce over document<\/h3>\n<p>This is the spine of the following chapters. The principle is simple: if something you want can be validated by good-ol&#8217; code instead of the LLM, then always prefer code. In other words: if a rule has a deterministic implementation, build it into the process. In other words, avoid LLMs always where possible!<\/p>\n<p>In my case, every instruction that is more than 1 line of code (e.g. cloning a worktree) resides in script files, not prose in skill files. Both of my loops started out as markdown files, but regularly pulling deterministic scripted parts out has two major benefits: 1) loading the skill takes up less context and 2) the model cannot screw up your script while running it. This is where I stand today with my main loops. <\/p>\n<figure class=\"kg-card kg-image-card\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-10.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"2000\" height=\"637\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-10.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-10.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/image-10.png 1600w, https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-10.png 2374w\" sizes=\"auto, (min-width: 720px) 720px\"\/><\/figure>\n<p>It is amazing how much you can actually do with deterministic checks built into your process. You can get a lot out of static code analysis and linters \u2013 when LLM writes code, there is no reason to not turn strictness up to 11.<\/p>\n<p>Some examples of scripts that work really well to keep the chaos level down in an LLM setting:<\/p>\n<ol>\n<li>Remove dead code \u2013 dead code is amazingly harmful for LLM planning processes, as their simple greps don&#8217;t understand whether code is live or dead.  <\/li>\n<li>Set up complexity caps \u2013 enforcing code simplicity will help both you and the LLM.<\/li>\n<li>Detect and block on code smells, they avoid bugs in both human and agent code.<\/li>\n<li>Turn your type system strictness to the max, and then some. Types are there to protect you from bugs, they are also documentation for the model. <\/li>\n<li>Make your markdown specs deterministic and executable. I&#8217;ve written about <a href=\"https:\/\/www.ivokund.com\/building-a-reliable-spec-driven-development-workflow-and-deleting-my-src-folder-to-test-it-claude-code-pt-2\/\">Gauge<\/a> before. One thing is prose about how the system behaves or spec about how the feature was planned. Another thing is human readable spec that actually runs your app and fails in CI on regressions.  <\/li>\n<\/ol>\n<p>This is a breakdown of the static checks my inner loops loop against, broken down by the loop and also type \u2013 a script is always better than a model doing the same work. And a human doing it obviously would be the worst. Log scale. <\/p>\n<figure class=\"kg-card kg-image-card kg-width-wide\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-11.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"2000\" height=\"1141\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-11.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-11.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/image-11.png 1600w, https:\/\/www.ivokund.com\/content\/images\/size\/w2400\/2026\/09\/image-11.png 2400w\" sizes=\"auto, (min-width: 1200px) 1200px\"\/><\/figure>\n<h2 id=\"creating-specs-from-plans-%E2%80%93-the-first-loop\">Creating Specs from Plans \u2013 The First Loop<\/h2>\n<p>This is all about capturing a problem and maturing it into a proper spec form. I&#8217;ve found that throwing in tokens here is worth it \u2013 so much better to discover an issue sooner than later. <\/p>\n<p>The first two steps are about formalizing what you want, then comes the automated loop for improvements.<\/p>\n<h4 id=\"idea-to-a-plan\">Idea to a plan<\/h4>\n<p>I&#8217;ve stayed true to the native Plan mode in Claude Code. Often there&#8217;s a discussion where I ask Claude to make a plan, and when it&#8217;s done I switch to Plan mode and ask it to do it <em>properly<\/em>. Plan mode forces it to do a bit more investigation and often it comes back with good improvements. <\/p>\n<figure class=\"kg-card kg-image-card\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-15.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"2000\" height=\"312\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-15.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-15.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/image-15.png 1600w, https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-15.png 2090w\" sizes=\"auto, (min-width: 720px) 720px\"\/><\/figure>\n<p>What you end up with in this step is almost always a horrible plan with holes, wrong base assumptions and sometimes an entirely wrong solution. If you take the bait and press Enter to &#8220;bypass and implement&#8221;, you will very very likely introduce new bugs into your system. But it&#8217;s a start. <\/p>\n<h4 id=\"plan-to-an-issue\">Plan to an issue<\/h4>\n<p>This is where the important planning parts actually happen. I have a skill called <code>\/github-issue-create<\/code> that makes sure the most important parts are handled properly:<\/p>\n<ol>\n<li><strong>Original user intent<\/strong> \u2013 adding this is one of the most valuable learnings I&#8217;ve had. Not having the original intent written down caused a lot of crazy over-engineering and scope drift in my earlier workflow versions. The skill specifically finds out (from the conversation) what I actually asked for, and writes it down in a few plain sentences. Every improvement to the plan is weighed against this. <strong>Non-goals<\/strong> are also an important part of this section.  <\/li>\n<li><strong>Acceptance Criteria<\/strong> \u2013 again, amazingly important for autonomous runs \u2013 describe how can the agent tell when to stop building. <\/li>\n<li><strong>Validation Steps<\/strong> \u2013 this is my latest addition to the workflow and it works wonders. Before this was introduced, the validation agent very often read and then misinterpreted or even skipped acceptance criteria items. AC can say that &#8220;This API needs to return X&#8221;, but who should test it and when? It sounds like an E2E test, so maybe it&#8217;s covered already? Or maybe the human will do it after deploy? This section specifies exactly which testing tier is used to test which exact ACs. The tiers are: unit test, E2E test, agent testing in local dev (in browser), agent testing in prod after deploy.<\/li>\n<\/ol>\n<p>After this, it is formatted as a detailed issue and posted to GitHub. <\/p>\n<figure class=\"kg-card kg-image-card\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-16.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"2000\" height=\"610\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-16.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-16.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/image-16.png 1600w, https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-16.png 2084w\" sizes=\"auto, (min-width: 720px) 720px\"\/><\/figure>\n<p>This is where the workflow for one single ticket ends. The plan is still in a very bad shape. If I ask an agent to find problems in it, it will find many. So this is what the next step automates. <\/p>\n<h4 id=\"issues-to-specs\">Issues to Specs<\/h4>\n<p>I start the loop by running my <code>\/prepare<\/code> skill manually. It starts a batch of 5-10 issues (only bounded by my token session limit), boots up parallel agents (from different vendors) that look at the GitHub issue and try to find problems with it. It runs in a loop, so if there are findings, they are collected and a new round of critique begins, until there are no high-severity findings. <\/p>\n<p>If you wonder why this is necessary and don&#8217;t have this loop up for your plans, then do this simple experiment. After you&#8217;ve come up with a good solid spec with an agent, copy the plan path, open up another Terminal window, paste in the file and ask a clean agent to critique and improve it. Paste the feedback to the original agent and let it improve the plan. Then do it again. See how many iterations it takes to arrive at a point where there are no horrors anymore in the plan. So this phase is moving validation &#8220;left&#8221;, as described above, but with automation. <\/p>\n<figure class=\"kg-card kg-image-card\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-17.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"2000\" height=\"1237\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-17.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-17.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/image-17.png 1600w, https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-17.png 2102w\" sizes=\"auto, (min-width: 720px) 720px\"\/><\/figure>\n<p>There&#8217;s a wrong way to do this, by the way. I made a huge and costly mistake here with an earlier iteration of this, here&#8217;s how:<\/p>\n<p>There&#8217;s an agent that summarizes the critique from all critique subagents every round (second box from the top on the image). In an earlier version this agent updated the original ticket content every round, which was then used as an input for the next round. This had the effect that the ticket&#8217;s scope started to drift badly as critics could no longer tell the original plan from requirements added by feedback. Most findings became about new stuff added during planning and this caused a lot of scope drift, over-engineering and super long planning iterations. <\/p>\n<p>Today&#8217;s version just appends all feedback to the ticket, listed under a &#8220;Feedback&#8221; section, so next critique agents can see the feedback and also the original ticket. One synthesizer agent looks at the entire thing in the end and creates a v2 of the ticket, considering all feedback and also the original intent. <\/p>\n<p>If there is still scope drift or some base assumptions are disproven, the plan is escalated to the human with options for how to continue. <\/p>\n<h4 id=\"putting-it-all-together\">Putting it all together<\/h4>\n<p>In essence this is how ideas reach the form of a spec in my factory.<\/p>\n<figure class=\"kg-card kg-image-card\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/Screenshot-2026-09-03-at-16.20.15.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"2000\" height=\"2449\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/Screenshot-2026-09-03-at-16.20.15.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/Screenshot-2026-09-03-at-16.20.15.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/Screenshot-2026-09-03-at-16.20.15.png 1600w, https:\/\/www.ivokund.com\/content\/images\/2026\/09\/Screenshot-2026-09-03-at-16.20.15.png 2130w\" sizes=\"auto, (min-width: 720px) 720px\"\/><\/figure>\n<p>There&#8217;s a specific reason why the Critique step does not automatically start after filing an issue. I like running the Critique loop on 5-10 issues together, so when it finishes, I can deal with all escalations together. A ticket takes an hour on average to plan, so I want my interruptions all at once. That&#8217;s one of the human chokepoints in the automated process. <\/p>\n<p>Issues end up on GitHub as tickets tagged <code>prepared<\/code>. This is food for the next loop. <\/p>\n<h2 id=\"specs-to-code-to-shipped-%E2%80%93-loop-two\">Specs to Code to Shipped \u2013 Loop Two<\/h2>\n<p>The goal of the previous step was to find all the surprises from the plans before implementation. If this was done well, then the next part is only about typing code as specced and getting it shipped \u2013 a process that can be automated pretty well. <\/p>\n<h4 id=\"configure-the-run\">Configure the run<\/h4>\n<p>It starts with the <code>\/fix-issues<\/code> skill that fetches open tickets and creates a dependency graph by ticket relations. I&#8217;m given a list of tickets to approve, and a preview of the graph of execution. Often there are 3-4 sequential batches, because tickets have dependencies.   <\/p>\n<figure class=\"kg-card kg-image-card\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-23.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"2000\" height=\"566\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-23.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-23.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/image-23.png 1600w, https:\/\/www.ivokund.com\/content\/images\/size\/w2400\/2026\/09\/image-23.png 2400w\" sizes=\"auto, (min-width: 720px) 720px\"\/><\/figure>\n<p>If specs contain production actions, like applying Terraform, running manual data heals, or substantial UX work, I can decide for each whether to do it myself, be there when the agent does it, or forgo control and authorize the agent to handle it fully. Once decisions are locked, the run is autonomous.<\/p>\n<h4 id=\"build-it\">Build it<\/h4>\n<p>For each selected issue, the following is done in parallel:<\/p>\n<ol>\n<li>A worktree is created: code, deps, a clean database<\/li>\n<li>Configuration assigned: port for running and browser-testing, and domains (as in DDD) that this issue touches and that need their expensive E2E tests run<\/li>\n<li>Agent writes code, opens PR<\/li>\n<li>Another agent reviews both code and functionality. It runs the app, takes screenshots as proof of working functionality. If there are findings, the PR is sent back to step 3<\/li>\n<\/ol>\n<p>Interestingly, the parallel ticket limit is set by my laptop&#8217;s RAM. Not because of models or LLMs, but because Chromium-based browser tests eat a lot of it and running more than 5 entire test suites in parallel will slow them down enough to start hitting timeouts, thus causing flakes and delays. So I&#8217;ve limited the workflow to do 5 in parallel max. <\/p>\n<p>As issues are completed, they are merged sequentially (following the original dependency graph) into a special &#8220;integration&#8221; worktree. It&#8217;s a worktree like all the others, so once all issues are merged, one final round of testing is done to make sure merges didn&#8217;t cause any change of behavior. <\/p>\n<p>So far all testing has been done on fixture data. There&#8217;s an optional step I can authorize (in the beginning), that will do a merge to my actual dev environment, so functionality can be tried out on more realistic data (by an agent or me). <\/p>\n<p>Visually, it looks something like this:<\/p>\n<figure class=\"kg-card kg-image-card\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-25.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"2000\" height=\"1824\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-25.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-25.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/image-25.png 1600w, https:\/\/www.ivokund.com\/content\/images\/size\/w2400\/2026\/09\/image-25.png 2400w\" sizes=\"auto, (min-width: 720px) 720px\"\/><\/figure>\n<h4 id=\"ship-and-review\">Ship and review<\/h4>\n<p>After work is validated locally, the run applies infrastructure changes (Terraform), merges code to <code>main<\/code> and monitors CI for a green run. Stuff comes up once in a while, so sometimes the agent applies follow-up fixes to make CI green. Mostly it&#8217;s new vulnerabilities that make the scan red, or flaky UX tests that behave differently in CI with lower CPU. <\/p>\n<p>Some changes are targeted at production data and require manual (in a sense that they are not in migrations) <strong>data heals<\/strong>. These are SQL statements that are run on the production database. They have been part of the plan and by this time, verified by the agent critic loops, and verified to work against production data. Optionally, I can review and watch them execute in this step.<\/p>\n<p>Now this next step was a major improvement in my workflow. I stopped being the tester for my agents and could hand over a truly E2E process. In the next step the agent runs <strong>prod smoke tests<\/strong> from all specs that had any. If a new UX flow was added, then the agent will open my browser and execute the flow, check data consistency, query latency, all that. It tries to actually use all the functionality that was added or modified, monitors Sentry and logs. This is the last batch of surprises that can come from infrastructure differences. <\/p>\n<figure class=\"kg-card kg-image-card\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-26.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"2000\" height=\"1272\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-26.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-26.png 1000w, https:\/\/www.ivokund.com\/content\/images\/size\/w1600\/2026\/09\/image-26.png 1600w, https:\/\/www.ivokund.com\/content\/images\/size\/w2400\/2026\/09\/image-26.png 2400w\" sizes=\"auto, (min-width: 720px) 720px\"\/><\/figure>\n<p>And now the most important step. I&#8217;m given a handoff file of everything that went wrong, every correction the agent had to make, every problem it randomly noticed. I give this to another planning agent to verify, prioritize and plan. And this then goes straight into the planning loop again.<\/p>\n<p>This is the part that compounds. This is the most important part if you want your machine and your mental health to improve. This is the medicine for the psychopath. Always take your medicine. <\/p>\n<h4 id=\"the-entire-thing\">The entire thing <\/h4>\n<p>If you are a visual person or actually didn&#8217;t read any text thus far, then this is the entire thing:<\/p>\n<figure class=\"kg-card kg-image-card\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-27.png\" class=\"kg-image\" alt=\"\" loading=\"lazy\" width=\"1254\" height=\"2478\" srcset=\"https:\/\/www.ivokund.com\/content\/images\/size\/w600\/2026\/09\/image-27.png 600w, https:\/\/www.ivokund.com\/content\/images\/size\/w1000\/2026\/09\/image-27.png 1000w, https:\/\/www.ivokund.com\/content\/images\/2026\/09\/image-27.png 1254w\" sizes=\"auto, (min-width: 720px) 720px\"\/><\/figure>\n<h2 id=\"final-thoughts\">Final thoughts<\/h2>\n<p>Is this perfect? Haha, no. Is this optimal for the thing I&#8217;m building with the resources I have? Seems like it. For some work it&#8217;s overkill and occasionally I skip either the planning or the parallel execution phase and just ask one agent to do a simple thing. Sometimes the work is heavily exploratory or sequential \u2013 in these cases I may lay back in nostalgia and watch one agent type code, which in some past era used to be my job. <\/p>\n<p>Do I still read code? Yes, not to validate agents&#8217; output, but to understand what&#8217;s going on. More often I browse around in the database to make sure what&#8217;s in my head matches what&#8217;s in the schema. <\/p>\n<p>The most annoying part? It&#8217;s definitely when reality comes back to destroy my nice little elegant plan with escalations from the real world. This is what I&#8217;d like to improve the most \u2013 delegating plan escalations more to the agent (in other words, have it not escalate, but deal with the surprises instead) \u2013 but so far, I don&#8217;t see that I can create enough context and alignment for this to work not horribly. So I&#8217;ve erred on the side of more feedback and more escalations. Power through the 8 escalations, dive deep into each one, spend an hour. But you&#8217;ll gain a set of good specs that will not fail you in implementation. <\/p>\n<p>What about security? For day-to-day code I rely on automated checks (around 10 plugins and skills scattered around the process) plus manual things (<a href=\"https:\/\/github.com\/vercel-labs\/deepsec?ref=ivokund.com\">deepsec<\/a>, <a href=\"https:\/\/docs.prowler.com\/getting-started\/basic-usage\/prowler-cli?ref=ivokund.com\">prowler<\/a>, <a href=\"https:\/\/www.checkov.io\/?ref=ivokund.com\">checkov<\/a>) I run once in a while (deepsec burns a lot of tokens). The manual ones have a huge amount of false positives, so I do it via an agent that also keeps track of them. I do occasional deep-dives to understand the security posture and get ideas for improvements. <\/p>\n<p>How to build your first factory? The trick is to do it piece by piece, balancing trust and autonomy at every step. Successful factories are <em>grown<\/em> from existing workflows, so you never give more autonomy than you can absorb back \u2013 i.e. validate. Validation is the bottleneck and the part that grows slowly. To delegate validation, you need trust and to build trustworthy loops, you need to go slow, absorb. Don&#8217;t turn off your lights at once and pray everything will work out fine (<a href=\"https:\/\/x.com\/dexhorthy\/status\/2080697380379427275?ref=ivokund.com\">it won&#8217;t<\/a>).<\/p>\n<p>Is this engineering now? Part of my work has definitely moved from building <em>things<\/em> to instead:<\/p>\n<ol>\n<li>Building things that <em>build<\/em> and validate things<\/li>\n<li>Focusing on <em>what<\/em> to build<\/li>\n<\/ol>\n<p><a href=\"https:\/\/www.ivokund.com\/lies-and-deception-of-the-yes-man-how-to-ensure-quality-in-agentic-workflows-part-1\/\">Alignment<\/a> is still important. Maybe more important than all this harness thing. Alignment will save you when your harness and planning both failed you. When planning slipped a requirement, the model needs to make a judgement call about scope, and your harness didn&#8217;t know this needs human escalation. Or it guessed a direction based on a gut feeling. A direction about security, about performance. It needs to know <em>you<\/em>, your problems, goals, your project, what your customers want.  <\/p>\n<p>Codebase structure and simplicity are hugely important. If your agent just says incredibly stupid things all the time, then most likely it cannot get enough context from your codebase. This is usually because most codebases are an incredible pile of spaghetti, without domain boundaries, encapsulation, documentation and so on. So if that hits home, fix <em>that<\/em> first, I&#8217;ve heard it also helps humans! <\/p>\n<p>The irony is that most of the codebase changes that make <em>agents<\/em> good at it, would have made <em>humans<\/em> better as well. But human timescales (both refactoring and measuring the impact of refactoring) are so long that, well, you just never go to that refactor, right? Now&#8217;s the time.   <\/p>\n<p>If you want the next one in your mailbox, then you can leave me your e-mail:<\/p>\n<div class=\"kg-card kg-signup-card kg-width-regular \" data-lexical-signup-form=\"\" style=\"background-color: #F0F0F0; display: none;\">\n<div class=\"kg-signup-card-content\">\n<div class=\"kg-signup-card-text kg-align-center\">\n<h2 class=\"kg-signup-card-heading\" style=\"color: #000000;\"><span style=\"white-space: pre-wrap;\">Sign up for Ivo&#8217;s AI blog<\/span><\/h2>\n<p class=\"kg-signup-card-subheading\" style=\"color: #000000;\"><span style=\"white-space: pre-wrap;\">Articles on AI and agentic software development, autonomous agent workflows, dark factories.<\/span><\/p>\n<p class=\"kg-signup-card-disclaimer\" style=\"color: #000000;\"><span style=\"white-space: pre-wrap;\">No spam. Unsubscribe anytime.<\/span><\/p>\n<\/p><\/div>\n<\/p><\/div>\n<\/p><\/div>\n<\/p><\/div>\n<p><a href=\"https:\/\/www.ivokund.com\/loops-graphs-harnesses-getting-quality-out-of-a-software-factory\/?utm_source=tldrnewsletter\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>It&#8217;s a mess, right? If you have not been thoroughly disappointed in AI capabilities, you have not tried enough. And the expectations grow. We&#8217;ve been through the ladder of prompt engineering, context engineering, harness engineering. Then Steinberg is tweeting about Loops one month, Graphs the next. Should your software factory be more autonomous, more dark? [&hellip;]<\/p>\n","protected":false},"author":16,"featured_media":23899,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[143],"tags":[],"class_list":["post-23898","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"_links":{"self":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23898","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/users\/16"}],"replies":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/comments?post=23898"}],"version-history":[{"count":0,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/posts\/23898\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media\/23899"}],"wp:attachment":[{"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/media?parent=23898"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/categories?post=23898"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/scannn.com\/lv\/wp-json\/wp\/v2\/tags?post=23898"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}