> **TL;DR:** In February I said the dark factory wasn't the point. Aim for lights-off anyway, but run it like a Michelin kitchen: sample the dishes, fix the kitchen when the same mistake shows up twice, then franchise.

In February I wrote that "[most teams do not need a 'dark factory.'](https://www.georgediab.com/posts/2026/the-dark-factory#:~:text=Most%20teams%20do%20not%20need%20a%20%22dark%20factory.%22)" Eight months later I have agents merging their own commits on one project, a review gate on another that won't let a model approve its own diff, and a game where the agent wrote 101 of 105 commits and checked its own work by playing it in a headless browser. Last week Lauren Tan got asked on [Matt Pocock's stream](https://www.youtube.com/watch?v=MN9dGgmLyso&t=3115s) whether what she runs is a dark factory. She's [poteto](https://x.com/poteto) on X, she ships thousands of PRs a month, and she wrote [pstack](https://github.com/cursor/plugins/tree/main/pstack). She said, "[It kind of is, yeah, actually.](https://www.youtube.com/watch?v=MN9dGgmLyso&t=3115s)" That's my answer now too, so I owe the February post a correction.

I still agree with most of that post. The part I'd take back is the advice to pick a stopping point, and that was pretty much what the title was saying.

## What I still agree with

The February argument was that [verification, not generation, is the bottleneck](https://www.georgediab.com/posts/2026/the-dark-factory#:~:text=Verification%20is%20still%20the%20bottleneck.), that "[if you want autonomy, you have to pay for evidence](https://www.georgediab.com/posts/2026/the-dark-factory#:~:text=If%20you%20want%20autonomy%2C%20you%20have%20to%20pay%20for%20evidence.)," and that "[no human review](https://www.georgediab.com/posts/2026/the-dark-factory#:~:text=%22no%20human%20review%22%20is%20only%20defensible%20if%20review%20is%20replaced%20with%20something%20stronger)" only makes sense when review is replaced by something stronger than a human skimming a diff. I kept a list of changes that should stay behind mandatory human review: security-sensitive code, migrations, concurrency, compliance, incident playbooks. I still run all of that today. Every loop below was built to produce evidence before it was trusted to run on its own, and the list of things I won't let an agent merge unsupervised hasn't gotten any shorter.

Poteto said the same thing on the stream: "[verification plus the environment. It's the combination of these two things that allow me to step away and say: agents, go off and merge your thing. I'll review it in the morning by looking at my commit history.](https://www.youtube.com/watch?v=MN9dGgmLyso&t=3221s)" February was mostly about the verification half of that.

## Start at dark and work backwards

I was writing about [Shapiro's autonomy levels](https://www.danshapiro.com/blog/2026/01/the-five-levels-from-spicy-autocomplete-to-the-software-factory/) when I said this: "[They are less helpful when treated as a maturity ladder every team must climb. Different products, risk profiles, and compliance requirements should lead to different stopping points.](https://www.georgediab.com/posts/2026/the-dark-factory#:~:text=They%20are%20less%20helpful%20when%20treated%20as%20a%20maturity%20ladder%20every%20team%20must%20climb.)"

It felt reasonable at the time. Pick the rung that fits your risk profile and stop there. The trouble is what happens once you stop. A team that's decided "agent drafts, human reviews" is good enough has no reason to build the checks that would let it go further, so it never builds them. The evidence engine I said every team needed is exactly what nobody funds when lights-off is optional. I got the order wrong. I wrote "[invest in the bottleneck before pushing for more autonomy](https://www.georgediab.com/posts/2026/the-dark-factory#:~:text=Invest%20in%20the%20bottleneck%20before%20pushing%20for%20more%20autonomy.)." In practice it works the other way around. Pushing for autonomy is what pays for the bottleneck work, because that's when missing verification shows up as something blocking you instead of a nice-to-have. Aiming for dark doesn't mean taking the gates out first. It means asking what evidence would let me take each one out, and then building it.

Poteto got there in that order: "[the big unlock for me for getting to 2,000 PRs is starting from the question and working backwards of: how do I get to the point where my agent can merge its own code?](https://www.youtube.com/watch?v=MN9dGgmLyso&t=2378s)" If you start from that question and work backwards, everything you have to build along the way is the evidence system. If you start with the evidence system and try to work forwards, it tends to end up as a backlog item.

So the dark factory is the point after all, because aiming for it is what gets the rest built. Whether you ever fully turn the lights off is a different question, and for most of my work the answer is no.

## Call it a kitchen

The other thing I'd change about February is the word. Poteto has "[never really liked the term software factory](https://www.youtube.com/watch?v=MN9dGgmLyso&t=680s)," because "[a lot of people when they think factory, they don't necessarily equate that with quality or craft](https://www.youtube.com/watch?v=MN9dGgmLyso&t=685s)." She calls what she runs a "[Michelin kitchen](https://x.com/poteto/status/2106140831904399851)," and I've come around to that too. When people hear factory they picture mass production. Software is still creative work, and the finished dish still needs someone's judgment before it goes out.

In a kitchen the agents are the cooks, and you divide the stations on purpose so their work comes together into one good dish. The kitchen itself is your engineering environment. For us, the sharp knives and mise en place are skills, tools, code structure, verification, and lint rules. How well that's set up decides what the cooks can reliably produce.

The head chef owns everything that leaves the kitchen, no matter how many cooks worked on it. Poteto said it directly: "[You are not writing the code yourself anymore. You have agents. But you as the human are still responsible for the final outcome.](https://www.youtube.com/watch?v=MN9dGgmLyso&t=880s)"

That's how I square the kitchen with aiming for dark. For me, lights-off means the cooks don't need me hovering at every station, and I still taste what comes out.

## What it looks like when you aim for dark

Here are three of my kitchens.

Scrapwake is a browser space-arcade game I built with Cursor's cloud agents, driven from a Grok Bot project. As I write this it has 108 commits across all branches, and the agent wrote 104 of them. I wrote three, the last one is a merge by Cursor's bot, and all nine PRs open under my GitHub account. The verification loop was easy to describe and took most of the effort. I taught the agent how to play the game so it could confirm its own work, and had it take screenshots so I could check too. `npm run verify` builds the game and runs 154 unit tests. Then it runs Playwright smoke tests that drive the real game in headless Chrome through test hooks, with real keyboard events and real collision frames, and saves the screenshots. A coverage test fails the build if any gameplay contract in the table is missing one of its four required cases, so the agent can't quietly skip a scenario. The agent also runs revert checks on its own. It breaks the feature on purpose, confirms the test goes red, and restores it. Six of those are documented for one shield mechanic.

The loop lives in a skill the agent maintains called "Verify before done." It opens with "Do not tell the user a change is done, deployed, or ready to play until something has observed the behavior they asked for." Its rules are the ones I'd have written after a decade of code review if I'd been honest about what review was for. "Test hooks only set up state," never re-implement game logic. "Check rendered state, not saved state." "Green reached by deleting a check is a fail, so restore the check and fix the game." And the rule I like most, because most teams skip it: "When the user finds a bug verify missed, treat it as two bugs: the feature bug, and the gap in the checks. Fix both, add a structural guard so that kind of gap can't come back, and add the lesson to this skill." The skill grew every time I found something it hadn't. My job was playtesting the deployed build and sending feedback like "too fat, not cool looking enough." The skill even leaves that part to me: "Claiming the game feels good. Feel stays with the user's playtest." Those playtests are the only review I do on that project.

[Tightbeam](https://georgediab.com/posts/2026/running-an-organization-of-agents-through-tightbeam) is the org of agents I wrote about in September. Its rule is that nobody approves their own code. When a Claude coder says it's done, that's a claim, and it stays on hold until a reviewer in a different session, usually a Codex model, files a clean verdict. One branch went through four review rounds before it became a PR. Another took six rounds with agents run by [Mike Manzano](https://x.com/bffmike), who built Tightbeam. On the first night the org ran on its new machine, a Codex reviewer checking a Claude fix removed the fix, confirmed the gate failed without it, put it back, and only then signed off. Nobody told it to do that. Between August and early October, my agents sent 270 changes to an independent reviewer in another session, and about 4 in 10 came back with changes requested before they passed.

The Autotrader is where I hold the line hardest, because it moves money. The verification there is a drill. Kill the executor between order submit and broker acknowledgement, restart it, and prove there is exactly one broker order, one database row, and one decision marked acked. The evidence files live in the repo next to the code. The rule in that repo's agent instructions is "a claim without the command and its output is not verification," and it applies to me as much as to the agents.

So each project sits at a different distance from dark. Scrapwake runs lights-off apart from my playtests. The Autotrader never will, because it moves money.

## Where this breaks

Whether the domain can be verified decides almost everything here, and I don't have a trick for domains where running the code can't tell you it's right. Matt asked Poteto about one-way-door PRs, the kind that cause data loss you can't revert. She was honest about it: "[for domains where the work is verifiable, this is easier, and the one-way doors become two-way doors in a sense](https://www.youtube.com/watch?v=MN9dGgmLyso&t=3392s)," and then "[it's a great question that I don't really have the answer to](https://www.youtube.com/watch?v=MN9dGgmLyso&t=3455s)." Neither do I. A game is verifiable because you can play it. A migration is verifiable if you can run it against a copy and diff. A compliance workflow often isn't, and aiming for dark doesn't change that.

Rework is the other cost, and Dex Horthy is right to keep pointing at it. His line is that "[the lights off factory does not work](https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md#:~:text=the%20lights%20off%20factory%20does%20not%20work)," and his reason is that a PR needing 20% rework, which he calls generous, is "[both an intellectual burden and an emotional burden on both the submitter and the reviewer](https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md#:~:text=both%20an%20intellectual%20burden%20and%20an%20emotional%20burden)." He's describing a human reviewer absorbing agent rework, and in that setup he's correct. My loops hand the rework to a second agent and leave the human sampling. That only works when the second agent can run the code, so it comes down to verifiability again.

Tokens cost real money too. Poteto's full autopilot spawns verifier agents that click through the app looking for regressions and fix what they find, and she says it's "[quite token intensive](https://www.youtube.com/watch?v=MN9dGgmLyso&t=3206s)." I wasn't tracking tokens per branch, so I can't tell you what Tightbeam's long review loops cost.

## Review turns into lint rules

What carries over from my own setups to engineering teams is how review changes once you're sampling instead of blocking. A head chef can't taste every dish, especially once there's a second or third kitchen. Poteto put it this way: "[You don't want to be in a position where you're not tasting your food ever again. But you also, for scale, you cannot be tasting every single dish that comes out of your kitchen, especially if you have multiple restaurants. So it becomes more about sampling.](https://www.youtube.com/watch?v=MN9dGgmLyso&t=2975s)" Here the factory gets a little credit back. A quality inspector on a production line doesn't check every unit either. They pull samples, and when something's off they fix the line.

In practice that means picking a few PRs, reading them closely, and looking for what your kitchen keeps producing. Poteto's description is the best I've heard: "[you look at the code that the agents are writing and you scrutinize it very rigorously, and you think about all the inefficiencies, the bad patterns, and then you think about how to course-correct the environment. Not that single agent.](https://www.youtube.com/watch?v=MN9dGgmLyso&t=3009s)" If several agents take the same shortcut, or keep spreading a pattern you don't want, that's her cue to "[amend your kitchen](https://www.youtube.com/watch?v=MN9dGgmLyso&t=3045s)." The fix might be a new skill, a lint rule, or a type constraint. Once it's in, the check catches that mistake every time, including when the next person you hire makes it. The Scrapwake coverage test works this way. A gameplay contract with a missing case fails the build, so I never have to spot a skipped scenario in review. A PR comment fixes one instance and teaches nobody.

Do that regularly and you can step away for longer, because you trust the kitchen is constrained well enough to do the work without you catching its mistakes. Poteto reading her commit history in the morning is that kind of tasting, and so were my Scrapwake playtests. Neither one is a full inspection.

There's a management version of this. Poteto, who was at Netflix before Meta, put it as "[context, not control](https://www.youtube.com/watch?v=MN9dGgmLyso&t=2498s)." You can get the outcome you want by micromanaging, or you can give the engineer enough context to be self-sufficient and stop being in every thread. I spent years saying that about people. It turns out to work for agents too, and I'd bet the teams that get good at it with agents are the ones that were already good at it with people.

The plan is to build one kitchen that produces dependable results, where a repeated mistake turns into a change to the kitchen instead of a comment on a plate, and then franchise. Poteto is already there: "[I've spent the time building one kitchen and one restaurant. And now I'm in a position where I don't actually have to be there anymore.](https://www.youtube.com/watch?v=MN9dGgmLyso&t=2040s)" Opening a second restaurant doesn't get me out of the first, and I still answer for what goes out of both. Quality is the goal, and throughput is what you get once the automation is worth trusting.