AI can run the software factory. It can't turn the lights off yet
StrongDM says code must not be written or reviewed by humans. We build GetPullRequest with coding agents every day. Here's what broke in our software factory when nobody was looking.
In 2018 Elon Musk admitted that excessive automation on the Model 3 line had been a mistake, and added: "Humans are underrated." Software is now having the same argument.
In January Dan Shapiro wrote up five levels of AI coding. Level 5 is the dark software factory, "a place where humans are neither needed nor welcome." A couple of weeks later StrongDM described how their AI team works, and Simon Willison wrote it up. Their rules:
Code must not be written by humans. Code must not be reviewed by humans. If you haven't spent at least $1,000 on tokens today per human engineer, your software factory has room for improvement.
We build GetPullRequest with coding agents, every day, and have since May (the launch post has the backstory). Most weeks it does feel like a software factory. It has never been dark, and below are a couple of the reasons why, straight out of our commit history.
TL;DR: An AI software factory, where agents do most of the work from ticket to pull request, is real, and we run one. The dark version, where no human writes or reviews anything, breaks on something manufacturing never had to solve: every unit of software is a new design, and working out whether it's right is the hard part. Automate the line, keep people at the few stations where judgment happens, and let them get there without sitting at a desk.
Manufacturing repeats a design. Software makes a new one every time
Lights-out manufacturing is real. FANUC has a plant where robots build robots and can run for long stretches with nobody on the floor, and there are a lot of videos of Chinese "dark factories" going around.
It works because the part was designed once and gets made a million times. The spec is a drawing with tolerances. Did the bracket come out at 40.00 mm, plus or minus 0.05? A gauge tells you, every time, for almost nothing. The design work happened months earlier, in another building, with the lights on.
Software doesn't have the million-copies part. git clone is the entire production line. Every ticket is a new design, so a "dark software factory" really means automating design and the checking of design. That's the part real factories never automated.
What broke when nobody was looking
Here's what that looks like in practice. Every one of these comes straight from our own commit history.
The agent said "done". There was nothing in it.
When an agent finished a task without changing any code, GitHub refused to open a PR, because there was nothing to compare. We'd decided early on that this wasn't really an error, since the task had run and nothing had crashed. So GPR quietly skipped the PR step and moved the task to Done anyway.
What we saw on the board was a green task, a "PR opened" event in its timeline, and no PR. Every automated signal agreed it was a success. The agent had exited cleanly and written a confident summary, and no test failed because there was nothing new to test. An agent that never touched the repo looked exactly like one that had finished the job, unless a person went looking for the diff. The checker had the same blind spot as the worker. In August we flipped it: no commits now means the task failed.
One wrong character hid two bugs.
Tasks were timing out after 90 seconds and we couldn't see why. The log line meant to explain it came out as an empty template, because it used %q, a Go formatting code, inside Python. Python choked on it, the logger swallowed the error, and the one number we needed (had the agent replied at all?) never got printed.
Worse, a test had been failing on exactly this for a while, and we'd been writing it off as an environment problem. Once the log was fixed, the next bug showed up behind it: under the right timing the daemon could start two agents for one session, and the extra one became a ~300 MB process nothing would ever clean up. Production showed one session spawned twice within 2.2 seconds. Even after both fixes, the original timeout still wasn't explained. We'd only made the next one visible.
A pipeline that just wants CI green, with nobody reading logs, would have called that test flaky and rerun it. That's what we did too, until someone read the log.
StrongDM's answer to this kind of problem is smart. They keep scenarios, end-to-end user stories, outside the codebase like a holdout set so the agents can't game them, and they score "satisfaction" probabilistically rather than pass/fail. Right instinct. But someone writes the scenarios, and someone decides what satisfied means. That's review, moved upstream.
The digital twin doesn't know the world changed
StrongDM's other big idea is the Digital Twin Universe: behavioural clones of Okta, Jira, Slack and Google Docs, so agents can test at volumes the real APIs would rate-limit. Hitting real Slack in a loop is miserable, so we get it. Two of our worst integration bugs this year would have gone straight through one, though.
Jira changed under us.
In September Atlassian removed an old search endpoint and our Jira calls started coming back 410 Gone. A Jira twin built in August would have kept happily returning results. The fix was itself a Jira ticket, PET2-4, which GPR picked up and turned into PR #500, tests passing first try. So the factory part was fine. It needed someone to notice the 410 and write the ticket.
Slack ate our screenshots.
Slack delivers every @gpr mention twice, and the attached files are only on one of the copies. We were reading the other one, so the screenshots people attached to bug reports silently disappeared (more in the Slack post). Nobody puts that in a clone, because nobody knows about it until a real person's screenshot goes missing.
Our real environment wouldn't fit in a sandbox.
Before GPR we tried hosted sandboxes, with Daytona. Getting our backend, Flutter app, database and Go daemon to come up together in there was a project of its own, and we gave up (more in the pillar post). Then there's Windows, where connections went half-open without the daemon noticing and our pairing QR code printed as an empty box in Windows Terminal. No twin was going to tell us any of that.
As for the $1,000 a day: Simon called it "more of a business model exercise", and it isn't the real problem anyway. Spend that much with no reliable way to tell whether the output is right and you've bought a faster way to ship things nobody checked.
| Manufacturing dark factory | "Dark" software factory | What we run | |
|---|---|---|---|
| Each unit is | A copy of one design | A new design | A new design |
| "Is it right?" | A gauge, cheap and certain | Agents grading agents | Agents first, then a person merges |
| When the world changes | Rarely | The twin doesn't know | Whoever's on the board notices |
What our software factory looks like
It is a factory. Tickets come in from GitHub issues with a label, from Jira, from a Slack mention. A daemon on one of our machines picks each up, makes a git worktree, starts whichever agent suits it, runs plan, implement, verify, and opens a PR. 370 commits in our repo carry Co-authored-by: GPR, and the branch list is mostly gpr/task-tsk_…. The conversation minimap in the task screen was PR #318, which GPR opened against its own repo.
The lights are on at three stations. Somebody writes the task, which is where the design happens; a vague ticket gets you a confident, wrong PR. Somebody answers the agent when it stops to ask, which happens more than you'd think and is honestly the most useful thing it does. And somebody reads the diff and merges.
Most of what a person does at those stations is small. Answer "should I keep the old endpoint as a fallback?", approve a command, read a 40-line diff, merge. None of it needs a desk. That's basically the product: agents run on your machine, your phone buzzes when one needs you, and you review and merge from there. If you only ever drive one Claude session, Claude Code Remote Control does the phone part fine.
It's not frictionless. PR #502, worktree badges leaking into unrelated chat sessions, took four agent rounds. PR #490, inline comment gestures not registering on diffs, took five. Somebody looked at every round. That's a normal week.
Where we might be wrong
Models will get better at checking their own work, and holdout scenarios are a real step. For a CRUD service with great coverage and no third-party APIs, a mostly dark pipeline could probably work today. And if review really goes away, the phone half of our product matters a lot less. Worth saying, since we're the ones selling it.
Our bet is that the stations move and don't disappear. Fewer people typing, more people deciding what gets built and saying yes or no to what comes back.
Keep reading
- Your AI coding agent is stuck in a terminal you have to babysit. So we built GetPullRequest
- We ditched cloud AI coding agents. Our laptops were faster, and already paid for
- We merge pull requests from our phones. It's less reckless than it sounds
- A Jira AI agent fixed our Jira integration
- Devin AI alternatives: why pay for another AI engineer?
- Agent Client Protocol (ACP) is the most important AI coding standard nobody's talking about
Try the lit version
GPR is free for projects, sessions, and the GitHub and Slack connections; you pay only for more workspaces or more tasks running at once. Get the app on Google Play (iOS app coming soon), then on a Mac, Linux or Windows machine with one coding agent installed:
curl -fsSL https://getpullrequest.com/install | bash
cd ~/code/your-repo
gpr setup # scan the QR code with the app
Give it something small and boring first, then go do something else until your phone buzzes.