Founder of Revin. Engineer by training, specialist in software development and digital products.

Automated review approves the shape of the code, never the intent behind it
AI code review is already good at the boring half: broken conventions, duplicated helpers, calls with no error handling, a new dependency riding in unnoticed, tests that run and assert nothing. In that half, the agent beats a tired human on their fourteenth approval of the day, and it reviews at 2am without complaining. What it does not do is tell you why that file exists, which customer asked for the exception living on line 212, and what breaks in Monday's operation if the rule changes.
So the practical cut, if you are designing this flow right now, is simple: the agent approves what a rule can verify, and a person approves whatever changes business behaviour, customer data access or an obligation to a third party. You can take the human off both sides. The bill for that decision just does not arrive in month one. It arrives around month fourteen, when somebody has to explain a rule nobody wrote down and nobody remembers.

When the question is why the rule exists, someone had to be in the room
The Pragmatic Engineer reported this week that 37signals, home of the creator of Ruby on Rails, moved to generating nearly all its code with agents, and that the next bet there is dropping code review altogether. It is a serious move by serious people, and I believe the results they report. In the same stretch, Simon Willison wrote that agents make engineering harder rather than easier, and Will Larson published his experiment with the software factory pattern. All three are true at once. Only the first one made headlines.
What rarely makes it into the reading is the boundary condition. 37signals runs a small, senior team that has maintained the same products for close to two decades, with very low turnover and a founder who still reads code. When a human reviewer steps out of the loop there, the knowledge stays in the building. When a human reviewer steps out of the loop at your company, they usually step out of the company.
Do the arithmetic before you copy the model. How many people on the current team were around when the billing module was written? If the answer is none, you do not have the context their decision rests on. You have the opposite: a system where code is the only source of truth and nobody can explain the intent behind it.
That is the scenario that shows up in most of the rescues that reach us. A codebase in the thousands of lines, three previous owners, and the most-edited file carrying a history that only says "fixes". AI will now write on top of that faster than any previous team could. It will also inherit the same blind spots, because it reads what was written and has no idea what was agreed in a meeting.
That is the mechanism by which technical debt from AI coding agents looks cheap now and expensive later. The agent is not what creates the debt. The missing person who answers for the decision is.

On a construction site, technical sign-off still carries a person's name
It does. And I use it for exactly that. It catches race conditions that slip by, error handling that swallows exceptions, a query inside a loop, an import nobody needs anymore. That is real value, and whoever is not using it is paying full price for a worse review.
The trouble starts when the thing that writes and the thing that reviews are the same class of tool. Two models with similar training tend to agree precisely where they are both wrong. I have seen the pattern in audits: lovely test coverage with not a single assertion that checks behaviour, and a health endpoint returning a hardcoded 200 while the database was down. Both sail through any automated review, because the shape is immaculate. It is a rotten shell with a quality badge.
Would an experienced human have missed it too? Maybe. The difference is that a human can be asked, and answers. A tool answers for nothing, and on audit day it does not sit in the room.
Here is what I would ask before signing off on the design:
Before software, I signed off as the responsible engineer on steel structure assembly. The logic there is blunt: machines cut, machines weld, instruments measure, and there is still a natural person's name on the technical responsibility record. Nobody ever let me sign with "the equipment approved it". Software will get to that point too, through the legal route rather than the technical one.
When generation and review both go automatic, what you are buying from a vendor changes in nature. You used to buy hours from people who understand the domain. Now you risk buying machine throughput plus a layer of process. And throughput is easy to demo in a two-week pilot.
Two things I would check in any proposal promising agent-speed delivery. First, who reviews what the AI writes and against what criteria, written into the contract, with names and seniority. Second, what happens on exit: if the vendor disappears in a year, what stays with you beyond the repository.
What we do runs the other way from generate-automatically plus approve-automatically. The agent comes in as a tool the team holds, every pull request keeps a named owner, and architecture decisions go through whoever will maintain the thing next year. Fewer lines per day. More system standing in month fourteen.
There is one cheap indicator worth tracking from today, and it has nothing to do with code quality. Measure how long the team takes to answer a business question about its own system. Something like: why do customers on the legacy plan not qualify for this promotion? If that answer took ten minutes a year ago and takes two days now, something important left the process, however green the delivery dashboard looks.
That delay degrades in silence. Nobody opens a ticket to complain about lost knowledge.
I cannot tell you whether the same holds for a three-person team that wrote everything from scratch in the past six months. There, the author's memory is still the backup, and automated review may be enough for a while. In a five-year-old system with many hands on it, removing the human reviewer erases the last person who knew why things were built that way.
So the question I would take into your next engineering meeting is not whether AI reviews well. It is this: today, if yesterday's pull request takes down tomorrow's revenue, who at your company can explain the decision in one sentence? Write the name down. If there is no name, you have not automated review yet. You have postponed a conversation.
7 read minutes
Article content: