The review bottleneck: When AI produces faster than the team can judge
Generation has become cheap, judgment is still expensive. The asymmetry moves the bottleneck from production to review, and if the tech lead becomes the only quality gate, everything stops. The solution is mechanisms that scale evaluation capacity, not control.
Two speeds that used to match
The arithmetic of a development team has always had two sides: how fast code is produced, and how fast it can be judged. For decades those two sides stayed in rough balance, because writing code took time and the review queue kept pace with it. That balance is gone.
Generation has become nearly free, while judgment costs what it always did: human attention, context, and time. You can already see the result in your team. The PR queue grows and reviews turn shallow – or they stay thorough, and then everything stops.
The bottleneck has moved, and it has moved to the only place that can’t be automated away: understanding.
The new asymmetry problem
A developer now produces code faster than ever. How much faster is hard to measure: METR’s controlled study found in 2025 that experienced developers were 19 percent slower with AI, and the February 2026 follow-up points to about 18 percent faster for the same developers – a number METR itself calls unreliable, because developers now refuse to work without AI. But the reviewer still has to understand every line, and understanding got no turbocharger.
Worse: AI-generated code is often heavier to review than handwritten. It’s longer. It’s more “complete” – with error handling and edge cases the author hasn’t thought through. And it lacks the most important signal a reviewer has: the certainty that someone has already thought.
This is measured, not assumed. GitClear measured 623 million code changes from 2023 to 2026 so far, and compares them with 2022, the last year before AI. Back then, 21 percent of changed lines were moved code – the mechanical trace refactoring leaves behind – and 9.4 percent were copy-pasted. So far in 2026, 3.8 percent are moved and 15.7 percent copied. The crossover came in 2024, the first year more was copied than moved. The 2026 figures are preliminary; the direction is not. It is not a verdict on any single line. It is a description of what your review queue is filling up with.
And the queue has been measured too. LinearB’s 2026 benchmarks, 8.1 million pull requests from 4,800 teams, find AI-assisted PRs about two and a half times larger than unassisted ones – over 400 lines against 157 at the 75th percentile – and AI-generated PRs waiting more than 16 hours for a first reviewer, against about 200 minutes for human-written work. Under a third of them merge within 30 days; unassisted PRs do 84.5 percent of the time. LinearB’s own conclusion is this article’s: code review is the critical constraint in AI-assisted development.
The asymmetry creates a choice no one wants to make out loud: lower the pace, or lower the quality of the judgment. Most teams choose the latter without ever deciding to, one PR at a time.
The trap: Tech lead as the only gate
The instinctive response is centralization: “all larger PRs go through me.” It’s understandable. And it’s exactly wrong.
I have watched a principal engineer do it. His queue grew until one round of review could take up to a week, and it stayed that way for a couple of months. Amazon did the same at scale in March 2026: after several serious outages, an internal memo named Gen-AI assisted changes, and the answer was that such changes must be approved by a senior engineer. Amazon later said that only one of the incidents involved AI at all, and that the cause was a user error the systems had let reach too far. In Sonar’s 2026 survey, 38 percent of developers already say AI code takes more effort to review than a colleague’s. That effort now moves to the fewest and most expensive people. Faros’ AI Engineering Report from spring 2026, two years of telemetry from 22,000 developers, measures that move: where AI adoption is highest, median time in review is up 441.5 percent, and 31.3 percent more PRs merge with no review at all. The report calls it the tax on senior engineers: the code looks like it was written by someone who knows what they are doing, and the failures sit beneath the surface. The September 2026 follow-up, The Speed Trap, looks at teams that already use AI heavily, and as use deepens the rise in PRs merged without review has gone from 31.3 to 76.3 percent, time in QA is up 300.6 percent, and time in review remains heavily elevated. Faros’ conclusion has not moved since April: the greatest leverage lies where the code is written.
As I wrote in Tech lead in the AI age: when complexity rises, it’s tempting to tighten with more control and more approval. But that creates bottlenecks – and we’re now talking about a bottleneck that’s already overloaded.
You cannot review your way out of this. No one can. You have to build mechanisms instead.
Mechanisms that scale judgment
1. Automate the first line
Everything a tool can check, a tool should check: formatting, naming, known pattern violations, security rules. Analyzers and CI gates aren’t new, but their value has multiplied. Every automated check frees human attention for what only humans can do: judging whether the code belongs.
2. Move quality upstream
The cheapest review is the one that isn’t needed. Context files and shared prompts make generated code resemble the system from the start. Divergence is what costs review time.
3. The duty to explain lowers the cost of judging
A PR where the author has explained the choices makes the reviewer’s job far lighter. The reviewer no longer has to reconstruct the reasoning, only evaluate it. That’s why “understand before you merge” is as much a speed optimization as it is a quality principle.
4. Size as a norm
Large generated PRs are where judgment breaks down. Set a clear norm: break the task down until it fits in one small PR. In practice, a medium-sized PR takes three to five rounds of review. The same change split into three small ones gets approved in one or two. AI makes it cheap to produce a lot. So the team must make it normal to deliver a little at a time.
5. Distribute the review competence
If only two people can evaluate architectural deviations, you have two bottlenecks. Use pair reviews and rotating reviewers deliberately – not for fairness, but to build evaluation capacity across the whole team.
This is about protecting the pace
Note what this is not: an argument for slowing down. It’s the opposite. Teams that ignore the asymmetry get one of two outcomes: a growing queue that kills the flow, or a growing pile of code no one understands, which kills it later. Both are speed losses, just with different delays.
The mechanisms above exist to preserve the speed AI provides. Same logic as in Decision latency: momentum is protected by structure, not by effort.
A gate that needs no guard
The bottleneck in software development has moved, from writing code to understanding it. That’s not a problem solved by smarter reviewers or longer days. It’s solved by mechanisms: automation, context, the duty to explain, small units, distributed competence.
The tech lead’s job is not to stand in the gate. It’s to build the gate so it doesn’t need to be staffed by one person.
Evaluation capacity has become the team’s scarcest resource, and it only grows when more people than you can read a PR and say what’s missing.