Code review in the AI age: Judging code no one wrote
When the code in a pull request was generated in seconds, the reviewer’s role changes fundamentally. It is no longer about typos and better names, but about judging something the author themselves may not fully understand. What happens to code review when code is no longer written, but chosen?
What review took for granted
Code review has always rested on an unspoken assumption: that the author understands the code better than the reviewer. The author has thought the problem through, weighed the alternatives and made a choice; the reviewer arrives with fresh eyes and asks questions.
That assumption no longer holds. Today a PR can contain code the author generated in thirty seconds. The code is tidy, the names make sense, the tests pass. But when you ask “why did you choose this approach?”, the room goes quiet. That’s not laziness. It’s a new reality – and it demands a new review practice.
When both parties meet the code for the first time
A developer has error handling generated for an integration. The code uses tenacity with exponential backoff:
@retry(stop=stop_after_attempt(5),
wait=wait_exponential(multiplier=1, max=60))
def fetch_invoice(id):
return client.get(f"/invoices/{id}", timeout=30)
Elegant. Considered. Approved.
Three weeks later the integration fails in production, and no one knows why the retries behave the way they do – not the author, not the reviewer, not anyone. The code was never understood, only accepted – and that happened twice.
Look at what it actually says. Five attempts means four waits of 1, 2, 4 and 8 seconds, 15 in total. max=60 never engages; the cap would not bite until the seventh attempt, and there is no seventh attempt. The parameter looks considered and does nothing.
The timeout, on the other hand, is 30 seconds per attempt. In the worst case that is 15 seconds of waiting plus five attempts of up to 30 seconds each: close to three minutes in a single call chain. Put that behind an HTTP endpoint with a 30-second timeout of its own and the client has long since given up, while the retry logic keeps working for a caller that has long since gone.
And then the opposite problem, on the same line: with Tenacity’s default retry policy, @retry fires only on exceptions. client.get() may not raise on a 503 Service Unavailable, depending on the client; it may hand the response back as it is. So the one response that actually deserved another attempt never gets one. Meanwhile @retry, by default, catches every exception that does occur, including ones a second attempt will never help: a misspelled hostname, an expired certificate.
So a 503 is not retried, and a wrong hostname is retried five times.
None of these numbers is wrong in itself. The problem is that the parameters were never really chosen for this scenario.
As I wrote in Tech lead in the AI age: the problem is rarely a lack of answers. It’s a lack of structure around them.
From “is this good?” to “does this belong here?”
AI-generated code can look very good locally, and that is exactly the danger. The model is trained on best practice from all over the world, but it doesn’t know your system. It doesn’t know that you already have a retry pattern, that logging must carry correlation IDs, or that this particular module has a history that explains why things are done differently.
Everything can look right and still break the whole. Review therefore has to shift its focus:
- Not “is the code correct?” – tests and static analyzers often catch that better
- But “does it fit?” – does it follow our patterns, or introduce a new one?
- And most importantly: “does the author understand it?”
The duty to explain
The most effective move I’ve seen is also the simplest:
Whoever opens a PR must be able to explain it.
It’s a norm, not an interrogation, and in practice that means:
- complex choices are explained in the PR description, in your own words
- new patterns are flagged explicitly: “this deviates from how we do it – deliberately”
- “the AI suggested it” is never a justification
This is the same principle as in Standards that actually stick (in Norwegian): expectations must be clear in the everyday, not in a document nobody reads.
None of this needs machinery. A single line in the PR template goes a long way:
“Is any part of this change generated? Explain the choices you approved.”
It is not a harsh rule. The Linux kernel, which has one of the most demanding review cultures there is, has the same thing in its process documentation: you are expected to understand and be able to defend everything you submit, and if you can’t, don’t submit it. Submit anyway, and maintainers are entitled to reject the contribution without detailed review. In spring 2026 the guidance for AI assistants was added: generated code should carry an Assisted-by: line, and only a human can sign it off, because only a human can be held responsible. The terminal project Ghostty says it even more briefly in its AI policy: if you can’t explain what your change does, and how it interacts with the rest of the system, without the aid of AI tools, don’t contribute. Kubernetes went furthest in June 2026: if you cannot personally explain changes that AI helped generate, your PR is closed, and reviewers expect to discuss with a human, not with a model.
The reviewer’s new responsibility
The reviewer can no longer assume that someone has thought the code through. Distrust has nothing to do with it; the questions simply have to be different ones:
- “What happens if this call fails?”
- “Why this pattern and not the one we use in the neighboring class?”
- “Can you walk me through the flow here?”
Note the form: not “this is wrong”, but “help me understand”.
If the author can answer, all is well, regardless of who or what wrote the code. If they can’t, that is the review’s most important finding.
The tech lead’s role
As a tech lead you should not review everything. It doesn’t scale, and it creates bottlenecks – the same mechanism described in The Quiet Responsibility.
Your job is to build the mechanisms:
- PR templates that make understanding explicit
- example code that shows what “good” means here
- a culture where “I don’t understand this” is a legitimate review finding
- automated enforcement of what can be automated, so humans can spend their time on what can’t
Review after code got cheap
Code review was never only about finding bugs. It was also about shared understanding, and that is the part AI doesn’t change, only makes more visible. When code can be produced without being understood, review becomes the last place understanding can be secured.
Then the question is no longer whether the code is good enough to merge, but whether anyone in the room can explain why it looks the way it does.