The Quiet Responsibility

THE QUIET RESPONSIBILITY

ARK 2026: The documents from the talk

The collection page for the talk “The tech lead in the AI age” at ARK 2026: the context file, the one-pager “Pass Code Review on the First Try” in anonymised form with the before-and-after figures, the line in the PR template, the TDR example, the full fixed-and-free table, the agent-profile excerpt, four drift signals and four things to do on Monday.

This is the collection page for the talk “The tech lead in the AI age: How to keep the architecture together when AI raises the speed” at ARK 2026, the Norwegian Computer Society’s architecture conference, October 14, 2026. I give it as a senior consultant at Ensō. The talk is in Norwegian; this is the English version of the page.

The talk shows four mechanisms, each with one or two documents on screen: the context file and the one-pager, the line in the PR template, a TDR, and finally the table of who owns which rule, with the profile that gets the rules through to the agent. Here they are in one place, so you don’t have to copy them out on Monday. The context file, the TDR and the table are examples to be replaced with the team’s own. The one-pager is from my own team, with the figures as they were.

On this page: The context file · The one-pager · The duty to explain: the line in the PR template · Size · The TDR example · Fixed and free: the full table · The profile · Four drift signals · Monday morning · Further reading

The context file

The first mechanism is standards where the work happens, and that is where the agent reads. A coding agent has no colleague who frowns at a 300-line function, and no gut feeling for “we don’t do it that way here”, so all of that has to be written in the file it reads. This is the file, nine lines; the lines in ⟨⟩ are the team’s own to fill in.

# AGENTS.md · ⟨team⟩ · under 60 lines, reviewed like code
Structure: vertical slices under src/features/. No cross-slice imports; a CI test fails otherwise.
Logging: structured (⟨library⟩), always correlationId. Never personal data.
Error handling: ⟨one pattern⟩, ProblemDetails toward HTTP. New pattern = TDR in the same PR.
Integration: only via ⟨API platform⟩; never directly against other domains' databases.
Tests: cover the behavior at the boundaries: input we don't trust, errors from services we call.
New libraries for mapping, caching, validation: TDR in docs/decisions/, in the same PR.
Org: ⟨one rule from the profile, as text⟩. The rest: read ⟨the profile⟩ first. Deviations require an exception.
Don't repeat what can be read from the code.

The tools read different file names: Copilot reads copilot-instructions.md, Claude Code reads CLAUDE.md, several read AGENTS.md. Keep one file, and let the others point to it. This is how the file grows on our team: every time the agent misunderstands us, we learn what we actually meant, and replace the line.

The one-pager: “Pass Code Review on the First Try”

The most useful one-pager on my team is called “Pass Code Review on the First Try”, and the word AI is not in it. It came about when I read six months of review threads from my own team back to front, March–August 2026, to see what the reviewers had to say again. The internal page names colleagues and the project, so this is an anonymised version: the figures, the five categories and the ten items are the same.

  • 474 pull requests
  • 726 threads opened by reviewers
  • 61% merged without a single comment

Without a single comment is not the same as without review: nearly nine in ten of them carry an approval from someone other than the author, mostly an “LGTM”. The page is about the other 39 percent, and the five that block most often are below, with thread counts. None of the five is about AI.

Blocks most often Threads
Assumptions instead of measurements 44
Stale branch 33
Hardcoded environment values 31
Acceptance criteria not met 22
Documentation contradicting the code 21

Ten items you check on your own branch before you press “Create”. Item five is the test line from the context file, in a stricter form. The list is adjusted as the threads change. All ten are decided before the reviewer opens the diff.

- [ ] Every acceptance criterion is either satisfied by this diff, or changed in the ticket to what I actually delivered.
- [ ] The branch is current with main as of this morning.
- [ ] Nothing in the diff is outside the ticket: no drive-by renames, no reformatting, no adjacent refactors.
- [ ] No secret, no personal data and no raw exception reaches a log, a response or the repository.
- [ ] At least one new test fails when the change is reverted, and I have watched it fail.
- [ ] The pipeline log shows that test running, by name.
- [ ] The build is linked, the environment named, the probes quoted.
- [ ] The README no longer says anything this diff has made untrue.
- [ ] The title is under 50 characters, imperative, with exactly one ticket key.
- [ ] The description reads as the commit body it is about to become.

So did it work? A little. Counted over all PRs, August against September 1 to October 8: rounds per PR went from 2.0 to 1.8, and 5 of the 7 authors with PRs in both periods need fewer rounds than before. Without the colleague behind the PR with 14 rounds, the drop is 1.9 to 1.8. Whether it is the authors bringing fewer things, or those of us who review asking for less, the figures cannot tell. The review bots were running in both periods. The share of PRs with no comment fell too, but mostly because more people on the team started commenting; threads per PR are nearly unchanged.

All PRs Before (August, 116 PRs) After (September 1–October 8, 166 PRs)
Rounds per PR 2.0 1.8
Threads per PR 2.2 2.1
Share of PRs with no comment 35% 17%

A round is a push after a reviewer comment, plus the first one. The share with no comment is lower than the half-year’s 61% because the review bots started commenting on many PRs in August.

“A little” is what I have learned to expect from anything written down: it can take care of some of what repeats, and the rest is still judgment. Your items will differ from mine; the method is the same. Look for what your reviewers are saying for the third time. That is the standard. Write it down where the person writing the PR is, before the PR is opened, and link to it from the review comment.

How I counted

  1. Fetch every PR in the period, with all its comment threads, straight from the PR tool’s API. For me: created March 1 to August 31, 2026, across 33 repos.
  2. Leave out machine traffic. For me that was 182 automatic access-request PRs; 474 remained.
  3. Attribute each thread to whoever opened it, and count only the threads a reviewer opened. For me: 726.
  4. Group the threads by what the reviewer had to ask for. An agent can do the first pass, but read the groups yourself before you trust them.
  5. Count rounds per PR before and after you start using the checklist, and count again after a month.

The duty to explain: the line in the PR template

The second mechanism is review that asks whether the code fits in, and with generated code: whether the author understands it. No apparatus is needed for that; one line in the PR template goes a long way.

## Generated?
Is any part of this change generated? Explain the choices you approved, in your own words.

The line gives the reviewer the right to ask “help me understand: why this pattern, and not the one in the neighboring class?” It is also the size limit: if you can’t explain 300 lines, split them. Everything that can be linted gets linted first. Give the AI reviewer the same context file, so it reviews against the team’s rules.

Size

Size is measured on our team too, on the team’s completed PRs from March to October 8, 2026, 580 in all. Four out of five PRs with one to three files went through on the first try. Of those over 25 files, three in ten.

Files in the PR Merged on the first try
1–3 80%
4–10 49%
11–25 40%
26+ 29%

The table says what it costs not to split. The limit is still the same: what you can explain in your own words.

The TDR example

The third mechanism is the decision trail: writing down why. A TDR answers three questions: what did we choose, why, and what did we reject. It is a markdown file where the code lives, in the same PR as the code. The aim is half a page and an hour. The TDR below is an example and not a real decision; feel free to replace it with one from the team. The three lines about the model, in the middle, are the ones I add when AI has been involved: what the model was told, what you checked afterwards, and what you still assume.

TDR-014 · Caching of product descriptions · status: decided · ⟨date⟩
Context: p95 over 1 s against the registry; read:write about 40:1; data may be 5 min old
Choice: local cache, TTL 5 min. Price and stock are fetched separately.
Rejected: shared cache (one more service to operate); event-driven invalidation (no events)
What the model was told: traffic profile, no real-time requirement
Checked afterwards: load test with the expected number of instances; TTL against the rate of change
Still assume: instances may show different versions for 5 min; nobody has measured how often
Revision point: after the first release, or when the registry starts emitting events

“Still assume” is the line I never drop, and behind it stands one more question: what don’t we know yet? “We don’t know” is a responsible answer. Then you need an investigation, and a date.

Fixed and free: the full table

The fourth mechanism is clear frames. On screen there were four rows; here is the full table. The top half is owned by the organization, the bottom half by the team, and the right-hand column says where the row lives. The org rows are from the organization’s own standards, translated and shortened; the infrastructure row is from a place I have been. “Free” and the team rows are mine; “free” is written nowhere, and that is part of the point.

Level Fixed Free Lives in
Org · identity and secrets Approved platform or library capabilities for authentication, authorization, cryptography and secrets; custom mechanisms require specialist review How roles map to permissions in the domain Secure coding standard (wiki; the profile points there)
Org · data and logs Collection, storage, transfer, logging, test use and deletion of sensitive data follow the security and privacy requirements Log format within the platform Secure coding standard (wiki; the profile points there)
Org · boundaries Untrusted input is validated at a defined trust boundary; authorization is enforced at the authoritative boundary Where the validation sits in the code Coding and design standard (wiki; the profile points there)
Org · new libraries New material dependencies are assessed for security, support, license, operations and supply chain; not every package needs committee approval Small utility packages within a row Dependencies and builds standard (wiki; the profile points there)
Org · shared contracts Changes to public interfaces, persisted data, configuration, events and shared components are assessed for compatibility, migration and rollback Sync/async within the domain Coding and design standard (wiki; the profile points there)
Org · infrastructure (example, from a place I have been) Shared infrastructure in the platform repo, once per environment, behind a contract The app’s own resources, within the contract The platform repo
Team · logging Structured, correlation ID, field names Log level per module AGENTS.md
Team · error handling One pattern, one mapping to HTTP Retry parameters, justified in the PR AGENTS.md + TDR
Team · new patterns TDR in the same PR as first use, with a revision point Experiment on a branch docs/decisions/

“Free” means a choice within the frame; the security requirements still stand. Look at the right-hand column (on a phone, swipe the table): every org row from the standards lives on the wiki. That is the gap we are working on.

The profile: how the agent should behave

On my team there is an agent profile in one repo that all the team’s repos are connected to, and the agent is told to read it before technical changes when it is started from the control repo; start it directly in one repo, and the profile isn’t there. Here is an excerpt, translated and shortened:

# agent-profile.md · excerpt, translated · 157 lines in the control repo
Authority: the organization's code maintainability standards, with seven supporting standards.
Owner: the maintainers of the agent setup. Reviewed at least annually; if the date has passed, report it.
A repo may strengthen the profile. It cannot weaken a mandatory control.
Agents cannot approve their own production changes. Independent review is done by another person.
Missing or unreliable checks are a governance gap. Report it instead of inventing evidence.
A debt record, a TODO or "the way we've always done it" is not an exception.

The profile opens by saying the wiki is still the governance authority and that the profile says how the agent should behave. One of our repos has taken the next step: its context file points to a reconciled copy of the standards in the repo. For the rest, the next step is getting the rules in as text, with only what an agent can act on in a diff.

A proposed TDR is not a yes. Where I work, the standard says an exception is created before a mandatory control is bypassed, with an accountable owner and approver, compensating controls and an expiry date, and the profile tells the agent to stop when the exception is missing or expired. The TDR is the reasoning; the exception is the yes. Security often owns a row at the top, and that is the row I would never move to “free”. You can tell a frame works when the team stops asking you for permission and starts asking whether the frame is still right.

Four drift signals

Everything above is about preventing drift. Four signals for seeing it, all cheap; the last two are for those of you who look across teams.

  1. Same need, several solutions: search across the repos and count variants. Example: rg -l ⟨library⟩ ⟨repo⟩ | wc -l, once per variant and repo. On our team, a code search on the default branches and a look at new packages in completed PRs, March to September 2026, found four caching libraries across five repos, three logging setups, at least two retry patterns and two ways of mapping. Nobody chose that; each of them was reasonable when it was written.
  2. Are many review comments about things the author could have checked? Then the author has not seen the standard before the PR, and it is the one-pager that needs to get better. On our team, about four in ten threads were like that, sorted by a model, so indicative only.
  3. The lockfile gets a new library with no TDR next to it. Across many repos: count only libraries new to the whole organization, in the categories you have a row for.
  4. TDRs with status “proposed” that are never decided. When they concern an org row, they are deviations in production that the board has not seen. Tag them with the row, and a job can collect them into one list before the board meets. In a regulated business, the owner of the row must approve the deviation before it goes to production.

And before all this: repos without a single TDR. On our team that was 30 of 38 in October.

Monday morning

Four things to do without asking anyone for permission, and the last one is for the architects.

  1. Start in one repo. Fix one line in the context file, or write the first one if the file is missing.
  2. Add the line to the PR template: “Explain the choices you approved, in your own words.”
  3. For the next design review: bring one rejected approach and one assumption that still needs checking.
  4. Architects: take the rule the board repeats most often. Write it into the profile, as text. Then: a pipeline out to every repo, and one number for the board: how many repos have the current version.

And measure two things, before and after: rounds per PR, and how many variants you have of the same need. The first shows the queue, the second the architecture. Send the two figures, without code or names, to hei@tauseansvaret.no, and I will publish them anonymised in November, whatever they show.

Further reading

The argument behind the talk is in Tech lead in the AI age: From expert to frame-setter. How a one-pager like this becomes part of the everyday is in AI usage guidelines people actually follow, and the TDR form with a longer example in TDRs in the AI age. If you want the texts as they come, you can sign up for The Tuesday Letter, one short email a week.

Read the Norwegian original →