The Write-Up Is Not The Change

The most dangerous pull request is not the one with an obvious bug. It is the one with a beautifully written description that confidently explains a change which, when you actually read the diff, was never made. I was reminded of this again recently, reviewing a change where the accompanying notes described a specific capability as implemented and verified, complete with a named test that had supposedly passed. The prose was clear, specific and completely plausible. The field it described simply did not exist anywhere in the code.
That gap used to be rare enough to ignore. A human writing a PR description from memory might round a detail or overstate a test, but the story and the diff rarely diverged by much, because writing the story took almost as much effort as making the change. That relationship has quietly broken. AI-assisted contributions can now produce a summary of intended work that is more polished, more specific and more confident than most humans would bother to write, and that polish is no longer evidence that the underlying change happened. It is evidence that language models are very good at describing what a sensible engineer would have done next, whether or not anyone did it.
This matters because code review has always leaned on the description as a shortcut. Nobody re-derives a change from first principles every time; you read the summary, form an expectation, and then skim the diff to confirm it matches. That shortcut works when the description is a lossy compression of real work. It fails badly when the description is a fluent guess about what the work should contain. The two look identical from the outside. The only way to tell them apart is to stop treating the description as a summary and start treating it as a claim that the diff either supports or doesn't.
The practical shift is small but has real teeth: read the diff first, or at least read it as though the description didn't exist, and only afterwards check whether the two line up. If they don't, the interesting question isn't "which one is right" but "why did something confident get written about something that didn't happen." In the case I was looking at, the honest answer was that the description had been drafted from the intended scope of the ticket rather than from the actual commit, and nobody had gone back to reconcile the two before asking for review. That is an easy trap to fall into precisely because the description reads as more authoritative than the code. A paragraph of well-structured English feels more trustworthy than forty lines of diff, even though the diff is the only one of the two that actually runs.
There's a second, quieter version of the same problem that shows up in test evidence rather than descriptions. A build log showing green, a soak-test number, a line claiming something was "verified on an isolated host" — these read as proof, but a proof only means something if you know what was actually run and against what. I've come to prefer asking a blunt question before accepting any of these as settled: what specific input, on what specific version of the code, produced this specific result? If nobody can answer that in one sentence, the evidence is closer to a vibe than a verification, however detailed it looks on the page.
None of this is an argument against using AI to draft descriptions, write release notes, or summarise a diff for a reviewer who doesn't have time to read four hundred lines. Those are genuinely useful uses of the tools, and refusing them on principle just pushes the work back onto people who have less time for it than the model does. The argument is narrower: once a description can be generated with effort close to zero, its existence stops being evidence of anything, and reviewers who haven't adjusted for that are extending trust on the old terms in a world that has changed the exchange rate. The fix isn't cynicism about every PR that arrives with a well-written summary. It's separating two questions that used to collapse into one — "does this sound like it was done properly" and "was this actually done" — and insisting on answering the second one from the diff, every time, regardless of how convincing the first one is.
There's a broader version of this that applies well beyond code review. Any process that lets someone (or something) describe its own work before that work is checked is vulnerable to the same failure mode, and it gets worse, not better, as the describer gets more articulate. A contractor's progress report, a vendor's implementation status, a team's sprint update — all of them can be accurate summaries of real progress, or fluent descriptions of intended progress that hasn't landed yet, and the writing quality tells you nothing about which one you're holding. The organisations that will handle AI-assisted work well aren't the ones that ban the tools or the ones that trust every output uncritically. They're the ones that already had a habit of checking the artefact instead of the story, and who now apply that habit a little more deliberately, because the story got a lot better at sounding true.
What made this particular case fixable, rather than just embarrassing, was that the diff was small enough to actually read end to end before approving. That's worth saying plainly, because it's the part that scales badly: as changes get larger and reviewers get busier, the temptation to trust a good description instead of reading the code only grows, exactly as the cost of that trust being misplaced grows with it. If there is one habit worth tightening now, it's keeping changes small enough that reading the actual diff remains the fast path, not the one you skip because the summary already told you what you wanted to hear.


Share your thoughts