← all articles

The Resolve Button Isn't The Agreement

The Resolve Button Isn't The Agreement

A merge request with a passing pipeline and every discussion marked resolved is not the same thing as a merge request everyone actually agrees is ready. Those are two different claims wearing the same green ticks, and mixing them up is one of the quieter ways good teams slow themselves down without noticing why.

I've been reviewing a run of merge requests this week where the mechanical signals all point the same way: pipeline green, requested changes addressed, comments marked resolved. And yet several of them are still sitting open, because a resolved thread turned out to mean "someone clicked resolve," not "we're aligned on this." One reviewer asks for another look after a rebase. Another discussion gets closed by the author before the reviewer has actually confirmed the new approach solves the original concern. The tooling reports the same thing in both cases: no unresolved discussions, ready to merge. The tooling is not wrong. It's just answering a much narrower question than the one people are actually asking when they glance at a merge request and decide whether to trust it.

This is worth naming because the failure mode is invisible from the interface. A checkbox and a shared understanding produce identical pixels. The whole point of turning a conversation into a status flag is to compress it into something scannable, and compression always throws something away. What gets thrown away here is the difference between "this objection has been addressed" and "this objection has been dismissed, deferred, or quietly worn down." All three close the thread. Only one of them means the work is actually settled.

I've watched the same pattern show up one layer up the stack, in ticket-tracking status rather than review threads. A ticket sits in a state that used to mean "changes required," the actual changes get made, the pipeline goes green, and the status field never catches up, because nothing in the workflow forced someone to revisit it. Nobody lied. Nobody skipped a step. The state just stopped being load-bearing the moment the real work moved faster than the field describing it. That's the general shape of the problem: any status flag that can be true without anyone re-checking the thing it claims to summarise will, eventually, be true for the wrong reason.

The instinct when you notice this is to add more rigour to the flag itself — require a comment before resolving, add an explicit "reviewer confirms" step, bolt on another gate. That helps a little and costs more than it looks like it costs. Extra ceremony around a status flag doesn't close the gap between the flag and the underlying agreement; it just adds friction that experienced people learn to route around under deadline pressure, which puts you back where you started with one more box to tick. The fix isn't a stricter flag. It's remembering, at the point of use, that the flag was always a pointer to a conversation and never a substitute for having had it.

In practice that means treating "resolved" as an invitation rather than a verdict. When I open a merge request and see every thread closed, my first move isn't to check the box count — it's to skim what actually changed in response to each comment and ask whether that change is what the original concern needed, or just the smallest edit that would make the comment go away. Those look identical in a diff. They only diverge when you read for intent instead of for presence. It costs a few extra minutes per review. It saves the much larger cost of a merge request that gets re-opened, re-argued, and re-reviewed three days later because the first resolution was cosmetic.

There's a trust dimension here too, and it cuts both ways. If a reviewer routinely resolves threads without re-reading the fix, authors learn that "resolved" is just something you eventually do to comments, not a signal that carries any information. Once that happens, everyone starts having the real conversation somewhere else — a chat message, a stand-up aside, a hallway comment — because the tool that's supposed to hold the record of the decision has stopped being where decisions actually get made. That's a worse outcome than a slightly slower review, because now the reasoning behind a change lives in nobody's memory and nowhere searchable, and the next person who touches that code has to reconstruct context that used to be sitting right there in the thread.

None of this is an argument against status flags, pipelines, or resolve buttons. They're genuinely useful, precisely because most of the time a resolved thread does mean what it claims to mean, and being able to trust that most of the time is what makes the flag worth having at all. The point is narrower: the flag is evidence, not the conclusion. It tells you where to look, not what you'll find when you get there. Teams that treat it as the conclusion end up optimising for the appearance of readiness, because that's the thing being measured, while teams that treat it as a pointer keep optimising for the readiness itself, because that's what they actually check.

The practical habit is small and easy to describe, even if it's easy to skip under deadline pressure: before you trust a green state, ask what specific evidence would have to be true for that state to be honest, and glance at whether that evidence actually exists. For a resolved review thread, the evidence is that the change matches the concern, not that the concern no longer has an open marker next to it. For a status field, the evidence is that someone with the full picture looked at the current state and confirmed it, not that no automation happened to disagree. That one habit — treating the flag as a question rather than an answer — is most of the difference between a team that trusts its own tooling and a team that's quietly learned not to.

Matthew Ratcliffe, software developer and architect, Ballarat
Senior Software Engineer & Architect

20+ years across the technology stack — from greenfield builds to brownfield rescues. Based in Ballarat, VIC, focused on AI, healthcare and high-risk data systems. Full resume →

Share your thoughts