← all articles

Prove Your Brake Before You Find The Cause

Prove Your Brake Before You Find The Cause

When someone on the other side of an integration phones to say "please stop sending", the first job is not to work out why. It is to prove, with evidence you can point at, that you have actually stopped. Diagnosis can wait ten minutes. Uncertainty about whether the tap is still running cannot, because every minute of that uncertainty is a minute in which the problem may be growing, and the person on the phone is the one living with it.

I've come to treat this as its own step, separate from investigation. Stopping is an action, and like any action it has an outcome that needs checking. "We switched it off" is a claim. "Outbound counts are flat at zero, every sender is disabled, and nothing is waiting in the queue" is evidence. The first feels the same as the second right up until it turns out to be wrong.

What proving the brake looks like

In a recent incident, a downstream partner reported that they were receiving a stream of unexpected inbound traffic from us and asked us to halt it. The natural instinct, especially for engineers, is to open the logs and start theorising. Instead the first pass was boring on purpose: confirm outbound volume was zero, confirm every sending component was disabled, confirm nothing was sitting in the pickup area waiting to go.

All three checks came back clean. That is a useful result even though it sounds like the opposite of a breakthrough, because it changes the question. We were no longer asking "how do we stop this?" We were asking "why does the other side still see traffic that, by our evidence, we are not sending?" Those are completely different investigations, and you can only choose between them once the brake is proven.

It is also where discipline gets tested. The partner was still reporting a few dozen items. Our own archive held a similar number, and when we looked at the timestamps they sat inside a single window of roughly ten minutes, which lined up with a deployment. So the likeliest story was a burst that had already ended, with the other side working through the backlog. That is a hypothesis, though, and I'd rather label it one. Test and production looked identical on the numbers we had, so we honestly did not yet know whether the cause was a mismatch in a schema, a difference in database version, or something in a single field of the payload. We had a candidate list, not a cause.

Why the order matters

There are three reasons I would defend this ordering to anyone.

The first is that it protects the person who called. Their problem is the traffic, not our root cause. A prompt, specific "we've verified we are not sending anything, here is what we can see on our side" is worth far more to them than a confident theory delivered an hour later. It also keeps the conversation honest. You can tell a counterparty exactly what you have verified and exactly what you have not, and you do not need to guess at a cause to be useful.

The second is that unproven brakes poison diagnosis. If you start investigating while some part of the flow might still be live, every observation is contaminated. A count that changes while you are looking at it cannot be reasoned about. Freezing the system into a known state is what makes the evidence stable enough to think with.

The third is the one that is hardest to see from inside the incident: confidence is cheap under pressure. When a partner is waiting and colleagues are asking for an answer, a plausible explanation arrives very quickly, and it is tempting to let it stand in for verification. The explanation might even be right. But a fix built on a plausible explanation and an unproven stop is two untested assumptions stacked together.

Prefer a repair you can run twice

Once the flow was understood well enough to act, the temptation was to reach in and tidy things by hand: move items, delete a few, patch the odd one. Manual surgery on live data is the sort of thing that feels responsible because you are being careful, and then turns out to be the thing nobody can reproduce or explain later.

The plan the team settled on was different. Rather than moving anything, they would re-run an existing, repeatable routine to correct what had been sent, and they would exercise it against a test environment before it went anywhere near production. That has a quiet commercial logic. A repair you can run twice, review afterwards and describe to a partner in a sentence is cheaper than one that depends on someone's memory of what they typed. It also leaves a trail, which matters when the data belongs to people who trusted you to handle it carefully.

The mood in the room mattered too. Someone pointed out that similar work earlier had been painful, and that was treated as a reason to assume nothing would go to plan, not as a reason to feel bad about it. I like that stance. Being slightly pessimistic about the procedure is a good habit when the cost of being wrong lands on somebody else's systems.

The uncomfortable part

None of this is glamorous. Proving a brake produces no insight and no story. The best outcome is a handful of zeros and an empty folder, which look like nothing. There is real pressure, particularly from people watching, to skip straight to the clever part where the cause is found.

I think we sometimes confuse being useful with being fast at diagnosis. In an incident that touches someone else's systems, the useful thing is usually to stop making the problem bigger, show your evidence for that, and tell the truth about what you do not yet know. The cause will still be there after you have done that. The trust you spend by guessing early may not be.

So the habit I'd suggest is small. When asked to stop, stop. Then write down, even if only in the incident thread, what evidence says you have stopped. Only then go looking for why.

Matthew Ratcliffe, software developer and architect, Ballarat
Senior Software Engineer & Architect

20+ years across the technology stack — from greenfield builds to brownfield rescues. Based in Ballarat, VIC, focused on AI, healthcare and high-risk data systems. Full resume →

Share your thoughts