A 401 Is Better News Than It Sounds

After an infrastructure change, the most useful thing you can do is compare how each service fails, not how badly. A request that comes back "unauthorised" is a service telling you the whole path to it is intact. A request that comes back "not found" from the same front door is telling you something quite different. If you read those two responses side by side before you touch anything, you can often shrink a frightening outage into a single misbehaving rule in a few minutes.
I watched this happen this week. We had applied a batch of changes to the networking layer sitting in front of several services in a staging environment, and shortly afterwards something returned a 404. My first instinct, and I suspect most people's, was that the change had broken traffic across the board. That is the natural reading. Changes to shared plumbing touch everything, and a failing request feels like a failing platform.
Then someone sent a request to another service behind the same entry point and got a 401. Strictly speaking that is also a failure. The request was refused. But it is a failure of a very particular kind, because to refuse you for lacking credentials, the application has to be reached, has to be running, and has to be sophisticated enough to ask who you are. A 401 is the system working. It had never heard of me, and it said so politely.
That one response redrew the problem. It was no longer "did our change break the load balancer?" It became "why does one route behave differently from its neighbours?" The other service stayed healthy. The shared infrastructure was plainly passing traffic through. Only one path consistently produced a 404, the same answer every time.
Why the difference matters more than the failure
A 404 from a balancer in front of an application is ambiguous in a way people underrate. It can mean the application was reached and has no such resource. It can also mean the request never found a home, because no rule matched it, or because a rule matched in the wrong order, or because a condition on a host name or path no longer described the traffic that was arriving. Those are very different faults, and they live in different places.
What narrowed it for us was the consistency. A fault in the application or its network path tends to be intermittent, or slow, or shaped by load. A request that is refused the same way, instantly, every time, smells like configuration: a decision made deterministically by something that was told, wrongly or incompletely, what to do. Deterministic problems are the good kind. They are the ones you can find by reading rather than guessing.
So the work that followed was dull in the best way: look at the rule for that one route, look at the condition that decides which requests it claims, look at its priority relative to the others, and look at where it forwards. We did not need to reboot anything, roll anything back wholesale, or convene a larger meeting. We needed to look at a short list of things one level down.
The cost of the wrong first reading
There is a commercial angle here that I think gets lost. If we had treated this as a platform-wide break, the sensible-sounding response would have been to revert the whole change. Reverting is not free. It undoes the parts that were working, it burns a deployment cycle, it makes the next attempt harder to attribute, and it quietly teaches the team that touching shared infrastructure is dangerous. That last lesson is the expensive one. Teams that come to fear their own plumbing change it less often, batch more into each change, and make every change riskier. It becomes a loop that feeds itself.
Narrowing first costs almost nothing and pays off whichever way the investigation goes. Even if the cause had turned out to be bigger than one rule, we would have known it by evidence rather than by mood.
A habit worth building
I have come to prefer a simple discipline when something breaks after a change: before forming a theory, gather the same probe from several places and look for the one that disagrees. Same request, different service. Same service, different route. Same route, different environment. You are not trying to explain anything yet. You are trying to find the smallest thing that is different from its healthy neighbours, because the cause is almost always sitting in that difference.
It also helps to treat the status codes as vocabulary rather than as pass or fail. People tend to sort responses into green and red, and in doing so throw away most of the information. A 401 and a 404 are not both just "red". One says the door is there, locked, and it knows you are knocking. The other says there may be no door at all, or that you are at the wrong address. A timeout says something else again. Learning to hear the difference is cheap and it compounds.
I should be honest about the limits of this. A healthy neighbour proves less than it feels like it does. The route that returned 401 shares the entry point but not necessarily every setting, so it narrows the search rather than ending it, and we still had to confirm our reading by actually inspecting the faulty rule and re-testing after any correction. A good clue is not a diagnosis. It is a reason to look in a smaller place.
But that is exactly why I like it. Under pressure, with a change freshly applied and everyone half-convinced they broke everything, the instinct is to widen the search and act fast. The better move is usually the opposite: take the one response that doesn't match the rest and ask what it knows that the others don't. More often than not, the good news was sitting in the error message the whole time.


Share your thoughts