← all articles

Every Wait Needs A Way Out

Every Wait Needs A Way Out

Every spinner you ship is a promise that something will end the wait, and most spinners are built as though success is the only thing that can keep that promise. If a loading screen can only be dismissed by the happy path, it will eventually be shown to someone for whom the happy path never arrives. The fix is not a cleverer loading screen. It is deciding, before you write the wait, every way it is allowed to finish: success, failure, timeout, cancellation, and "something I did not predict". A wait with only one exit is a trap with a nice animation.

The shape of the bug

The pattern I keep seeing lives in startup gates. An application, on first load, needs to answer a question before it shows anything useful. Is there a session? Is the connection ready? Has the stored state been read? While that question is open, the user sees a spinner. Once it is answered, the gate either lets them through or sends them somewhere else.

That sounds sensible until you look at what the gate actually does. In the version I have been thinking about, the check ran once, when the component first appeared. If the gate decided the user needed to be redirected, the redirect happened inside the application, which meant the same gate instance stayed mounted, still showing its spinner, and never ran its check again. The user landed on the correct page in the address bar and the wrong state on screen. A second refresh fixed it, because a full reload created a fresh gate. That is the worst kind of intermittent fault: it goes away when you try the obvious thing, so everyone learns to refresh and nobody files a bug.

The second flaw was quieter. The calls the gate made had no time limit, and the error handling caught the failures the authors had imagined: the API returning an error, the network returning a bad status. It did not catch cancellation, and it did not catch a call that simply never answered. Those are not exotic. A request can stall behind a flaky connection. A browser can abort work when the user navigates. Reading a stored session can fail for reasons that have nothing to do with the application. In each case the code did not take the success path and did not take the failure path either. It just stopped, with the spinner still up.

Why this keeps happening

Nobody writes a spinner thinking, "I would like this to hang forever." It happens because we write the code in the order we think about it. We imagine the request, we imagine it succeeding, we add handling for the errors the library documents, and the spinner comes down at the end of that story. The states outside the story are not so much decided against as never decided at all.

There is also a commercial reason it survives. A hang that clears on refresh costs the business almost nothing that it can measure. It does not throw an exception in the logs, because nothing threw. It does not show as an error rate, because nothing errored. It shows up as a customer who gave up, or a support message that says "it was just stuck", which is very hard to reproduce and therefore very easy to defer. Silence is not evidence that the wait is safe. It can be evidence that the wait is failing in the one way your instrumentation cannot see.

And the people who pay are not the ones who wrote it. Someone opens the application to do something that matters to them, and what they get is a screen that neither works nor explains itself. They cannot tell whether to wait, refresh, or assume they have been locked out. Uncertainty is the real cost here. A clear error would be more honest than an indefinite spinner, because a clear error tells a person what to do next.

Design the exits first

The practical habit I have come to prefer is to write down the exits before the happy path. For any wait that blocks the user, ask what ends it when the answer arrives, when the answer is no, when the answer never arrives, when the work is cancelled underneath you, and when something fails that you did not anticipate. If you cannot name an exit for each of those, the wait is not finished being designed.

Each of those exits has a straightforward implementation. Put a bound on every call that can stall, so that "never" becomes "not within a reasonable time". Make the timeout generous enough that a slow but healthy connection still succeeds, and short enough that a person has not already given up. Catch cancellation explicitly, because in many runtimes it is a different kind of failure from the errors you already handle and slips straight past a catch block written for the others. And when a gate re-uses the same component across in-application navigation, make it re-check when the location changes, rather than assuming that being mounted once means being correct forever.

Fail towards showing something

The choice that matters most is what the gate does when it does not know. There are two defensible defaults. Fail closed, and the user is blocked until the question is answered. Fail open, and the user is shown the page and the page is left to report its own problems. For a gate that exists only to decide where to send someone, I would lean towards the second. A page that loads and says "I could not confirm your session" is a better experience than a spinner, and it keeps the failure where the person can see it and act on it.

That is not a universal rule. If the gate is protecting something sensitive, failing open is the wrong answer, and the right answer is a clear, visible refusal. The principle is not "always let people through". The principle is that an unknown state should resolve to something visible and honest, not to an indefinite pause. For anything involving entrusted data, a refusal with an explanation is a perfectly good exit. A silent hang is not.

Test the hang, not just the success

The part I find most useful is how this changes testing. A hang is awkward to test precisely because nothing happens. You have to build the situation on purpose: an API that never responds, a stored session that cannot be read, a redirect that happens mid-check, a call that times out. Each of those is cheap to simulate once you decide to, and each one proves a specific exit exists. Without them, the only evidence that your spinner leaves is that it usually does.

I would also be wary of treating "I refreshed it a few times and it was fine" as verification. With an intermittent hang, repeated manual refreshing mostly tests your patience. It tells you the happy path is common, which you already knew. The tests worth having are the ones that force each unhappy path and assert that the user ends up somewhere sensible.

The small version of a bigger habit

It is tempting to file all this under front-end polish, but the same shape appears wherever one thing waits for another: a job queue consumer waiting on a downstream service, a deployment waiting on a health check, a person waiting on an approval that nobody knows they are supposed to give. Anything that waits needs a way to stop waiting. If it does not, the system has quietly handed an open-ended obligation to whoever is looking at it.

So the next time you add a loading state, spend a few minutes on how it ends in the cases you are not thinking about. It is a cheap piece of design, it costs almost nothing against the first support request it prevents, and it is the difference between software that is merely fast when things go well and software people can trust when they do not.

Matthew Ratcliffe, software developer and architect, Ballarat
Senior Software Engineer & Architect

20+ years across the technology stack — from greenfield builds to brownfield rescues. Based in Ballarat, VIC, focused on AI, healthcare and high-risk data systems. Full resume →

Share your thoughts