A Rolling Deploy Needs Room To Roll

When a deployment fails on a timeout but the release turns out to have worked, don't just raise the timeout and move on. The timeout is telling you something about headroom. A rolling deploy briefly runs the old version and the new version side by side, so it needs spare capacity that normal running never uses. If the cluster has none, the platform has to go and make some, and that takes far longer than your pipeline was prepared to wait. Fix both halves: give the swap room to happen, and make the wait honest about how long it can really take.
The failure that wasn't one
I recently chased a deploy that kept failing at the step that waits for a rollout to finish. The error said the old replicas were pending termination, and then that the wait had timed out. Read literally, that sounds like the old version refusing to die. You start looking at shutdown hooks, stuck connections and drain behaviour.
None of that was the problem. The old instance was doing exactly what it should: staying alive until its replacement was ready to take over. The replacement was the one that was late. The message named the wrong party, which is a small but real way a tool can send you off in the wrong direction. It describes what the system is waiting to do, not what is blocking it.
The replacement was late because it had nowhere to land. The existing machines were already carrying most of their memory in reservations, so the extra instance the rollout needed at the peak of the swap didn't fit. The autoscaler noticed, and went off to add a machine. That took about three minutes before the new instance could even begin fetching its image. The whole rollout, start to finish, ran a little over three minutes. The pipeline had been told to wait two.
So the release succeeded and the pipeline reported failure. Both were telling the truth about different things, and that is the worst kind of failure to live with, because the person who sees the red tick has to decide whether to trust it.
Headroom is part of the deploy design
It's easy to think of capacity as a runtime concern and deployment as a separate, procedural one. A rolling update is where they collide. At the moment of the swap you are paying for two copies of something, and the cluster has to be able to afford that without help.
The reservation was the lever. Each instance had asked the scheduler to set aside a generous amount of memory, and that ask is what the scheduler uses to decide whether something fits, regardless of what the process actually consumes. Halving the request meant the extra instance could fit on the machines already running, so the swap no longer had to wait for new hardware. The ceiling on what the process was allowed to use stayed exactly where it was, which matters. A request is a promise about planning, a limit is a promise about safety, and it's worth being deliberate about which one you're changing.
I'd also raise the timeout, but for a different reason than "make the red go away". Even with headroom restored, there will be days when the platform legitimately needs to add a machine: after a scale-down, during a busy period, when someone else's workload has taken the space. A wait that's only long enough for the happy case turns the occasional ordinary event into a failed release. The honest timeout is the worst realistic case plus a margin, not the median.
What I hadn't measured
Here's the uncomfortable part. I lowered a memory number without knowing how much memory the service really uses. The cluster had no metrics collection installed, so there was no usage history to read. The new figure was a reasoned reduction, not a measured one.
I think that's acceptable only if you say so out loud, and say what would prove you wrong. If the new request is too low, the symptom is instances being evicted under memory pressure, and the response is to raise the request to match real use. Writing that down next to the change matters more than it sounds. A guess that's labelled as a guess has an owner and a tripwire. A guess that looks like a decision gets inherited by whoever touches the file next, who has every reason to assume someone checked.
This is where the commercial reality sits too. Over-reserving resources is cheap insurance that gets quietly expensive: more machines, a larger bill, and as this case showed, slower and more fragile deploys. Under-reserving is cheap right up until it isn't. Measuring is the way out of that trade-off, and installing the thing that measures is usually a smaller job than the arguments about what the numbers probably are.
The pattern worth keeping
A few habits fall out of this, and they apply well beyond one cluster.
Treat a timeout as a claim about how long something takes, and check that claim against the slowest thing the system is allowed to do, not the typical thing. Treat a failed deploy of a working release as a signal about the environment, not noise to be retried. Be suspicious of error messages that describe intent rather than cause, and go looking for what the waiting party is actually waiting on. And when you change a number you can't measure, record that you couldn't, and what you'd expect to see if you were wrong.
None of this is glamorous. Nobody's day is improved by a deploy that quietly takes three minutes instead of one. But the people on the other end of a red pipeline are making a decision at the worst moment: do I roll back, do I retry, do I wake someone up? A deploy that reports its own outcome accurately is one of the cheaper ways to spare them that decision, and a team that trusts its pipeline ships more calmly because of it.


Share your thoughts