← all articles

A Migration Is The Cheapest Time To Ask Why

A Migration Is The Cheapest Time To Ask Why

The best moment to question why a piece of infrastructure exists is not when it breaks and not when someone finally schedules a review. It is the moment you are already forced to touch it, because that is the only time the cost of asking drops below the cost of not asking. Most organisations never get that moment cheaply, which is why so much of what runs their systems was never actually decided by anyone. It simply survived.

I was reminded of this recently while helping a team plan a move from one hosting provider to another. Partway through, someone doing the network planning noticed that the address their public staging environment resolved to was, by coincidence, the same address their VPN gateway had been sitting on for years. Nobody had put those two things next to each other deliberately. Nobody had put them next to each other accidentally either, in the sense of a single bad decision. They had simply both been built at different times, by different people, solving different immediate problems, and the overlap had sat there unnoticed because nothing ever forced anyone to look at the whole picture at once. The migration did that for free. You cannot move infrastructure without describing where everything currently lives, and the act of describing it is often the first time anyone has done so in years.

That is the real value of a migration, and it is usually invisible in the business case. The business case talks about cost, about vendor lock-in, about a newer platform's features. The quieter benefit is that a migration is an enforced audit. You cannot lift something you refuse to look at, so for a short window the incentive to understand a system finally outweighs the incentive to leave it alone.

Most infrastructure complexity does not arrive through a bad decision. It arrives through the absence of one. Someone needed a staging environment quickly, so they cloned an existing one rather than designing a fresh topology. Someone needed a second front end for a slightly different customer, so they stood up a parallel deployment rather than asking whether the same one, driven by the same API, would have done. Each of those calls was reasonable on the day. None of them were reviewed as a set. Years later you have three environments where one would do, two front ends serving identical logic, and a public-facing address that happens to sit on the same subnet as something it should never have been near. Nobody chose that outcome. It accreted.

The uncomfortable part is that removing accreted complexity is almost never worth doing on its own. Telling a team "let's spend two weeks rationalising our environments" is a hard sell, because the return is diffuse and the risk of breaking something that quietly works is concentrated and immediate. Nobody gets promoted for merging staging and test into one environment that used to be two. But when you are already rebuilding the thing from scratch on a new platform, that calculation flips. You are already going to stand up an environment. The marginal cost of asking whether you need one, two, or three of them is close to zero, because you were going to do the work regardless. The only thing that changes is whether you do it thoughtfully or by habit.

I watched exactly that conversation happen in real time on the project I mentioned. The prompt was mundane: someone asked whether the new platform needed two separate front-end deployments, given both were going to call the same API. It is the kind of question that sounds almost too obvious to ask out loud, which is probably why it usually does not get asked. The answer, once someone actually sat with it, was no. The two front ends existed because at some point two environments existed, and each one had picked up a deployment to match, and the pairing had simply been carried forward every time since. Nobody had verified in years whether the split still earned its place. Migrating gave the question a legitimate reason to be asked, and once asked, it answered itself in about a minute.

The same pattern showed up with staging and test. Two environments had existed for long enough that everyone assumed there was a reason, right up until someone asked what the reason actually was and the room went quiet. That silence is worth paying attention to. It is not evidence that the environments are unnecessary; sometimes there is a genuine constraint that nobody has bothered to restate recently. But if a room full of people who work in the system every day cannot immediately produce the reason something exists, that is a strong signal the reason has decayed even if the thing itself has not.

This is where I think a lot of technical debt conversations go wrong. Debt gets treated as something that only lives in code, tracked in a register, paid down through refactoring sprints. Infrastructure debt is the same phenomenon wearing a different coat, and it is often more expensive precisely because it is less visible. Nobody reads infrastructure the way they read a file. It just runs, quietly, until someone has to move it, and only then does its actual shape become obvious. The register approach helps with code because code sits still long enough to be catalogued. Infrastructure needs an event to force the catalogue into existence, and a migration is one of the few events that does that honestly, because you cannot fake having looked at what you are moving.

None of this is an argument for using every migration as an excuse to redesign everything. That instinct is just as dangerous as the opposite one, and I have seen platform moves collapse under their own ambition because someone decided a change of hosting provider was also the moment to rewrite the authentication model, restructure every environment, and rename half the services. A migration earns you the right to ask why something exists. It does not obligate you to change everything whose answer turns out weak. The discipline is in separating the two: ask the question everywhere it is cheap to ask, because you are already there, and act on the answer only where the change is genuinely low risk to make alongside everything else already moving.

What made the environment conversation on that project work was restraint. Nobody proposed collapsing every environment at once, or rewriting the deployment pipeline mid-migration. They asked one specific, answerable question about one specific duplication, got a clear answer, and made a small, contained decision before moving on to the next piece of infrastructure that needed describing anyway. That is a modest way to run a migration, and it is also, in my experience, the only way that actually leaves a system simpler than it started rather than merely relocated.

The coincidence of a staging address sharing a subnet with a VPN gateway was never going to cause an incident on its own. It was a symptom, not a fault. But symptoms like that are exactly what you are paying to surface when you migrate anything properly, and ignoring them because the primary goal was "get it running on the new platform" would have wasted the one moment the question was ever going to be cheap.

Matthew Ratcliffe, software developer and architect, Ballarat
Senior Software Engineer & Architect

20+ years across the technology stack — from greenfield builds to brownfield rescues. Based in Ballarat, VIC, focused on AI, healthcare and high-risk data systems. Full resume →

Share your thoughts