← all articles

Your Machine Has Too Much Installed To Test That Fix

Your Machine Has Too Much Installed To Test That Fix

If a build fails because something is missing, don't prove the fix on a machine where nothing is missing. It will pass, you will feel good about it, and the real failure will be waiting for you on the build server. The environment that verifies a fix has to be at least as poor as the one that broke. A developer laptop is almost always the richest environment in the whole delivery chain, which makes it the worst possible judge of a missing-dependency bug.

I got to relearn this recently on a mobile app that targets both Android and iOS from a single project. The setup is common: one project file lists two target platforms, and the build pipeline runs one job per platform. The Android job runs on a Linux machine with only the Android tooling installed. The iOS job runs on a Mac with only the iOS tooling installed. Sensible, cheap, and exactly how you'd want it.

Then both jobs failed, each complaining about the other platform's tooling. The Android runner was asking for an iOS component it had no reason to own, and the Mac was asking for Android.

The plausible fix that wasn't

The first diagnosis was reasonable. The restore step was processing the whole project, so it was resolving both platforms. The build command already accepted a flag naming a single platform, so the obvious move was to drop the separate restore and let the scoped build do its own. I made the change, ran the build locally with that flag, and it worked cleanly. No warnings, no errors, and a satisfying description of why it should work.

It failed identically in the pipeline.

The explanation is that the flag narrows what gets built, not what gets read. Before any compilation, the build tooling evaluates the project file, and it evaluates every platform listed in it, including the one you've told it to ignore. That evaluation tries to load the tooling for each platform, and that is where the missing iOS component caused trouble on a machine that never needed to build for iOS. The flag was scoping the output. The failure lived in the input.

None of that is exotic once you've seen it. What interests me more is why I didn't see it sooner, and the answer is embarrassingly simple. My laptop has both sets of tooling installed. Every step of the evaluation I thought I was testing found exactly what it was looking for, so the fix looked correct because the one thing the pipeline lacked was the one thing my machine had in abundance. I had written a fix for an absence and tested it somewhere nothing was absent.

The test environment is part of the claim

This is a different problem from the usual "it works on my machine". That phrase usually means the machine has something the target doesn't, and you find out when you deploy. Here the machine did have more, and I knew it, but I still treated a local pass as evidence. The reasoning felt airtight, which is the dangerous part. A clear, logical explanation of why a change should work is not the same thing as watching it work somewhere it could have failed.

A useful question before trusting any verification is: what would this test have done if I were wrong? If the honest answer is "passed anyway", it isn't telling you anything. A build that succeeds on a fully loaded machine would have succeeded whether or not the fix did anything, so the pass carried no information about the thing I cared about.

For a missing-dependency problem, that points to a fairly practical rule. Reproduce the poverty first. Run the failing job again, or something equivalent, in an environment that lacks what the failing one lacks: a clean container, a fresh runner, a machine you've deliberately stripped back. If the original error appears there before your change and disappears after it, you have evidence. If you can only show the pass, you have a story.

What the real fix looked like

Once I accepted the pipeline as the authority, the actual fix was small. Instead of asking the build to ignore a platform, I made the list of platforms itself overridable by a custom property, with the full list as the default. Each pipeline job then passes in just its own platform, so the other one is never in the list for tooling to trip over. The name matters too. The obvious choice, overriding the standard platform property directly, flows into every referenced project and breaks the ones that were fine. A property with a name only this project reads stays out of everyone else's way.

The check this time was different. I asked the tooling to report the platform list it would use with the override applied and confirmed it showed one entry rather than two. Then I let the real pipeline run on both runners and watched each go green. The Mac build couldn't be verified on my own machine at all, which in hindsight was a feature, not a limitation. The only way to test it was the poor environment.

Why this is more than a build-system quirk

The commercial cost here is small but real, and it's the sort that accumulates. Each wrong guess cost a pipeline run, some waiting, and a round of confidence that had to be taken back. In a larger team, a fix that is merged on the strength of a local pass produces a particular kind of damage: everyone downstream assumes the problem is solved, so the next failure is investigated as something new. The false pass doesn't just waste a cycle. It misleads the people reading the history later.

The same shape appears well outside tooling. A permission fix tested with an administrator account. A data-handling change checked against a clean database that has none of the awkward legacy rows. An offline feature tried on an office network with perfect signal. A security control assessed by someone who already holds every entitlement. In each case the person testing is, by virtue of building the thing, the richest user the system will ever have, and the people who'll actually live with it are mostly not that person.

I've come to treat my own machine as a convenience for writing code and a poor witness for proving it. The place where the failure was reported is the place where the fix has to be shown to work, and the cheapest way to honour that is to make the test environment deliberately less capable than the one I'd naturally reach for. It's slightly humbling to rely on the less comfortable machine, but it's the one that tells the truth.

Matthew Ratcliffe, software developer and architect, Ballarat
Senior Software Engineer & Architect

20+ years across the technology stack — from greenfield builds to brownfield rescues. Based in Ballarat, VIC, focused on AI, healthcare and high-risk data systems. Full resume →

Share your thoughts