← all articles

Make Bad Input Impossible, Not Just Unlikely

Make Bad Input Impossible, Not Just Unlikely

The safest way to stop a bad value from causing damage is to make the system incapable of accepting it, rather than trusting that someone will notice it went in wrong. That sounds obvious stated plainly. It is much less obvious in the moment, because the alternative — a loose convention that "usually works" — is faster to build, easier to explain, and passes every test you happen to think of writing. The problem only shows up later, in someone else's pipeline, dashboard or decision, once nobody is looking at the part that quietly misbehaved.

I watched this play out this week in a piece of shared build infrastructure a colleague was reworking — the kind of component dozens of other projects include without thinking much about it, the way you don't think much about the electricity in your house until it does something strange. The original version configured itself through a handful of loosely typed flags, string-matched at runtime. One controlled whether a build step should skip rebuilding a solution it had already built. Type the flag correctly and set it to the literal string "true", and it worked. Misspell the flag, and nothing complained — it was simply treated as an unrecognised setting and ignored, so the step quietly did the opposite of what was asked. Set it to "yes" or "1" instead of "true", and the same thing happened: accepted without complaint, interpreted as the default. Every failure mode here has the same shape. The system did not reject the bad input. It accepted it, decided what it probably meant, and moved on. The build stayed green. The person who mistyped the flag had no reason to think anything had gone wrong, because nothing told them it had.

The fix wasn't a smarter runtime check. It was removing the ambiguity at the boundary. The rework replaced the loose flags with a typed, declared input schema — the build tool itself now validates what a consumer is allowed to pass before the pipeline is even created. A misspelled input isn't silently ignored; it simply isn't a valid input, so the pipeline refuses to start. A value outside the accepted set is rejected rather than interpreted charitably. The failure moved from "quiet and downstream" to "loud and immediate," which is a strictly better place for a failure to live. Nobody enjoys a red pipeline. Everybody prefers it to a green one that lied.

That distinction — reject what you don't understand, rather than guess at what was probably meant — is worth generalising past build tooling. It's the same argument for schema validation on an API boundary instead of quietly coercing whatever arrives. It's the same argument for a strict parser over a lenient one, for a config loader that refuses unknown keys instead of ignoring them, for a feature flag service that errors on an unrecognised name instead of falling back to off. In every one of these cases, the lenient version looks like better developer experience right up until it isn't — right up until someone depends on behaviour nobody actually intended, and now you can't tighten the rule without breaking them. Leniency doesn't remove the cost of a mistake. It just relocates the cost to whoever discovers it later, usually with less context than the person who could have caught it at the boundary.

There's a second, quieter version of the same failure that I think is actually more dangerous, and it showed up in the same piece of work. The build step in question reports a code coverage percentage, extracted from its output with a regular expression. The expression had a subtle bug: for a result like "35.7%", it was capturing only the fragment after the decimal point, so the recorded figure was "7.0" instead of "35.7". That's a wrong number, not a missing one, and wrong numbers are worse. A missing number gets noticed — someone sees a blank cell and asks about it. A wrong number gets trusted. It looks exactly like a real observation, so it gets typed into a spreadsheet, promoted into an "approved" column, and used as the basis for the next decision, all without anyone re-checking the thing it was supposed to represent. In this case, the fix had already landed for every other project using the shared component, but the register that tracked coverage across the business still listed one project's old, wrong figure as verified — because the row had been marked done from a pipeline run that predated the fix, and nobody had gone back to confirm the number against a fresh run. Structurally correct process, factually wrong content. It took a second, careful reviewer actually tracing each row back to its source pipeline to catch it.

The lesson isn't "always double check everything," which is true but useless as advice — nobody has time to re-verify every number by hand, and a business that tried would grind to a halt. The useful lesson is narrower: know which numbers in your systems are self-correcting and which ones are trust-and-forget. A missing value, a failed build, a broken link — these announce themselves and get fixed because someone trips over them. A silently wrong value that looks perfectly plausible does not announce itself. It survives exactly as long as nobody has a specific reason to distrust it, which in a busy team can be a very long time. If a number is going to feed a decision — a coverage threshold gating a release, a metric determining whether a feature ships, a dashboard someone quotes to a client — it deserves either an automated check against its own source, or a standing habit of tracing at least the important entries back to raw evidence before they're treated as settled.

The part of this I found genuinely encouraging wasn't the schema or the regex fix — it was how the review handled a smaller edge case that came up along the way. Two related build components, one for Linux and one for Windows, both defaulted to the same internal job name. Include both together in one pipeline without explicitly renaming one, and GitLab would silently merge them into a single job, quietly dropping test coverage for whichever platform lost. The first instinct, reasonably, was to fix it with documentation — a comment warning future consumers to set the name explicitly. The reviewer pushed for something stronger: don't warn people about a landmine, move the landmine. The Windows component's default name was changed so the collision literally cannot happen unless someone goes out of their way to force it. A comment only helps the person who reads it before they get hurt. A structural fix helps everyone, including the people who never read anything.

That's the through-line worth carrying into other systems: a warning is a bet that someone will read it at exactly the right moment. A structural constraint isn't a bet at all. When you're building something other people will depend on without necessarily understanding it deeply — and most of what any of us build eventually falls into that category — the higher-leverage question isn't "have I explained this clearly enough." It's "have I made the wrong version of this impossible, or merely unlikely." Those look similar from the outside. They behave very differently the day someone, in good faith, gets it wrong anyway.

Matthew Ratcliffe, software developer and architect, Ballarat
Senior Software Engineer & Architect

20+ years across the technology stack — from greenfield builds to brownfield rescues. Based in Ballarat, VIC, focused on AI, healthcare and high-risk data systems. Full resume →

Share your thoughts