Hardening Is Only A Claim Until You Deploy It

Hardening you have never run in production-shaped conditions is a claim, not a control. Strip a container down to the least privilege it needs, ship a Release build that behaves differently from the Debug build you test on, and you have made two bets at once. Neither bet is wrong. Both should be placed early, on purpose, while a failure is cheap, instead of arriving together on the day you most want a quiet deployment.
I was reminded of this by a rollout that went red in the most thorough way possible. Every service in the environment crashed on startup. That sounds like one big problem, and the first instinct is to hunt for the single change that broke everything. There wasn't one. There were three separate causes, in three separate places, and the only thing they shared was that nothing had ever run the system in the shape it was about to be shipped.
The build you test is not the build you ship
The first failure was in the application itself. During development, the framework generated some of its own plumbing at startup, compiling it on the fly. That is convenient, and it is why the feature worked on every developer machine. The Release image was built differently: the plumbing was generated ahead of time, and the heavyweight compiler dependency that does the on-the-fly work was deliberately left out, because a production image has no business carrying it.
Both decisions were sensible. The gap was a single setting that told the application which mode to use. In development the default happened to match reality. In the shipped image it did not, so the application went looking for a compiler that was never packaged and fell over on boot.
The uncomfortable part is that nothing here was careless. Each choice, taken alone, is what you would want: fast iteration locally, a lean and ahead-of-time image for production. The defect lived in the seam between them, and seams are exactly what a test run on the developer-shaped build cannot see. Green tests told us the logic was right. They said nothing about whether the artefact we were about to deploy could start.
Least privilege is a hypothesis about someone else's software
The other two failures were the same story in a different costume. Both services run stock web-server images, and the deployment configuration dropped every Linux capability from the container, which is a good default. Start from nothing and add back only what you can justify.
The trouble is that "what you can justify" requires knowing what the software actually does when it starts. The stock images do things at startup that are entirely reasonable and entirely invisible from the outside: change ownership of cache directories, switch to an unprivileged user, bind to a low port. Without the capabilities to do those things, one server refused to start with a permissions error on a cache directory, and the other exited with a bare "operation not permitted" before it had said anything useful.
The fix was to add back a short list of capabilities to each. The application we wrote ourselves, an unprivileged listener on a high port, needed none of them and kept the strict setting. That contrast is the interesting bit. The code we understood best needed the least, and the software we had only ever treated as a black box was the thing asking for more than we had assumed.
I don't read that as a reason to loosen the baseline. I read it as the baseline doing its job, because it forced the question of what each component really needs. The mistake was not locking things down. The mistake was treating the lock-down as finished when it was written, instead of when it had been exercised.
Why the cost is not evenly spread
There is a commercial argument here that is easy to miss. A failed rollout in a pre-production environment costs an afternoon and a slightly bruised ego. The identical failure discovered during a customer-facing release costs trust, a rushed fix, and a decision made under pressure about whether to roll back or push forward. The defect is the same size in both cases. The price is not.
That is the whole case for rehearsing deployment early and often, and it is why I have become wary of anything labelled hardening that lives only in a configuration file. A security posture that has never been deployed is a document. It might be an excellent document, but its accuracy is untested, and the people who find out are whoever happens to be on call.
There is also a quieter human cost. Three simultaneous crash loops produce a particular kind of fog. When everything is failing, the temptation is to assume a common cause and to start undoing things. Separating the three failures, reading each one's actual error, and fixing them independently was calmer and faster than guessing. The habit that helped was refusing to let one dramatic symptom stand in for one cause.
What I would do differently, and what I would keep
I would run the production-shaped artefact, with production-shaped restrictions, as an ordinary part of the build, long before a release depends on it. Not a full deployment ceremony, just enough that "does this image start under the rules it will run under" has an answer that was written down before it mattered. If that check had existed, all three problems would have been found by a machine, quietly, with no one watching.
I would also keep the strict defaults. Dropping every privilege and adding back only what you can name is a better position than the reverse, because each addition becomes a small, reviewable decision with a reason attached. The list of capabilities we added back is now short, specific, and explained. That is worth more than a permissive default nobody remembers choosing.
And I would write down what each crash taught us, in the place the next person will look. Every failure in a rollout is a sentence missing from your documentation. The cheapest time to add it is while the error message is still on the screen.
None of this is glamorous. It is the unremarkable discipline of making sure the thing you describe is the thing that runs. But when people trust you with a system, that gap between description and reality is where most of the surprises live. Closing it early is a kindness to everyone who has to deploy, support, or depend on what you build.


Share your thoughts