← all articles

Emulation Is Not Evidence

Emulation Is Not Evidence

When an AI assistant tells you it has reproduced a bug, that is a hypothesis, not a finding. Treat it as one right up until someone watches the failure happen on the same hardware a customer would actually be holding. I was reminded of this again this week, watching a team chase a device-specific defect that an AI tool had apparently recreated with impressive fidelity, right up until nobody could make it happen on the physical device it was supposedly describing.

The pattern is worth naming because it is going to happen to every team using these tools, probably repeatedly. A defect had shown up on a ruggedised handheld scanner used out in the field, the kind of device with its own firmware quirks, its own scanning stack, its own way of handling interrupts and background processes that no laptop or emulator replicates faithfully. Someone fed the symptoms to an AI coding assistant and asked it to figure out what was going on. It came back not just with an explanation but with something that looked like a working reproduction of the failure. The reaction in the room was somewhere between delight and unease: "this thing is getting too good."

That unease was the correct instinct, and it is worth being precise about why. A language model asked to reproduce a bug is, at bottom, being asked to construct a plausible story that fits the symptoms it has been given. Sometimes that story is exactly right. Sometimes it is a very well-written piece of fiction that happens to satisfy every constraint you handed it, without touching the actual mechanism. The tell is usually in what's easy to imitate versus what's hard. Application logic, timing sequences, state transitions, API call orders, all of that a model can construct convincingly because it lives in the same medium the model was trained on: text describing software behaviour. Driver-level quirks, firmware version differences, the exact way a particular scanner engine handles a malformed barcode, the physical timing of a hardware interrupt, none of that is something a model has actually observed. It can only infer it from descriptions, and inference dressed up in confident prose is very easy to mistake for verification.

That's exactly what happened here. The reproduction had been built and tested in a browser tab, standing in for the device. It looked complete. It matched the reported symptoms. And when someone went and ran the same sequence on the real handheld, nothing happened. The bug that had just been "replicated" wasn't there. What had actually been reproduced was a plausible narrative that fit the available evidence in an environment that shared almost nothing with the one where the original defect lived.

This is not really a story about AI being unreliable. The tool had done something genuinely useful: it had generated a coherent, testable theory quickly, faster than a human would have pieced the same theory together from a bug report and a stack trace. The mistake would have been treating that theory as the finding itself instead of as the first step toward one. The team didn't make that mistake, because someone in the room had the reflex to say "I don't trust it" and go and check on real hardware before anyone started writing a fix for a bug that, as far as the evidence now showed, might not exist in the form described.

That reflex is worth protecting deliberately, because it runs against the grain of how these tools make you feel. Fluency reads as competence. A confident, well-structured explanation feels like evidence, especially when it arrives faster than you could have produced it yourself and uses all the right terminology. But confidence and correctness are produced by different processes, and nothing about a model's tone tells you which one you're looking at. The more capable these tools get at sounding authoritative, the more deliberate you have to be about separating "this is a good explanation" from "this is a confirmed cause."

The practical rule I've settled on is simple enough to actually survive contact with a busy sprint: any diagnosis or reproduction that touches hardware, firmware, network conditions, or anything else outside the software's own control gets one non-negotiable step before it's believed, which is running it on the real thing. Not the emulator, not the browser tab standing in for the device, not the developer's laptop that happens to share an operating system with production. The actual device, in conditions as close to the field as you can manage. If it reproduces there, you have a bug. If it doesn't, you have a theory that needs revising, and you've saved yourself from shipping a fix for a problem that never existed while the real one keeps quietly happening to whoever is holding the actual scanner.

There's a broader version of this worth sitting with too. As these tools get better at constructing plausible accounts of systems they can't directly observe, the discipline of checking against reality doesn't become less important, it becomes the whole job. The parts of debugging that used to take the most time, forming a hypothesis, writing exploratory code, tracing through logic, are getting faster and cheaper. The part that hasn't changed at all is the part where you go and look at what the system actually does. That step was never really about intelligence. It was always about contact with the real thing, and no amount of capability on the other end changes what it takes to make that contact.

Matthew Ratcliffe, software developer and architect, Ballarat
Senior Software Engineer & Architect

20+ years across the technology stack — from greenfield builds to brownfield rescues. Based in Ballarat, VIC, focused on AI, healthcare and high-risk data systems. Full resume →

Share your thoughts