← all articles

The Agent That Knew What Wasn't Its Call

The Agent That Knew What Wasn't Its Call

I watched an automated coding agent do something more useful than writing good code this week: it stopped. Partway through a piece of work, it hit a decision that wasn't its to make, and instead of guessing, or quietly expanding its own mandate to cover the gap, it wrote up exactly what it had found, raised a properly scoped ticket, and left the actual judgement call sitting with the human reviewer. That is a small moment, and I think it's a preview of the skill that will separate useful autonomous systems from dangerous ones over the next few years.

The instinct most people have about agentic tools is that the risk is capability. Can it write correct code, can it reason well, can it avoid hallucinating an API that doesn't exist. Those questions matter, but they're not the ones keeping me up. The uncomfortable part is that a system can be extremely capable and still be untrustworthy, because capability says nothing about scope. A junior engineer who can write flawless code is still dangerous if they don't know which changes need a second opinion. The same is true of a model.

What I saw was closer to what you'd want from a good senior engineer handling an ambiguous ticket. The work had been deliberately deferred earlier for reasons that, on reflection, weren't quite the right reasons — technically defensible, but not the real ones. Rather than silently patching that over, the agent restated the deferral with the actual reasons attached, which is a more honest position even though it's slower and less flattering. It then noticed a second, related problem nearby: same bug class, different part of the system, no safety net of a real pipeline to validate a fix against. Instead of folding that into the same piece of work because it was sitting right there and the temptation to just fix it while it was in scope must have been enormous, it split it out. New ticket, left unassigned so it would go to triage instead of landing on somebody's desk by accident, with the affected areas, the shape of a fix, the reasons it had been deferred, and acceptance criteria that explicitly required a real run through a real pipeline before anyone could call it done.

Then it left a code review thread open. Not because it forgot, and not because closing threads is hard — closing a thread is trivial. It left it open because resolving that particular thread wasn't a decision the agent had the standing to make. That was the reviewer's call, and the agent said so, plainly, rather than assuming that "technically correct" meant "cleared to close."

None of that required brilliance. It required knowing the difference between "I can do this" and "this is mine to decide." That distinction is exactly the one a lot of automation quietly erases, and it erases it in the direction that feels helpful right up until it isn't. A tool that fixes the adjacent bug while it's already in the file looks efficient. A tool that resolves its own review comments looks tidy. A tool that reframes an awkward deferral as a clean success looks impressive in a status update. Every one of those moves trades a small amount of visible tidiness for a real erosion of who actually owns the decision. Multiply that across a team running dozens of these agents concurrently and you get a system that is very fast at producing work nobody quite remembers approving.

I think about this the same way I think about access control in the systems I help design for organisations that hold sensitive data. The safe default when a permission is ambiguous is to deny, not to assume access, because assuming access is the mistake that turns a missing role into an open door. Scope works the same way for autonomous agents. The safe default when authority is ambiguous is to hand the decision back, not to assume the mandate stretches to cover it, because assuming it does is how a code-fixing tool quietly becomes a code-approving one. Nobody designs that outcome on purpose. It accumulates one reasonable-looking shortcut at a time, and by the time someone notices, unwinding it means auditing months of decisions that were never really decisions, just defaults nobody chose.

There's a commercial angle here too, and it's worth being honest about it rather than pretending trust is purely a values question. A team that has to double-check everything an agent touches hasn't actually gained capacity, it's gained a second job: supervising the thing that was supposed to remove work. The value of these tools isn't raw throughput, it's throughput you don't have to re-verify. An agent that occasionally does slightly less — that defers, that splits scope, that leaves something for a human — is worth more than one that always finishes, because "always finishes" is usually a sign that ambiguous calls are getting made silently somewhere in the middle. The version that stops and says "this part isn't mine" is the version you can actually build a process around, because you know where the human checkpoints genuinely are instead of hoping they're wherever the agent happened to leave them.

It also changes what "good ticket" means when an agent, not a person, is writing it. The write-up I saw wasn't a status update dressed as a summary. It carried enough — the affected scope, the shape of a fix, the reasons for deferral stated honestly, the conditions for acceptance — that whoever eventually picked it up wouldn't need to reconstruct the reasoning from scratch. That's the actual bar for handing work back well: not "I stopped," but "I stopped in a way that didn't cost the next person the thinking I'd already done." A lot of human-written handovers don't clear that bar either, so this isn't really an AI standard, it's a professional one that AI happens to make more visible, because when a model does it consistently, the absence of it in a person's handover suddenly looks like what it always was: a shortcut.

None of this needs the tool to be more careful in some abstract sense. It needs the system prompting it, and the team reviewing its output, to have decided in advance which classes of decision are reversible enough to leave to the agent and which ones aren't — the same conversation you'd have onboarding a contractor, except now you can actually write the boundary down and have it enforced every time instead of hoping it's remembered under deadline pressure. Get that boundary right and the agent stopping mid-task stops looking like a limitation and starts looking like the feature that makes the rest of its output usable.

I don't think this scales down to a tidy rule, which is a little unsatisfying to admit in an article that's supposed to land on one. But if there's a version of it worth carrying around, it's this: judge these systems less by how much they finish and more by whether they can tell you, clearly and without prompting, what they didn't. That's a harder thing to fake than good code, and it's the thing that actually determines whether you can hand more work to a system like this next month than you can today.

Matthew Ratcliffe, software developer and architect, Ballarat
Senior Software Engineer & Architect

20+ years across the technology stack — from greenfield builds to brownfield rescues. Based in Ballarat, VIC, focused on AI, healthcare and high-risk data systems. Full resume →

Share your thoughts