A Retry Needs Something New To Try

A retry is only worth having if the second attempt can differ from the first. If you send the same request, to the same model, with settings that make it nearly deterministic, you are not building resilience. You are asking the same question again and hoping for a different answer. When a long-running AI pipeline stalls on the same item every time, the fix is rarely a bigger retry count. It is to work out what the failure is telling you, keep whatever the failed attempt did manage to finish, and change something about the next attempt.
I ran into this recently while feeding long documents through a small, fast language model to divide them into sections. Most of the document got through. Then the job stalled a few parts short of the end with a polite message saying the model had answered but its output couldn't be used. Retrying did nothing, and that was the clue. With the randomness turned right down, which is exactly what you want for structured work, the same input produced the same oversized reply on every attempt. Three retries and thirty retries would have ended identically.
What the failure was actually saying
The cause was mundane once I looked at the replies rather than the error. The small model was starting a new section on almost every line. That is a perfectly understandable misreading of the instruction, but it meant some replies were long enough to hit the output limit and get cut off mid-structure. A truncated reply doesn't parse, so the pipeline threw the whole thing away, including the many sections that had been completed perfectly well before the cut.
That is the first lesson: a reply that was cut off is not the same as a reply that was wrong. Everything before the cut was valid. Discarding it treats a partial success as a total failure, and it forces you to pay again for work you already had. Once the parser kept the sections that were finished and only dropped the dangling one, the oversized replies stopped being fatal.
The second lesson was less comfortable. Even the replies that did parse weren't good. They produced one-sentence passages under headings copied straight from the source text. Nothing had errored, so nothing looked wrong in the logs, but the output was thin and unhelpful to whoever would eventually read it. A pipeline that only measures whether the call succeeded will happily report a green run full of mediocre results.
Change the attempt, not the count
The most useful change was also the least clever: stop asking the small model to do the job it was clearly bad at. Splitting documents into sections needs a sense of where an idea starts and finishes, and that judgement is closer to what the larger model does well. Routing that one task to the bigger model, through configuration rather than code, did more than any amount of prompt tuning on the small one. The small model stayed where it was cheap and good enough.
This is the commercial reality of mixing models. The fast one is cheaper and quicker, and using it everywhere is tempting. But the cost that matters is the cost per usable result, not per call. A cheap call that fails on every retry, or succeeds and gives you something you have to redo, is the expensive option.
I also added a modest guard on the output side. A section that comes back shorter than a handful of lines is merged into the one before it, unless it genuinely marks a new chapter or a change in what's being kept versus skipped. That second clause matters. A blunt rule like "merge anything short" would have quietly glued together things that were short for a good reason. The rule needed to know why a boundary might be real.
Make the fix cheap to apply to the past
The last piece is the one I'd put on any list of things to do before shipping an AI pipeline. The model's answers had been saved. When the splitting logic changed, I didn't want to send every finished document back through the model and pay for it all again. So the logic was versioned, and anything already processed was re-packed from the saved answers. Nothing went back to the model.
That only works if you treat the model's raw output as an asset in its own right, separate from how your code later interprets it. Interpretation will change. Parsers get smarter, merge rules get refined, and you will want to apply those improvements to history. If the only thing you kept was the final tidy result, every improvement means starting again.
What I'd carry forward
When a step fails the same way repeatedly, treat the repetition as information rather than bad luck. Ask what a second attempt could change: the model, the size of the request, the instruction, the way you read the reply. If the honest answer is nothing, a retry is just delay with a log entry.
Keep what a failed attempt did finish. Measure whether the output is useful, not only whether it arrived. And keep the raw answers, because the day you improve your understanding of them will come sooner than you expect.
None of this is specific to one model or one task. It is the same instinct that makes a good engineer look at a stuck job and ask what it's trying to say, rather than simply pressing the button again.


Share your thoughts