← all articles

A Two-Person Team That Shipped 500 Pull Requests In 14 Days

A Two-Person Team That Shipped 500 Pull Requests In 14 Days

Have you met a dev team of two that has shipped 500 pull requests in 14 days? I have. It's mine. That number isn't a typo or a vanity metric pulled from a dashboard nobody trusts — it's what happens when you stop treating AI coding tools as a single chat window you babysit and start treating them as a fleet of workers you direct, review and hold accountable, the same way you'd run any team.

The honest reason most people don't get anywhere near that throughput isn't the tools. Claude Code and Codex are both good enough now to carry serious, independent work. The reason is that most people still use them one conversation at a time, waiting for each reply before deciding what to do next. That's the equivalent of hiring five contractors and only ever speaking to one at a time while the other four sit idle. The moment you flip that — multiple agents, multiple worktrees, multiple pull requests in flight simultaneously — the ceiling moves somewhere else entirely: your own ability to review, decide and unblock.

I run Claude Code and Codex in parallel, each working a distinct branch in its own worktree so they never trip over each other's changes. I keep a simple board — I use YouTrack for this — so every agent, and every session of me, knows exactly what's next without a conversation to explain it. And I built devStudio specifically to make this manageable: a self-hosted console that runs Claude Code, Codex and OpenCode agents side by side, gives each one a repo and a worktree, and lets me chain them into scheduled workflows instead of manually kicking off every task. None of that replaces judgement. All of it removes the friction that stops judgement from being applied often enough.

The mentality that makes this work matters more than the tooling. It's "let's dive in and get it done" rather than "let's plan this properly first." That sounds reckless until you notice what it actually changes: the cost of being wrong. When an agent can draft a full implementation, tests and a pull request in the time it used to take to write a design doc, the doc stops being the cheap option. Diving in, seeing the shape of the real problem in code, and course-correcting from there is now often faster than reasoning about it in the abstract. Planning still has its place — but it's earned by evidence of actual difficulty, not owed up front to every task by default.

The place this breaks, and where I've broken it myself, is trusting the throughput instead of verifying it. Five hundred pull requests is a count of proposals, not a count of correct code. Every one of them still needs a human who can tell good judgement from confident-sounding nonsense, because agents are extremely good at producing both in the same tone of voice. My review discipline hasn't gotten looser as volume went up — it's gotten more structured, because it has to. I lean on smaller, more frequent pull requests specifically because a 40-line diff is something I can actually hold in my head and check properly, where a 2,000-line one just becomes something I skim and hope about. Velocity without a review bottleneck that scales alongside it isn't velocity — it's technical debt with better marketing.

The other trap is confusing pull request count with value delivered. It's trivially easy to make an agent look busy: more branches, more commits, more churn. None of that matters if it isn't attached to something a customer or the business actually needed. I keep the count honest by tying almost every pull request back to a ticket that describes a real outcome, not a task that exists to give an agent something to do. If I can't say in one sentence why a piece of work matters to someone using the product, it doesn't get picked up, however cheap it would be to build.

What I'd actually recommend, if you want a taste of this rather than the full setup, is smaller than it sounds. Start with two agents instead of one, each on its own branch, working genuinely independent pieces of work — not the same feature split awkwardly in half, which just creates merge conflicts and coordination overhead you didn't have before. Give each one a clear, self-contained brief the way you'd brief a competent contractor who won't ask follow-up questions: what outcome you want, what constraints matter, what "done" looks like. Then hold your review bar exactly where it was before, and let the pull request count tell you whether you've actually gained capacity or just gained noise.

None of this makes engineering judgement optional. If anything it makes it the scarcest resource on the team, because it's now the one thing that doesn't parallelise. What it does do is let two people sustain a pace that used to need a team many times the size, without pretending the code review, the architectural decisions or the responsibility for what ships got any smaller. The team didn't get bigger. The distance between deciding something is worth building and having a reviewable pull request in front of you got a lot shorter. That's the whole trick, and it's genuinely learnable — it just starts with diving in.

Matthew Ratcliffe, software developer and architect, Ballarat
Senior Software Engineer & Architect

20+ years across the technology stack — from greenfield builds to brownfield rescues. Based in Ballarat, VIC, focused on AI, healthcare and high-risk data systems. Full resume →

Share your thoughts