Human Review Is the Bottleneck: Designing Review for Parallel Agents

Run agents in parallel and review becomes the constraint. An operating model, statuses, and review gates: agents move to Ready for review, humans move to Done, with a worked example from ticket to merged PR.

Ashley Innocent

Ashley Innocent

17 September 2026

Human Review Is the Bottleneck: Designing Review for Parallel Agents

The first time you run three coding agents at once, execution stops being the slow part. You hand a bug to Claude Code, a refactor to Codex, and a test-coverage gap to a third Runtime, and forty minutes later three branches are ready. Then the real wait starts, because all three are now waiting for you. You read the first diff, approve it, and the other two sit idle. You have parallel execution and serial review, so the parallelism buys you almost nothing.

This is the trap that catches every team that gets good at running agents in parallel. The bottleneck doesn’t disappear; it moves, from “how fast can an agent write the code” to “how fast can a person read, judge, and accept the result.” Adding a fourth agent makes it worse: a longer queue of finished work, the same single reviewer.

The fix isn’t to review faster or to trust more. It’s to design review as an explicit part of the workflow instead of an afterthought that lives in your head. That’s the workflow this article walks through, and it’s the problem Sharkly is built around: keeping the goal, the diff, the evidence, and the review of every agent run in one Task, so the review queue becomes something you can see and manage instead of a pile of open browser tabs.

Why parallel agents move the bottleneck to review

Review is a queue. Finished agent runs arrive at some rate, a person clears them at some rate, and if arrivals outpace clearances the queue grows without limit. That’s the whole story, and it’s why raw agent throughput is the wrong number to optimize. If ten agents finish in an hour and you can carefully accept three changes an hour, you’ve built a factory whose loading dock is one person with a clipboard.

Most tooling hides this. A stack of pull requests shows you the work, but not which items are genuinely waiting on a decision versus still running versus already merged. The queue is real, but it’s spread across tabs, notifications, and memory. The first design goal is to make it visible, because you cannot manage a bottleneck you can’t see. A companion piece on review capacity math works through what happens as concurrency climbs.

The operating model: who assigns, who runs, who reviews, who accepts

Before you touch statuses, name the four jobs. They are different jobs, and confusing them is what makes review feel chaotic.

Someone assigns the work: turns an intent into a scoped Task with acceptance criteria. Something runs it: an Agent, on a Computer, executing in a Runtime, producing a change and a record of what it did. Someone reviews the result: reads the diff and the evidence, decides whether it meets the criteria. Someone accepts it: makes the call to merge and release, which is a different decision from “the code looks fine.”

Sharkly states the contract in one line worth keeping on a sticky note: agents research, execute, test, and report; people set direction, grant authority, and accept the result. Automation stops where team judgment is required. Assign and accept are human by design. Run is the Agent’s. Review is the job that has to be engineered, because it’s where the queue forms.

A useful boundary: a merged PR is not the same thing as an accepted result. Merging is a git operation. Acceptance is a person deciding the change does what the Task asked. When those two collapse into one click, review quality drops to whatever the reviewer had energy for that minute.

Roles, statuses, and review gates

Give the queue a shape by giving Tasks a small set of statuses that agents and humans move between. A Task is the main unit of work: it holds the goal, the conversation, and the run. The point of statuses is that agents can advance a Task right up to the review gate, and only a person can move it past.

Status Who moves it here Meaning
Backlog Person Scoped, not started. No run happens here.
Ready Person Cleared to run; an Agent can claim it.
In progress Agent A run is executing in an isolated worktree.
Ready for review Agent Run finished, change and evidence attached, waiting for a person.
Changes requested Person Sent back with a comment; a follow-up run continues on the same Task.
Done Person Accepted, merged, released.

The gate is the transition into Ready for review and out to Done. Agents can reach the gate on their own; they cannot pass it. That single rule is what lets you run ten agents safely: they fill the Ready for review column in parallel, and you drain it at whatever pace keeps quality honest. The column length is your bottleneck, made visible.

Here’s the honest decision pair. Gate every change that ships to users or touches money, auth, data migrations, or public APIs. Don’t gate throwaway spikes, generated boilerplate you’ll regenerate anyway, or work behind a flag that no one can reach yet. A gate on everything just moves the bottleneck back onto the reviewer for changes that never needed a human. Approval gates covers where to place the stop for higher-risk work.

Design review so it scales

Reviewing parallel agent output is not the same as reviewing a teammate’s PR. You have more of it, it arrives in bursts, and you often know less about how it was made. Three moves keep the reviewer’s cost per item down.

Attach the evidence to the Task, not to your memory. A run should return the change summary, the typecheck and test output, the build result, and the known limits, all on the record. When you open a Ready for review item, you’re reading a case file, not reconstructing what happened. That’s the difference between reviewing with task-level evidence and re-running everything yourself to regain trust.

Review the decision, not the keystrokes. Machine-checkable properties, formatting, type safety, whether tests pass, should be gated by machines before the Task reaches you. Google’s code review guidance makes the same split: automate what a tool can verify so the human spends attention on design and fit. Your scarce judgment goes to “is this the right change,” not “did it compile.”

Cap work in progress. A hard limit on how many Tasks sit in In progress at once keeps the Ready for review column from becoming a wall you’ll never scale. This is the kanban WIP limit idea applied to agents: starting a sixth run when five are already waiting for you doesn’t add throughput, it hides the queue. It also helps to monitor the running agents so you can tell a stuck run from a slow one before it reaches the gate.

One worked example, from ticket to merged PR

Take a real-looking Task, SH-312: “Rate-limit the public search endpoint.” A person shapes it in the Backlog with acceptance criteria (429 after N requests per minute per key, documented header, a test) and moves it to Ready.

Three Agents are working the Space in parallel. One picks up SH-312, clones the configured repository onto a Computer, and runs in an isolated worktree so its branch never collides with the two other runs in flight. It writes the limiter, adds a test, runs typecheck and the suite, opens a PR whose title carries the Task ID, and moves SH-312 to Ready for review with the change summary, passing test output, and one known limit noted: the counter is in-process, not shared across instances.

You open the Task, not fifteen tabs. The evidence is right there, and the known-limit note is the useful part: the tests pass, but shared state across instances is a design call you have to make. You leave a comment and move it to Changes requested. That comment starts a follow-up run on the same Task; the Agent switches the counter to a shared store and returns updated evidence. This time you accept. Moving SH-312 to Done is your decision, recorded as a status change with your name on it, not a silent side effect of a git merge.

Meanwhile the other two Agents are sitting in Ready for review. Because the queue is a column and not your inbox, you clear them next, in order, without losing the thread.

The line you keep for people

Parallel agents are a throughput story right up until review becomes the constraint, and then they’re a queueing story. The teams that get value from many agents treat review as designed infrastructure: explicit statuses, a gate agents can reach but not cross, evidence on the Task, machine checks before human ones, and a limit on how much can be in flight.

You can build this by hand with worktrees, a branch-naming convention, and discipline. Or you can let Sharkly manage the workflow: assignable Agents, Tasks that carry the run and its evidence, and a board where Ready for review is a column you drain, so agents research, execute, test, and report while people keep the last decision. If you’re already running several agents, that’s one place to manage their tasks and results instead of counting on memory.

FAQ

Does running agents in parallel save time if review is serial? Only up to the point where finished runs arrive faster than you can accept them. Past that, extra agents grow the review queue without shipping anything sooner. The gain comes from making review visible and gating only what needs a human.

Should every agent change get human review? No. Gate changes that ship to users or touch auth, data, money, or public APIs. Let low-risk or regenerable work through with machine checks only. Gating everything moves the bottleneck back onto the reviewer for changes that never needed one.

Isn’t merging the PR the same as accepting the work? No. Merging is a git operation; acceptance is a person deciding the change meets the Task’s criteria. Keeping them separate preserves review quality when volume is high. See human-in-the-loop coding agents for where to keep the human decision.

Explore more

Mixed-Model Crews: Cheap Workers, an Expensive Reviewer, and the Handoff Between Them

Mixed-Model Crews: Cheap Workers, an Expensive Reviewer, and the Handoff Between Them

Cheap GPT-6 Luna workers plus a Claude Opus 5.5 reviewer only works if the handoff carries evidence, not a summary. The packet a worker returns, what the reviewer sends back, and the caching math.

23 September 2026

Ten Agents on GPT-6 Luna or One on Opus 5.5: Which Configuration Actually Ships More

Ten Agents on GPT-6 Luna or One on Opus 5.5: Which Configuration Actually Ships More

GPT-6 Luna is 93% cheaper per task than Claude Opus 5 on DeepSWE 1.1 while scoring 66.6%. Ten cheap agent runs or one expensive one: it depends on whether a test suite or a person picks the winner.

23 September 2026

Coding Agent Permissions: What an Agent Should and Shouldn't Be Allowed to Do

Coding Agent Permissions: What an Agent Should and Shouldn't Be Allowed to Do

Coding agent permissions in three layers: runtime tools, the machine and its credentials, and task scope. The failure mode at each step, plus a checklist for what an agent should and shouldn't do.

20 September 2026