A pull request lands in your queue with 340 changed lines across nine files. A coding agent wrote it. The diff does not tell you what the agent was asked to do, which approaches it tried and dropped, what it ran to convince itself the change works, or where it knows it cut a corner. So you do one of two things: skim and approve, or re-derive the whole change from scratch. The first is how bad code ships. The second means the agent saved you nothing.
Reviewing AI-generated code well is a different job from reviewing a colleague’s PR, because the author cannot sit next to you and explain their reasoning. The fix is to review the diff plus its evidence: the goal, the checks that ran, and the limits the agent already flagged. That is what Sharkly is built around. Sharkly sits above your coding agents (Claude Code, Codex, Gemini, and the rest) and keeps the goal, the change, the verification results, and the human review of every run in one Task instead of scattered across terminals and PR descriptions.

This guide walks through the operating model behind that: who assigns work, who runs it, who reviews it, who accepts it, the statuses and gates that separate an agent’s job from a human’s, and one worked example from ticket to merged PR.
Quick answer
Review AI-generated code at the Task level, not the diff level. Have the agent attach a change summary, the typecheck/test/build results it ran, and any known limits to a shared Task, then move that Task to a review status. A human reads the evidence alongside the diff, requests changes or accepts, and makes the merge decision. Agents produce and verify; people accept.
Why reviewing the diff alone fails
Reading code is harder than writing it. A well-documented reality of code review is that the reviewer holds less context than the author, and with an agent author that gap is total. A GitHub diff shows you the destination, never the route. You cannot see that the agent tried a caching layer first, hit a race condition, and backed it out. You cannot see that it ran the test suite and one flaky test failed unrelated to the change.
Without that context, review collapses into two failure modes. Rubber-stamping approves plausible-looking code that quietly breaks an edge case the agent never considered. Full re-derivation makes you reconstruct the reasoning yourself, at which point the agent’s speed advantage is gone. Neither scales when three agents are opening PRs at once. The way out is to make the agent’s work reviewable, not just its output.
What task-level evidence actually means
Task-level evidence is the record an Agent attaches to its Task: the goal it worked from, the change it produced, the checks it ran, and the limits it knows about. Instead of a bare diff, you get a diff with a provenance trail.
In practice the evidence bundle is four things:
- The goal and acceptance criteria, in the Task description, so review is measured against what was asked rather than what looks reasonable.
- The change summary, in the agent’s own words: what it touched and why.
- Verification results: the output of the typecheck, tests, and build the agent ran before handing off.
- Known limits: the edge cases, assumptions, and follow-ups the agent flagged instead of hiding.
All of it returns to the Task. Execution, blockers, results, and follow-up discussion live on one timeline, so a reviewer opens a single record instead of stitching together a PR, a Slack thread, and a half-remembered prompt. This is the discipline behind a traceable requirements-to-code workflow: the evidence is part of the deliverable.
The operating model: who assigns, who runs, who reviews, who accepts
Good review depends on clear roles. The contract is short: agents research, execute, test, and report; people set direction, grant authority, and accept the result. Automation stops where team judgment is required.
| Role | Who | Job |
|---|---|---|
| Assigner | A person (usually the Task Owner) | Describes the need, sets acceptance criteria, assigns the Task |
| Runner | An Agent, on a connected Computer, via a Runtime | Claims the Task, executes in an isolated worktree, attaches evidence |
| Reviewer | A person | Reads the diff against the evidence, requests changes or approves |
| Accepter | A person (the Owner or a designated approver) | Makes the final acceptance and merge/release decision |
Sharkly is not a replacement for Claude Code, Codex, or other execution tools. It adds the shared Task, Computer, context, and review layer around the tools your team already uses. The runtime writes the code; Sharkly decides whose queue the review lands in and what evidence has to be present before a human looks. If you want the deeper version of the assignment side, see who assigns and who accepts agent work.
Statuses and review gates: agents move to Ready for review, humans move to Done
The review gate is a status boundary. In Sharkly, every task status belongs to a fixed category that carries system meaning: Backlog (parked), Unstarted (ready, not started), Started (in progress), and Completed (terminal). Several visible statuses can share a category, so a task type’s workflow can run Development and Ready for review as two distinct Started statuses before anything reaches Completed.
That structure is what makes the gate real. An Agent can carry a Task from Development to Ready for review, and it can leave the Task waiting for human review. It does not get to move the Task into a Completed status on its own. Acceptance is a human transition, by design.
| Status | Category | Who moves it here |
|---|---|---|
| Backlog | Backlog | Person, when parking work |
| Todo | Unstarted | Person, when the Task is ready |
| Development | Started | Agent, when it claims and works the Task |
| Ready for review | Started | Agent, when evidence is attached |
| Done | Completed | Person, after accepting |
Configure the workflow so the transition into your Completed status requires a human, and you have a review gate that holds even when agents are producing faster than you can watch. This is the mechanical version of keeping a human in the loop for coding agents, and it is where you decide where to stop agents before they merge.
A worked example: from ticket to merged PR
Take SH-312: “Add rate limiting to the public search endpoint, 60 requests per minute per API key.”
- A person files SH-312 with that goal and two acceptance criteria: requests over the limit return HTTP 429, and the limit is per key, not per IP. The Task starts in Todo.
- An Agent claims it, moves it to Development, and runs in an isolated worktree so its checkout never collides with other agents working the same repo. (Worktrees are why parallel runs stay clean; see one checkout per task.)
- The Agent implements a token-bucket limiter, opens a PR, and attaches its evidence to SH-312: a change summary, the passing typecheck and unit tests, the build result, and one known limit: “The bucket lives in process memory, so limits reset on deploy and are not shared across instances.”
- The Agent moves SH-312 to Ready for review. It lands in a human’s queue with the diff and the evidence together.
- The reviewer reads the known-limit note first, decides the in-memory bucket is fine for now, and requests one change: return the
Retry-Afterheader on a 429. The Agent makes the change and re-attaches updated test output. - The reviewer accepts, merges the PR, and moves SH-312 to Done, with a follow-up ticket for a shared limiter already linked.
The review took minutes because the reviewer never reconstructed the change; the evidence did that work. It is the same loop as Claude Code project management from issue to review, with the evidence bundle making review fast instead of a bottleneck.
How to review the evidence fast, and when to go light
Not every change needs the full treatment. Use a decision rule.
Go heavy when the change touches auth, billing, data migrations, public APIs, or anything hard to roll back: read the diff line by line, re-run the tests yourself, and check the known-limits note against your own edge cases. Go light when the change is a copy tweak, a well-tested internal refactor, or a dependency bump with green checks: read the summary, scan the diff, confirm verification passed, and accept.
A fast heavy-review checklist:
- Does the diff satisfy every acceptance criterion, not just the happy path?
- Did the agent actually run the checks, or claim it did? Verify the attached output.
- Read the known-limits note. Is any limit a blocker rather than a follow-up?
- Is the change scoped to the request, or did the agent touch unrelated files?
If you are already running several coding agents, this is the workflow Sharkly gives you one place to manage: assign the Task, let the agent attach evidence, review at the gate. You can build the pieces by hand with worktrees, PR templates, and a status board, or Download Sharkly and get the Task, evidence, and gate as one system.
FAQ
What is task-level evidence for AI-generated code? It is the record an agent attaches to its Task alongside the diff: the goal, a change summary, the typecheck/test/build results it ran, and any known limits. Reviewing that bundle is faster and safer than reviewing a diff alone.
Can an AI agent approve its own code? No, and that is the point of a review gate. An agent can move a Task to Ready for review, but the transition into a Completed status is a human decision. Agents produce and verify; people accept.
How do I review AI code faster without rubber-stamping? Match effort to risk. Read evidence for every change, but reserve line-by-line review and re-running tests for high-risk areas like auth and migrations. See how review capacity scales when agents outpace humans.
Do I still use pull requests? Yes. Agents open real pull requests on GitHub; the Task layer adds the evidence and the gate around them so the PR is not the only place context lives.



