Mixed-Model Crews: Cheap Workers, an Expensive Reviewer, and the Handoff Between Them

Cheap GPT-6 Luna workers plus a Claude Opus 5.5 reviewer only works if the handoff carries evidence, not a summary. The packet a worker returns, what the reviewer sends back, and the caching math.

Ryan Mitchell

Ryan Mitchell

23 September 2026

Mixed-Model Crews: Cheap Workers, an Expensive Reviewer, and the Handoff Between Them

The price table from 2026-09-22 suggests an obvious shape. Put GPT-6 Luna at $0.10 per million input tokens on the work nobody enjoys, put Claude Opus 5.5 at $4 on the Agent that checks it, and you keep most of the throughput for a fraction of the bill. Plenty of teams are wiring that up this week, and most will build something that looks like review and is not.

The failure is not in either model. It is in what crosses the gap between them. A reviewer reading a worker’s summary of its own work is reading a claim, and the only thing it can return is a judgement of a claim. The diff passed, says the worker. Looks good, says the reviewer. A person approves it because two models agreed, which is the least informative kind of agreement there is.

So the interesting question in a mixed-model Crew is not which model goes where. It is what the handoff has to carry.

What a mixed-model Crew is

A mixed-model Crew is a group of Agents where the Runtime is chosen per role rather than per team, and at least one role exists only to check the output of the others.

A reviewer Agent is not an approval. It is a second opinion with a cost, produced before a person looks, so the person reads a shorter list. Agents research, execute, test, and report. People set direction, grant authority, and accept the result. Adding an expensive model to the middle of that sentence does not move the last clause.

The three tiers, as the vendors published them

Claude Opus 5.5 GPT-6 Sol GPT-6 Luna
API id claude-opus-5-5 gpt-6-sol gpt-6-luna
Input / output per 1M $4 / $20 $2 / $10 $0.10 / $0.50
Cached input $0.20 read, $5 write 90% off reads 90% off reads
Context 1M 872k 1M

OpenAI describes Sol and Luna as 50% cheaper than GPT-5.6 promotional pricing, and “promotional” is its own word: the baseline is a discounted rate. OpenAI also says GPT-6 Astra “continues to be our best model across the board”.

One caution before anyone reads a ranking into that split. OpenAI’s launch post benchmarks Sol and Luna against Claude Opus 5, because Opus 5.5 did not exist when the post was written and shipped hours later at 40% less to run than Opus 5 by Anthropic’s own account. Nobody has re-run that comparison, and the two vendors do not share a harness. Pick roles on price and on observed behavior in your own repository.

What each vendor does say about its own model is enough to assign roles. For Luna, OpenAI reports 66.6% on DeepSWE 1.1 at maximum reasoning effort, 93% cheaper per task than Claude Opus 5 and 96% cheaper than Fable 5, and at higher effort a factuality result matching GPT-5.6 Sol at about a hundredth the cost. That is a model you can point at volume. For Opus 5.5, Anthropic reports 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode v1.1, 81.8% on OSWorld 2.0, 128k maximum output, and staying on task 18 hours or more. That is a model you can point at judgement.

Reading is the cheap half of an expensive model

The reason this shape survives contact with an invoice is caching, and the asymmetry decides the design.

Opus 5.5 cached input reads at $0.20 per million. That is a twentieth of its own uncached input rate and twice Luna’s uncached rate. Its output is $20 per million, forty times Luna’s $0.50. So a reviewer that reads an enormous amount and writes twenty lines is close to free, and one that narrates at length is not.

That gives you a rule with a receipt behind it: the reviewer reads everything and says almost nothing. Give it the repository prefix, the conventions, the diff and the packet below, then cap what it returns.

The worker side caches too. GPT-6 raises hit rates by default with a 90% discount on cached input reads, adds explicit breakpoints so you choose where a cached prefix ends, and no longer invalidates the cache when you change reasoning effort or tool availability. GitHub reports more than 50% fewer prompt tokens needing fresh processing across billions of requests. For a fleet of Luna workers re-reading the same repository all day, that is most of the bill.

What the handoff has to carry

Here is the artifact, a handoff template built for a model reader. A worker returns it to the Task when it stops, and the reviewer reads only this plus the diff.

SH-312  Add retry with backoff to the webhook sender
Worker:   Agent "backend-worker" - Runtime Codex on gpt-6-luna
Attempts: 2 (attempt 1 failed on lint)
Diff:     4 files, +118 / -12   (worktree sh-312, commit a91f4c2)

Commands run, verbatim:
  $ pnpm lint            exit 0
  $ pnpm typecheck       exit 0
  $ pnpm test -- sender  exit 0   14 passed, 0 failed, 3 skipped
  $ pnpm build           exit 0

Not done:
  - The 3 skipped tests need a live webhook endpoint. Not run.
  - Did not touch the queue consumer, which has the same bug.

Assumptions made:
  - Max 5 retries, ceiling 30s. No spec said. Copied the values
    already used in payments/retry.ts.

Stopped because: acceptance criteria met.

Four properties make that packet worth its tokens.

It is artifacts, not narrative. Commands and exit codes can be re-run. “I verified the behavior” cannot be checked by anyone.

It states what was not done. The skipped tests and the untouched consumer are the two things a summary would have dropped, and they are the two things a reviewer needs most. Ask for absences explicitly, or a tidy summary will not contain them.

It records assumptions. Every bounded Task contains a decision the spec did not make. Written down, it is a review item. Left out, it is a surprise in production a month later.

It names a stop reason. Finished, blocked, out of budget, and failed twice look identical in a repository and mean different things on a board.

What the reviewer returns

REVIEWER  Agent "review" - Runtime Claude Code on claude-opus-5-5
Reads:  repository prefix (cached), the packet, the full diff
Writes: at most 20 lines

VERDICT      changes requested
CHECKED      4 commands re-run, all exit 0. Backoff ceiling honoured.
UNVERIFIED   3 skipped tests. No evidence the retry path ever runs.
BLOCKING     Retry wraps the 4xx branch. A 422 will be retried 5 times.
FOR A PERSON Queue consumer has the same bug and is out of scope.
             Split SH-312, or widen it?

The UNVERIFIED line is what earns the reviewer its price. A cheap worker will not tell you which of its claims rest on nothing, partly because it does not know. An expensive model reading a structured packet can separate what it confirmed from what it took on trust, and that separation is the product of the review step. The FOR A PERSON line is the other half: the reviewer does not decide scope, it returns the question to the Task.

The boundary that keeps this honest: a reviewer Agent can re-run commands, read the diff, and flag the difference between evidence and assertion. It cannot tell you whether this was the right feature. What a review Agent can and cannot verify is a fixed list, and it gets shorter, not longer, when the reviewer is expensive.

Two failure modes to watch

Latency is the first, and it is counter-intuitive. Artificial Analysis, a third party rather than either vendor, measured the “max” reasoning variants at 124.23 seconds to first token for Luna and 102.15 seconds for Sol. [VERIFY] Confirm those against the current listing before relying on them. The shape matters more than the digits: at high reasoning effort the cheap models are not the fast-to-start ones, so a mixed-model Crew has workers that look hung and a reviewer waiting on them. That is a status problem, not a model problem.

Review volume is the second. Ten Luna workers produce ten packets. If the reviewer approves nine and escalates one, you have converted ten agent runs into one human decision, which is the point. If it escalates six, you have paid for the review and still owe six decisions, and the review queue was already the bottleneck. Track the escalation rate per worker Agent for a Sprint. It tells you whether the packet is too thin or the Tasks too big, long before the invoice does.

When not to do this

Use one model for everything when one person runs one or two Agents on one repository. The packet is overhead, the reviewer is a second bill, and you are the reviewer anyway.

Use a mixed-model Crew when several Agents run in parallel, when more than one person can start a run, or when finished work has outgrown the time anyone has to read it.

And do not put the expensive model on review when the work itself is the hard part. A migration that must hold context for hours belongs on Opus 5.5 directly, with a cheap Agent writing the packet after it. Choosing per role means choosing per Task, not once per team, which is why matching models to agent roles is a rule you write down rather than a habit six people hold privately.

The number that did not move

Cost per completed task fell hard on both sides of this launch. OpenAI reports its median researcher spending over $600 a day on coding agents, with the 90th percentile above $7,000. Teams below that line can now afford several times more agent work than last week.

Reviewer hours are exactly where they were. A mixed-model Crew earns its keep by turning more cheap output into fewer human decisions, and that only happens if the handoff carries evidence instead of a summary. Keep the goal, the diff, the commands, the blockers, the unverified claims, and the human review returning to the Task from request to release, and the model split becomes a setting you can change next quarter without changing how anyone works.

Sharkly does not run the model and does not choose it for you. It is the shared Task, Computer, context and review layer around the tools that do, which is where a handoff has to live if a reviewer and a person are both going to read it. If you are already running several coding agents: connect one Computer, create one worker Agent and one reviewer Agent, and make the worker return that packet on something boring.

Explore more

Ten Agents on GPT-6 Luna or One on Opus 5.5: Which Configuration Actually Ships More

Ten Agents on GPT-6 Luna or One on Opus 5.5: Which Configuration Actually Ships More

GPT-6 Luna is 93% cheaper per task than Claude Opus 5 on DeepSWE 1.1 while scoring 66.6%. Ten cheap agent runs or one expensive one: it depends on whether a test suite or a person picks the winner.

23 September 2026

Coding Agent Permissions: What an Agent Should and Shouldn't Be Allowed to Do

Coding Agent Permissions: What an Agent Should and Shouldn't Be Allowed to Do

Coding agent permissions in three layers: runtime tools, the machine and its credentials, and task scope. The failure mode at each step, plus a checklist for what an agent should and shouldn't do.

20 September 2026

Human Review Is the Bottleneck: Designing Review for Parallel Agents

Human Review Is the Bottleneck: Designing Review for Parallel Agents

Run agents in parallel and review becomes the constraint. An operating model, statuses, and review gates: agents move to Ready for review, humans move to Done, with a worked example from ticket to merged PR.

17 September 2026