You have a fixed number of reviewers and, since 2026-09-22, a model bill that no longer decides anything for you. GPT-6 Luna lists at $0.10 per million input tokens and $0.50 per million output, API id gpt-6-luna. Claude Opus 5.5 lists at $4 and $20, API id claude-opus-5-5. That is a 40x spread on both sides of the meter, our arithmetic from the two vendors’ own list prices, and it turns a question that used to be theoretical into a weekly scheduling decision. For the same money, do you run ten cheap Agents or one expensive one?
The two configurations are not substitutes. What separates them is not model quality, and it shows up at the end of the run rather than the start: each one puts a different amount of work in front of a person.
The arithmetic that makes the question live
OpenAI publishes cost per task next to score, which is more useful than a price list. On DeepSWE 1.1, Luna at max reasoning effort scores 66.6%, which OpenAI describes as comparable to Claude Opus 5 and Fable 5 at medium effort, at 93% lower cost per task than Opus 5 and 96% lower than Fable 5.
Work the first figure backwards and one Opus 5 task is about fourteen Luna tasks. Ten Luna runs therefore cost roughly seven tenths of a single Opus 5 run on that harness. That multiplication is ours, from OpenAI’s stated 93%.
Anthropic says Opus 5.5 costs 40% less to run than Opus 5. Apply that and the ratio falls to somewhere near nine to one. That figure chains two vendors’ measurements taken on different harnesses, so treat it as a direction rather than a number: ten cheap attempts and one expensive attempt now cost about the same, whichever way you compute it.
| GPT-6 Luna | Claude Opus 5.5 | |
|---|---|---|
| API id | gpt-6-luna |
claude-opus-5-5 |
| Input / output per 1M | $0.10 / $0.50 | $4 / $20 |
| Cached input | 90% off reads | $0.20 read, $5 write |
| Context window | 1M | 1M |
| Vendor’s long-horizon claim | none stated | stays on task 18+ hours |
| Artificial Analysis Index | 37 | 58, ranked #1 of 212 |
Two cautions belong with that table. OpenAI’s headline is that Sol and Luna are each 50% cheaper than GPT-5.6 promotional pricing, and “promotional” is OpenAI’s own word: the comparison runs against a discounted rate, not a list rate. And the two launch posts are not a head-to-head. Both vendors report an AutomationBench score, at different harness versions and different effort levels, on the same day. Putting those two percentages in one row would produce a comparison neither lab ran. The cost-per-task ratios above survive because they come from inside a single vendor’s own measurement.
Three configurations that look like one trade
“Ten agents or one” is really three arrangements, and they behave differently at the review step.
| Configuration | What arrives at review | Who picks the winner | Works when |
|---|---|---|---|
| Ten cheap Agents, ten Tasks | Ten diffs, ten decisions | A person, ten times | The backlog is wide and each item is independently acceptable |
| Ten cheap Agents, one Task | One diff, one decision | Your test suite | A machine can rank attempts without a human reading them |
| One expensive Agent, one Task | One diff, one decision | A person, once | Long horizon, weak automatic signal, or a change that needs judgment throughout |
The middle row is the one the price cut actually unlocks, and it is the one most teams skip. Running the same Task ten times and keeping the attempt that passes typecheck, tests and the build is a strategy that only became affordable when an attempt stopped costing real money. At a benchmark pass rate near two thirds, a third of single attempts fail. If attempts failed independently, ten tries would almost never all fail. They do not fail independently: a task with an ambiguous requirement or a missing fixture defeats every attempt in the same way. Fan-out buys you the variance, not a guarantee.
The first row is the one most teams actually reach for, and it is the expensive one. Not in dollars. Ten Agents on ten Tasks produce ten diffs, and every one of them needs a person to read it, understand what the Agent chose, and take responsibility for merging it. That cost did not drop on 2026-09-22 and no model release touches it. We have written about why human review is the bottleneck in parallel agent work and about how many agent workdays one human workday can absorb; the arithmetic there is unchanged by cheaper tokens.
Ten Agents on Luna is not ten times the output. It is ten times the arriving work. A cheaper model buys attempts. It does not buy review capacity, and the two are easy to confuse when the invoice is the only number being watched.
Cost per accepted change is the figure that transfers
Cost per accepted change is what your team spends on inference for every diff a person actually reads and accepts. It is the only agent cost figure that connects to shipped work, because an attempt nobody reviews has not shipped anything.
That definition sorts the three configurations immediately. Fan-out on one Task divides ten attempts across one review decision, so cheaper attempts lower cost per accepted change in a straight line. Width across ten Tasks divides one attempt across one review decision ten times over, so cheaper attempts lower the inference line while the review line stays flat and becomes the constraint. Depth on one expensive Agent buys a higher chance that the single attempt in front of your reviewer is the one worth merging.
There is a second effect that only shows up in practice. Ten cheap attempts and one expensive attempt do not feel the same to work with. Artificial Analysis measures the max reasoning variants of these models at 124.23 seconds to first token for Luna and 102.15 seconds for Sol, third-party figures for the max variants specifically and worth confirming before you plan around them. Ten Agents that each spend two minutes silent before producing a character are not a fast pipeline. They are ten runs that all look hung at once, which is a status problem rather than a model problem.
The decision pair
Use ten cheap Agents when a machine decides. That means the Task has a real acceptance signal that runs without a person: a test that fails today and must pass, a typecheck, a build, a lint rule, a reproduction script. Triage, labelling, test scaffolding and mechanical refactors qualify. So does any change where “did it work” is a command rather than an opinion.
Use one expensive Agent when a person decides. If accepting the change requires reading the code and forming a view, extra attempts do not help, they multiply the reading. The same holds for long-horizon work: Anthropic states Opus 5.5 stays on task for 18 or more hours, and a run that holds a goal across an afternoon is worth more than ten runs that each restart from the same partial understanding.
Do not use fan-out to avoid writing the acceptance signal. Ten plausible diffs with no automatic judge is worse than one diff, because choosing between ten plausible diffs is a harder job than reviewing one, and it is a job that lands on the person you were trying to protect.
What the wide configuration needs from the layer above
Ten Agents running at once is a coordination problem before it is a cost problem. The pieces are known: each run in its own worktree so two Agents editing the same file is not an incident, one Assignee per Task so ownership is never ambiguous, and a status that distinguishes running from waiting from failed.
The Agent defines how the work should be handled. The Computer supplies the host and local resources. The Runtime performs the session. The Task stays the shared record. When ten runs finish within an hour of each other, that last part is what keeps the hour from becoming a scavenger hunt across ten terminals.
A finished Luna run should return something like this to its Task, not to a private session log:
SH-312 Add retry with backoff to the webhook dispatcher
Runtime gpt-6-luna (fan-out, 8 attempts, 3 passed the suite)
Selected attempt 6, smallest diff of the passing set
Evidence typecheck ok, 214 tests pass, build ok, coverage +0.4%
Known limit no test for the 429 path, dispatcher is not idempotent yet
Status Ready for review, waiting on a human
The selection line is the part that matters. It says what decided, so the reviewer knows whether they are confirming a machine’s choice or making one. Ten attempts with no selection record is ten times the work pretending to be one Task.
The answer
Ten cheap Agents ship more when a suite decides and the backlog is wide enough to keep them busy. One expensive Agent ships more when the acceptance decision is human, the horizon is long, or the automatic signal is weak. Most crews want both, routed by Task type rather than by preference, which is the argument in matching models to agent roles.
Neither configuration changes the last step. Agents research, execute, test, and report. People set direction, grant authority, and accept the result. The model prices moved 40x in a day, and the number of changes your team can responsibly accept moved not at all.
If you are already running several agents, Sharkly gives you one place to keep context, progress, blockers, results, and human review visible from request to release, whichever runtime produced the diff.



