Cheaper Opus 5.5 Does Not Mean More Shipped: Why Model Price Was Never Your Ceiling

Claude Opus 5.5 costs 40% less to run than Opus 5, and GPT-6 Luna less again. Cheaper models raise how fast agent work arrives, not how much your team can review and accept.

Sophia Carter

Sophia Carter

24 September 2026

Cheaper Opus 5.5 Does Not Mean More Shipped: Why Model Price Was Never Your Ceiling

Anthropic says Claude Opus 5.5 costs 40% less to run than Opus 5. Somebody on your team has already done the obvious sum: same budget, roughly 1.7 times as many agent runs, so roughly 1.7 times as much shipped work. That sum is arithmetically fine and operationally wrong, and the week it gets acted on is the week your board fills with finished runs nobody has read.

The constraint on running more agents in parallel was never the model bill. It was one person’s capacity to read a diff, understand what an agent decided, and accept it. A price cut does nothing to that number. It raises how fast work arrives at it.

TL;DR

A review ceiling is the number of completed agent runs a team can actually read, judge, and accept in a day, and it is set by people rather than by tokens. Cheaper models raise the arrival rate at that ceiling and leave the ceiling itself untouched, so the first visible effect of a price cut is a longer queue, not a faster release. Agents research, execute, test, and report. People set direction, grant authority, and accept the result. The useful response is to cap concurrency against reviewers instead of against budget, and to separate agent work that converges into one reviewable result from agent work that fans out into many. Sharkly keeps context, progress, blockers, results, and human review visible on the Task from request to release, which is what makes the queue countable in the first place.

What actually got cheaper

Two labs shipped on 2026-09-22. Here is the part of it that touches your invoice.

Claude Opus 5 Claude Opus 5.5 GPT-6 Sol GPT-6 Luna
API id claude-opus-5 claude-opus-5-5 gpt-6-sol gpt-6-luna
Input per 1M tokens $5 $4 $2 $0.10
Output per 1M tokens $25 $20 $10 $0.50
Cached input n/a $0.20 read, $5 write 90% off reads 90% off reads

Anthropic states Opus 5.5 costs 40% less to run than Opus 5 and produces output about 30% faster. The list rates above are 20% off the Opus 5 token price, which is our own arithmetic from two published numbers and a separate claim from the 40%. OpenAI states that Sol and Luna are each 50% cheaper than GPT-5.6 promotional pricing, and “promotional” is OpenAI’s own word, so that headline is measured against a discounted rate rather than a list one.

Cost per completed task moved further than cost per token. On AutomationBench 1.0.6, OpenAI reports GPT-6 Sol at extra high reasoning scoring 33.2% at $0.27 per task, against Claude Opus 5 at max reasoning scoring 26.9% at 11.1 times Sol’s cost. On DeepSWE 1.1, OpenAI puts GPT-6 Luna at max reasoning at 66.6% and 93% cheaper per task than Opus 5.

So the price of an attempt fell by a lot, across two vendors, in one day. Nothing else in your week changed.

The sum everybody did this week

Hold spend flat and divide. At 40% less to run, the same monthly figure buys about 1.67 times as many Opus 5.5 runs as it bought Opus 5 runs. Route the routine half of the work to Luna instead and the multiplier stops being interesting, because at $0.10 and $0.50 per million tokens the model line is no longer the thing you are budgeting.

Now do the second sum, the one that decides whether any of it ships. Take the number of completed runs one of your engineers can read properly in a day. Not skim, not approve on a green check: read the diff, understand what the agent chose to do, and accept responsibility for it. Whatever that number is, it was the same on 2026-09-21 and it will be the same next quarter. The review-capacity arithmetic works out badly even before a price cut, and a price cut only changes the input side.

A price cut is not a capacity increase. It is a change in how cheap it is to generate something that still requires a person.

The tell: cost was not what stopped you

If model cost were the binding constraint, teams would have been running the cheapest configuration that worked. They were not. OpenAI’s own numbers say the median researcher in its organization runs more than $600 a day of inference at API prices and the 90th percentile runs $7,000 a day. Its published benchmark table shows teams knowingly using an 11.1 times more expensive per-task configuration to get a lower score, because a result you trust is worth more than a result you have to check twice.

That is the shape of the real constraint. People pay a large multiple for output that needs less human scrutiny per unit. The scarce resource being bought is attention, not compute, and it was already being bought at a steep premium.

Which means a price cut lands asymmetrically. It makes the generating side cheaper and leaves the accepting side exactly where it was.

Fan-out that converges, and fan-out that lands on the board

Here is the distinction that makes cheap models genuinely useful instead of quietly expensive, and it is the one most teams skip.

Some parallelism converges. Three agents attempt the same migration in isolated worktrees, a test suite and a leader Agent pick the one that passes, and one result reaches a person. The cost tripled and the review load did not move. This is where a price cut buys you real throughput, and where the new economics are worth changing your setup for.

Some parallelism diverges. Ten agents pick up ten unrelated Tasks and produce ten diffs, ten change summaries, and ten acceptance decisions. The cost is the same ten runs. The review load is ten times one person’s unit of attention.

Converging fan-out Diverging fan-out
Example 3 attempts at one migration, best one wins 10 Tasks, 10 Assignees, 10 diffs
What cheaper models buy More attempts, higher chance one is good More finished work waiting
Review items produced 1 10
Safe to scale on price alone Yes No

Use cheap models to widen converging fan-out as far as your patience for retries goes. Do not use them to widen diverging fan-out past your reviewers, however affordable the runs have become. That is the honest decision pair, and the second half of it is the one that costs money when it is ignored.

What to change instead of the budget

Cap concurrency against reviewers, not against spend. Write the rule down where the Crew can see it: open Tasks awaiting human acceptance should not exceed the number of reviewers on shift multiplied by however many reviews one of them completes in a shift. Your two numbers, not ours. When the cap is hit, agents stop claiming new Tasks even though the budget would allow it. A written routing rule beats a preference here, because a preference quietly loses to a cheaper price.

Make “waiting for a human” a visible state. A run that finished and a run that stopped to ask something look identical from outside, and both are invisible if the only record is a terminal. Execution, blockers, results, and follow-up discussion should return to the Task timeline, so the queue is a thing you can count rather than a feeling you get on Friday.

Price the review, not the tokens. Cost tracking across agents that reports dollars per run answers a question that no longer binds. The number worth putting on a dashboard is human minutes per accepted change. If that went up after you switched to a cheaper model, the cheaper model is more expensive.

Spend part of the saving on evidence. A run that attaches its typecheck, test results, change summary, and known limits costs more tokens and less attention. At $0.20 per million cached input reads on Opus 5.5 and 90% off cached reads on GPT-6, the extra verification pass is close to free and it buys back the resource that was actually scarce.

Where this leaves the week

Two labs cut the price of an attempt on the same day, and one of them benchmarked against a model that was superseded hours later. Neither of them shipped anything that reads a diff for you. Human ownership of the work is the part that did not get 40% cheaper, and it is the part that decides what reaches production.

If you are already running several coding agents, the useful move this week is not to raise the concurrency number. It is to find out what your review ceiling actually is, then set the concurrency to match it. Sharkly gives you one place to see the runs, the blockers, the results, and the queue of work waiting for a person to accept it.

Explore more

Sonnet 5.5 or Opus 5.5: Which Agent Gets the Task?

Sonnet 5.5 or Opus 5.5: Which Agent Gets the Task?

Sonnet 5.5 costs half of Opus 5.5, but effort decides cost per task. A written routing rule for coding agents: which model, what effort, when to escalate.

29 September 2026

Model Handoffs Lose the Reasoning: What Claude Sonnet 5.5 Changes for Multi-Agent Work

Model Handoffs Lose the Reasoning: What Claude Sonnet 5.5 Changes for Multi-Agent Work

Claude Sonnet 5.5 binds thinking blocks to the model, conversation and account, so reasoning never survives an agent handoff. What the Task must carry instead.

29 September 2026

Claude Sonnet 5.5 vs GPT-6 Sol: Two Runtimes at the Same Price

Claude Sonnet 5.5 vs GPT-6 Sol: Two Runtimes at the Same Price

Claude Sonnet 5.5 and GPT-6 Sol both list at $2/$10. What the shared benchmarks show, why cost per task flips with effort, and how to test both on your code.

29 September 2026