The invoice is the easiest number to move and the least interesting one to read. Route your Crew’s routine work to GPT-6 Luna at $0.10 and $0.50 per million tokens and the model line collapses inside a day. Two weeks later the bill is still down, the board carries more finished runs than ever, and last Tuesday’s four changes are still waiting for somebody to accept them.
Nothing went wrong. You moved cost out of a line you were watching into a line nobody prices.
Cost per accepted change is the total money and human time spent on every attempt that led to one change a person accepted. It counts the runs that failed, the output nobody merged, and the minutes somebody spent deciding about both. No vendor publishes it, no benchmark measures it, and it is the only number that moves with what your team ships.
TL;DR
A cheaper model lowers the price of an attempt and raises three things that are paid in reviewer attention: attempts per shipped change, output per attempt, and decisions per shipped change. Attention did not get cheaper on 2026-09-22. The cheap tier is genuinely cheaper wherever a machine decides whether an attempt passed, and quietly expensive wherever a person has to. Agents research, execute, test, and report. People set direction, grant authority, and accept the result. The fix is an accounting change before it is a routing change: measure attempts and human minutes per accepted change, per Runtime, and make work waiting for a person a state you can count rather than a feeling you get on Friday.
Three units, and only one of them is yours
| Unit | Who uses it | What it tells you | What it leaves out |
|---|---|---|---|
| Per million tokens | Vendor pricing pages | What an input and an output cost | How many tokens your job takes |
| Per completed task | Benchmark tables | The cost of one scored attempt | Failed attempts, your codebase, the review |
| Per accepted change | Your team | What one shipped change cost end to end | Nothing, which is why nobody publishes it |
Both labs shipped on 2026-09-22, and both published in the middle column. Do not read those tables as if they were the right column.
What the cheap tier costs and what it scores
| GPT-6 Luna | GPT-6 Sol | Claude Opus 5.5 | |
|---|---|---|---|
| API id | gpt-6-luna |
gpt-6-sol |
claude-opus-5-5 |
| Input / output per 1M | $0.10 / $0.50 | $2 / $10 | $4 / $20 |
| Cached input | 90% off reads | 90% off reads | $0.20 read, $5 write |
| Context | 1M | 872k | 1M |
| DeepSWE 1.1 at max reasoning | 66.6% | 68.8% | Not in the launch material |
| Artificial Analysis index | 37 | 48 | 58, first of 212 |
Anthropic reports 40.0% on AutomationBench for Opus 5.5. OpenAI reports 33.2% for Sol at extra high reasoning on AutomationBench 1.0.6, at $0.27 per task. Two labs, two harnesses, one published cost per task. That pair is not a head to head, and being unable to make it into one is the point: no vendor table gets you to cost per accepted change.

Hidden term one: attempts per accepted change
OpenAI puts Luna at max reasoning at 66.6% on DeepSWE 1.1 and Sol at max at 68.8%. By our own arithmetic, that leaves roughly one Luna attempt in three without a passing result. Your repository is not DeepSWE and your rate will differ, but the shape holds: a cheap tier scoring a few points lower is not a few points worse per run, it is a few points more likely to hand you something that needs another lap.
The money side of that lap stays trivial. OpenAI reports Luna at max as 93% cheaper per task than Claude Opus 5 on the same benchmark. Double the attempts and, again by our arithmetic, you are still around 86% cheaper. The model line never becomes the problem.
The failed attempt is not free anywhere else. It produced a branch, a change summary, and a decision that belongs to a person. A failed run is not a zero. It is a review item with no change at the end of it.
Hidden term two: the cheap tier gets good by thinking longer
Read where Luna’s good numbers come from. The 66.6% is at max reasoning, and OpenAI’s factuality note is that Luna at higher effort matches GPT-5.6 Sol at about a hundredth of the cost. The cheap tier reaches acceptable quality by spending more effort, and effort arrives as output tokens and wall clock.
Artificial Analysis, a third party rather than either vendor, measures the max reasoning variants at 153.9 output tokens per second for Luna with 124.23 seconds to first token, and 115.2 tokens per second for Sol with 102.15 seconds. Treat those as indicative: they are third party, they cover the max variants specifically, and they move as the lab re-measures.
Two consequences land on your team, not on your invoice. A run that is silent for two minutes looks exactly like a run that has hung, so somebody goes and checks. And a longer trace with a longer diff is more to read, per attempt, before anyone can accept it. Reading is the scarce resource, and you just bought more of the thing that consumes it.
Hidden term three: decisions, not diffs
Count the human decisions between a Task being assigned and a change being accepted. On a strong Runtime the usual path is one: read it, accept it. On a cheap Runtime doing unverifiable work it branches, and every branch is a person:
- The run failed. Retry on the same Runtime, or escalate to a stronger one?
- The run passed its tests but took an approach you would not have chosen. Accept, or send it back?
- The run produced three partial results. Which one is the change?
- The run stopped to ask something. Who answers, and where does the answer live?
None of that appears in a token bill. All of it appears on the calendar of whoever owns human review.
The sum, with your numbers in it
cost per accepted change
= attempts per accepted change
x (model cost per attempt + reviewer minutes per attempt x loaded cost per minute)
+ escalation cost when the cheap tier hands the job back
Fill it in with your numbers, not ours. At $0.10 and $0.50 per million tokens the model cost inside the bracket rounds to nothing, so the expression is governed by reviewer minutes multiplied by attempts. Switching to a cheaper Runtime shrinks a line item that was already small and multiplies the one that was not. Sometimes that is still a good trade. It is never one you can evaluate from a pricing page.
When the cheap tier really is cheaper
Use a cheap Runtime when a machine decides whether the attempt passed. Typecheck, unit tests, a build, a lint rule, a schema check: the verdict arrives before a person spends attention, and the extra attempts are absorbed by the harness rather than by your Friday. That is most of the volume in a working Crew, and the honest case for routing the boring work to Luna.
Do not use a cheap Runtime when acceptance is a judgement. Architecture decisions, data migrations, anything security relevant, anything where “is this right?” cannot be answered by a command. There the score gap converts straight into reviewer minutes, and the cheapest thing you can buy is an attempt you only have to read once.
Write the escalation rule down instead of leaving it to preference, because a preference loses to a cheaper price every time. Model routing is a rule about verifiability first and cost second.
Why no vendor will tell you this
Both labs benchmarked against each other on the same day, and OpenAI’s comparison baseline was Claude Opus 5, which Anthropic replaced with Opus 5.5 hours later. Every vendor’s cheap tier is the top of its own funnel. None will publish the sentence “our budget model will cost your reviewers an extra pass”, and none is wrong to leave it out: the number depends on your codebase and your people.
Sharkly does not sell model tokens. Model usage continues through the subscriptions or API keys configured in the tools your team already uses, so nothing here rides on which Runtime you pick. That is the only position this argument can be made from honestly, and it is the same reason your workflow should survive the next cheap model without being rebuilt.
What to put on the dashboard instead
Human minutes per accepted change, per Runtime. If that went up after the switch, the cheaper model is more expensive and you now have the receipt. Dollars per run answers a question that stopped binding on 2026-09-22.
Attempts per accepted change, per Runtime. This term decides everything else and no vendor can give it to you. You get it by recording which Runtime produced each attempt, on the Task it belongs to.
Waiting for a person as a real state. Execution, blockers, results, and follow-up discussion return to the Task timeline, so the queue needing a human is countable instead of anecdotal. Require each run to arrive with its evidence:
Before a run reaches a person, the Task carries:
- the change summary
- typecheck, test, and build results
- known limits and what was not attempted
- the attempt number and which Runtime produced it
You do not need any of this at small scale. If one person runs two Agents and reads everything the same afternoon, keep doing that: states and queues are overhead on a problem you do not have yet. The accounting starts to matter when nobody can say from memory how many finished runs are waiting.
The part that did not get cheaper
Two labs cut the price of an attempt on one day. Neither shipped anything that reads a diff for you, and ownership of the change still ends with a person. A cheaper model is not a cheaper workflow. It is a cheaper attempt, and attempts feed a process whose expensive step is human.
If you are already running several coding agents, Sharkly gives you one place to see which Runtime produced which attempt, what it attached as evidence, and how much work is sitting in front of a person right now.



