One line in one config decides where most of your team’s tokens go this quarter. On 2026-09-22 both candidates for that line changed at once: Anthropic shipped Claude Opus 5.5, OpenAI shipped GPT-6 Sol and GPT-6 Luna. If you are running several coding agents in parallel, the question is not which model is smarter. It is which one your Agents use by default, and what that does to the queue of work waiting for a person to accept it.
There is a catch in the evidence, worth knowing before you read either launch post. OpenAI benchmarked Sol against Claude Opus 5. Not Opus 5.5. Opus 5.5 did not exist when that post was written, and it shipped hours later at $4 per million input tokens against Opus 5’s $5. Every “beats Opus 5 at a fraction of the cost” line in OpenAI’s post is therefore measured against a model that was superseded the same day. That does not make the numbers wrong. It makes them the last published numbers rather than the current ones.
TL;DR
A default Runtime is the model your Agents use when nobody chose otherwise, which in practice is most runs. Pick it on three axes: cost per completed task, how long it holds a task before it needs you, and what it does when it gets stuck. On published figures Sol is cheaper per completed task, though the only published comparison is against Opus 5, and Opus 5.5 is the one that stays on a long task without stopping. Nobody has run the two against each other on a common harness, so treat both launch posts as vendor claims and settle it on your own repository.
What each vendor actually published
| Claude Opus 5.5 | GPT-6 Sol | |
|---|---|---|
| API id | claude-opus-5-5 |
gpt-6-sol |
| Input / output per 1M | $4 / $20 | $2 / $10 |
| Cached input | $0.20 read, $5 write | 90% off reads |
| Context | 1M | 872k |
| Compared in its launch post against | Claude Opus 5 ($5 / $25) | Claude Opus 5, GPT-6 Astra ($10 / $50) |
Anthropic states Opus 5.5 costs 40% less to run than Opus 5, produces output about 30% faster, and stays on task for 18 hours or more. Maximum output is 128k tokens, and a fast mode is priced separately at $8 / $40.
OpenAI states Sol and Luna are each 50% cheaper than GPT-5.6 promotional pricing. The word promotional is OpenAI’s own and it matters: the comparison is against a discounted rate, not against GPT-5.6’s list rate. OpenAI also says GPT-6 Astra “continues to be our best model across the board”, which is a useful reminder that Sol is the value tier of its own family, not the flagship.
Axis 1: cost per completed task, not cost per million tokens
Token price is a proxy for the number you actually pay. A model at half the rate that needs three attempts costs more than one that gets it in a single run. OpenAI publishes cost per task alongside score, which is unusually direct.
| Benchmark, as OpenAI published it | Sol | Compared against |
|---|---|---|
| AutomationBench 1.0.6 | 33.2% at extra-high, $0.27 per task | Claude Opus 5 at max: 26.9%, 11.1x Sol’s cost per task |
| AutomationBench 1.0.6 | same | GPT-6 Astra at low: 30.3%, 3.9x Sol’s cost per task |
| Agents’ Last Exam V1 | 56.4% at max | above Opus 5’s best, at 60% lower cost per task |
| DeepSWE 1.1 | 68.8% at max | within 1.1pp of Fable 5 at extra-high (69.9%), about 80% cheaper per task |
| OSWorld 2.0 offline | 60.5% at extra-high | Opus 5 at medium: 60.3%, about 80% cheaper per task |
The first row is the one people are quoting. Sol at extra-high reasoning scores higher than Claude Opus 5 at max reasoning for about 9% of Opus 5’s cost per task. Multiply $0.27 by 11.1 and Opus 5 at max lands near $3 per task on that harness, which is our arithmetic from OpenAI’s two figures rather than a number OpenAI printed.
Now apply the catch. Anthropic puts Opus 5.5 at 40% less to run than Opus 5, so that $3 figure is stale. It does not simply fall by the same 40% either, because cost per task moves with how many tokens a model burns getting there, not with the list rate alone. The published gap was large, one side of it moved the same day by a stated amount, and nobody has re-run the comparison. If you want that number for your codebase, run the comparison on your own repository rather than waiting for a lab to run it for you.
For scale: OpenAI reports its median researcher spending over $600 a day on coding agents, and the 90th percentile $7,000 a day. Cost per task stops being an abstraction fast.
Axis 2: how long it holds the task
The two vendors answered the same problem in different places.
Anthropic answered it in the model. Opus 5.5 stays on task for 18 hours or more. A run that outlives your working day changes supervision: you stop watching it and start returning to it, which only works if the run reports somewhere that is not a terminal you closed.
OpenAI answered it in the cache. GPT-6 discounts cached input reads by 90%, raises hit rates by default, and, importantly for agent work, no longer breaks the cache when you change reasoning effort or the set of available tools. Explicit breakpoints let you choose where a cached prefix ends. GitHub reports more than 50% fewer prompt tokens needing fresh processing across billions of requests. A long agent run is mostly the same context read many times, so this is a cost-of-duration fix rather than a duration fix.
Use Opus 5.5 as the default when your typical Task is a migration, a wide refactor, or anything where stopping halfway costs more than the tokens. Use Sol as the default when your typical Task is bounded, repeated many times a day, and cheap to retry. Most teams have both kinds, which is why the default matters less than being deliberate about the exceptions.
Axis 3: what it does when it is stuck
This axis has no benchmark and it is the one that generates support pings.
Artificial Analysis, a third party rather than either vendor, measured the “max” reasoning variants at 115.2 output tokens per second for Sol with 102.15 seconds to first token, and 153.9 tokens per second for Luna with 124.23 seconds. Those are max-effort configurations and you should confirm them against the current listing before betting on them. The shape is what matters: at high reasoning effort, the cheap models are not the fast-to-start models.
An agent that is thinking for 102 seconds and an agent that has hung look identical in a terminal. So does one that finished 40 minutes ago and one that stopped to ask a question nobody saw. No model release fixes that. A status that returns to the Task does: In Progress, Waiting for human, Failed, Ready for release, with the log and the question attached. Agents research, execute, test, and report. People set direction, grant authority, and accept the result.
Give how each Runtime signals progress and failure the same weight as the benchmark table. You will read the status a hundred times a week and the benchmark once.
Do not put the two benchmark tables side by side
Both vendors published an AutomationBench number this week. OpenAI reported Sol at extra-high reasoning scoring 33.2% on AutomationBench 1.0.6. Anthropic reported Opus 5.5 at 40.0% on AutomationBench. Both reported OSWorld 2.0, at 60.5% and 81.8% respectively, and OpenAI’s figure is on the offline set.
Those are not a head-to-head. Different reasoning settings, at least one different benchmark variant, different harnesses, no shared cost-per-task figure, and each run by the party that benefits. The posts comparing Opus 5.5 to GPT-6 Astra on specific numbers are tweets, not vendor data. Artificial Analysis does publish one common index, putting Opus 5.5 at 58 and first of 212 models, Sol at 48, Luna at 37, and that is a third-party aggregate rather than a measurement of your codebase.
The decision that survives the next launch
Write the default down. Not as a preference someone holds, but as a rule attached to the Agent: this Agent runs this Runtime, these Tasks go to it, and here is the cheaper Agent that takes the routine ones. Something like this, kept where the team reads it:
Default Runtime: Claude Code on claude-opus-5-5
Tasks: migrations, wide refactors, anything that fails badly halfway
Second Runtime: Codex on gpt-6-sol
Tasks: bounded, repeated many times a day, cheap to retry
Escalation: a Sol Task that fails twice returns to the Task for a person,
it does not silently re-run on the expensive Runtime
That is what turns which model gets which task into throughput, and it survives the next launch, because when the answer changes you edit one rule instead of asking six people what they have been using. Two labs shipped their strongest value tier on one day and one benchmarked against a model that aged out hours later. A default Runtime you can change in one place is worth more than picking correctly this week.
The thing neither launch changed is the queue. Cost per completed task fell hard on both sides, so a fixed budget now buys several times more finished agent work than it did last week. It buys zero additional reviewer hours. Review capacity is the ceiling, and a price cut pushes volume against it rather than relieving it. Keep the goal, the run, the blockers, the diff, and the human review visible in one place from request to release, and the default Runtime becomes a setting rather than a bet.
Sharkly does not run the model and does not choose it for you. It is the shared Task, Computer, context and review layer around the tools that do. If you want to keep using multiple coding agents rather than committing to one ecosystem, that is the problem it is designed around. Connect one Computer, create one Agent on each Runtime, and give them the same boring Task.



