Your team already routes work between models. It just does it in people’s heads. One engineer runs everything on the expensive tier because it fails less often. Another switched to a cheap one last month and told nobody. A third is still on whatever the tool defaulted to at install. That is survivable while one person runs one agent. It stops the moment five Agents execute in parallel and nobody can say which is spending forty times more than it needs to on a docstring, or which is quietly failing a job it was never strong enough to hold.
On 2026-09-22 the spread between tiers got wide enough to matter. Anthropic shipped Claude Opus 5.5. OpenAI shipped GPT-6 Sol and GPT-6 Luna. The gap between the cheapest and the most expensive of the three is a factor of forty on input tokens, and every one of them is good enough to hand a real Task.
A routing rule is a written statement, not a preference
A routing rule is a written statement of which class of Task runs on which Runtime, attached to the Agent that executes the work rather than to the person who happened to start the run.
That is the whole idea, and the important half is “written”. A preference lives in one person’s terminal and expires when they go on leave. A rule lives next to the Agent, so an Assignee picking up a Task inherits it, a reviewer can see which tier produced the diff they are reading, and changing your mind next quarter is one edit instead of six conversations.
A routing rule is not a ranking of models. It does not claim a winner, and it does not need one. It says where each kind of work goes by default, and what happens when that default turns out to be wrong.
What the three tiers cost
| Claude Opus 5.5 | GPT-6 Sol | GPT-6 Luna | |
|---|---|---|---|
| API id | claude-opus-5-5 |
gpt-6-sol |
gpt-6-luna |
| Input / output per 1M | $4 / $20 | $2 / $10 | $0.10 / $0.50 |
| Cached input | $0.20 read, $5 write | 90% off reads | 90% off reads |
| Context | 1M | 872k | 1M |
For reference: Claude Opus 5 at $5 / $25, and GPT-6 Astra at $10 / $50. OpenAI describes Sol and Luna as 50% cheaper than GPT-5.6 promotional pricing. The word promotional is OpenAI’s own and it changes the meaning: the baseline for that headline is a discounted rate, not a list rate. OpenAI also says Astra “continues to be our best model across the board”, which is a useful reminder that Sol is the value tier of its own family rather than its flagship.
One practical check before you write a tier into a rule: OpenAI lists Sol and Luna in ChatGPT Work and in Codex for Plus, Pro, Business, Enterprise and Edu, with Free and Go getting Luna in the desktop app, and says neither is in Chat yet. Availability shapes a routing rule as much as price does, because a tier half your Crew cannot reach is not a tier.
Route on cost per completed task, not on the token rate
Token price is a proxy for the number you actually pay. A model at a fifth of the rate that needs four attempts and a human rescue is not cheaper. OpenAI published cost per task alongside score for this launch, which is unusually direct and is the right shape of evidence for a routing decision.
| Benchmark, as OpenAI published it | Result | Cost per task, relative |
|---|---|---|
| AutomationBench 1.0.6, Sol at extra-high | 33.2% | $0.27 per task |
| AutomationBench 1.0.6, GPT-6 Astra at low | 30.3% | 3.9x Sol |
| AutomationBench 1.0.6, Claude Opus 5 at max | 26.9% | 11.1x Sol |
| AutomationBench 1.0.6, Fable 5.1 with Opus 5 fallback at max | 31.4% | more than 8.9x Sol |
| Agents’ Last Exam V1, Sol at max | 56.4%, above Opus 5’s best | 60% lower than Opus 5 |
| OSWorld 2.0 offline, Sol at extra-high | 60.5% against Opus 5 at medium on 60.3% | about 80% cheaper |
| DeepSWE 1.1, Luna at max | 66.6% | 93% cheaper than Opus 5 |
| DeepSWE 1.1, Sol at max | 68.8%, within 1.1pp of Fable 5 at extra-high | about 80% cheaper |
Two things are true of that table at once. The margins are large, and they are measured against Claude Opus 5, not Opus 5.5, which did not exist when OpenAI’s post was written and shipped hours later at 40% less to run than Opus 5 by Anthropic’s own account. Nobody has re-run the comparison. Treat the rows as the last published numbers, and the ranking they imply as a hypothesis to test on your own repository.
What survives that caveat is the shape, and the shape is what a routing rule needs. The cheap tiers hold their own on bounded, repeatable work at a small fraction of the cost. Anthropic’s own claims for Opus 5.5 are not about price per task at all: 40% less to run than Opus 5, output about 30% faster, 128k maximum output, and staying on task for 18 hours or more. Those describe a model you point at long, ugly work, not one you point at a hundred small jobs.
The rule, written out
Here is the artifact. It is deliberately boring, and it belongs where the team reads it, not in a private config file on one laptop.
ROUTING RULE (owner: platform, review: end of Sprint)
Tier 1 Runtime: Codex on gpt-6-luna
Triage, labelling, changelog and release notes, first-pass review
comments, test scaffolding, dependency bumps, doc edits.
Rule of thumb: if a wrong answer costs a reviewer one minute, route here.
Tier 2 Runtime: Codex on gpt-6-sol <- default
Bounded feature work, bug fixes with a reproduction, refactors inside
one module, anything repeated many times a day and cheap to retry.
Tier 3 Runtime: Claude Code on claude-opus-5-5
Migrations, cross-cutting refactors, work that must hold context for
hours, anything where stopping halfway costs more than the tokens.
ESCALATION
A Tier 1 or Tier 2 Task that fails twice returns to the Task with its
log and does NOT silently re-run on Tier 3. A person moves it.
EXCEPTIONS
Named on the Task, with a reason, by the Assignee. Not in a DM.
Three lines do the real work. The default tier is named, so a Task nobody thought about lands somewhere sensible. Escalation is explicit, so a cheap failure cannot cascade into an expensive one unnoticed. Exceptions are recorded on the Task, which is the difference between a rule and a suggestion.
Escalation is the half teams skip
The savings in the table above are real per attempt. They evaporate if a cheap attempt routinely becomes an expensive retry plus a human rescue, because you then pay for both runs and the interruption. This is the failure mode of automatic fallback chains: they hide the cost of the cheap tier inside the bill of the expensive one, and nobody can see the exchange rate.
So make the escalation manual and visible. A failed run returns to the Task with its log, its diff so far, and the question it got stuck on. A person decides whether to re-route it up a tier, split it, or clarify it. Agents research, execute, test, and report. People set direction, grant authority, and accept the result. If your escalation path is a retry loop, you have automated exactly the decision that needed a human.
One more thing worth watching before you route latency-sensitive work down a tier. Artificial Analysis, a third party rather than either vendor, measured the “max” reasoning variants at 115.2 output tokens per second for Sol with 102.15 seconds to first token, and 153.9 tokens per second for Luna with 124.23 seconds. Confirm those against the current listing before you rely on them. [VERIFY] The shape is the point: at maximum reasoning effort the cheap models are not the fast-to-start models, and an Agent that is thinking for a hundred seconds looks exactly like an Agent that has hung. Whatever tier you route to, how the Runtime signals progress and failure decides whether anyone notices the difference.
When not to write a routing rule
If you run one agent on one repository, do not do any of this. Pick the tier that fits your work and move on. Routing costs a document to maintain and an argument to have, and it pays back only when several Agents run at once and their cost or failure pattern differs.
The threshold is roughly this: write the rule when two or more people can start a run, or when one budget covers several concurrent Agents. Below that, the rule is overhead. Above it, its absence is the reason nobody can explain last month’s invoice.
The number that does not move
Cost per completed task fell hard on both sides of this launch. For scale, OpenAI reports its median researcher spending over $600 a day on coding agents, with the 90th percentile at $7,000. Teams below that line just got several times more finished agent work for the same budget, which is worth tracking per Runtime rather than per invoice.
They got zero additional reviewer hours. Routing is a cost lever and a reliability lever, and it is not a throughput lever, because the queue of work waiting for a person to accept it is unchanged. What routing can do is make that queue legible: which tier produced this diff, what it cost, how many attempts it took, and whether the Task that failed twice is sitting in somebody’s Inbox or nowhere at all. Keep the goal, the run, the cost, the blockers and the review on the Task from request to release, and the routing rule becomes a setting the team edits rather than a habit six people hold privately.
Sharkly does not run the model and does not pick it for you. It is the shared Task, Computer, context and review layer around the tools that do, which is where a routing rule has to live if it is going to outlast the person who wrote it. If you are already running several coding agents, which model gets which task and which agent gets it are the same decision seen from two sides. Connect one Computer, create one Agent per tier, and give the cheap one something boring.



