Model Routing for Agent Crews: Which Tasks Go to Luna, Sol, or Opus 5.5

GPT-6 Luna, Sol and Claude Opus 5.5 are forty times apart on input price. Write down which Task class goes to which Runtime, where escalation stops, and why routing is a cost lever, not throughput.

Lucas Hayes

Lucas Hayes

23 September 2026

Model Routing for Agent Crews: Which Tasks Go to Luna, Sol, or Opus 5.5

Your team already routes work between models. It just does it in people’s heads. One engineer runs everything on the expensive tier because it fails less often. Another switched to a cheap one last month and told nobody. A third is still on whatever the tool defaulted to at install. That is survivable while one person runs one agent. It stops the moment five Agents execute in parallel and nobody can say which is spending forty times more than it needs to on a docstring, or which is quietly failing a job it was never strong enough to hold.

On 2026-09-22 the spread between tiers got wide enough to matter. Anthropic shipped Claude Opus 5.5. OpenAI shipped GPT-6 Sol and GPT-6 Luna. The gap between the cheapest and the most expensive of the three is a factor of forty on input tokens, and every one of them is good enough to hand a real Task.

A routing rule is a written statement, not a preference

A routing rule is a written statement of which class of Task runs on which Runtime, attached to the Agent that executes the work rather than to the person who happened to start the run.

That is the whole idea, and the important half is “written”. A preference lives in one person’s terminal and expires when they go on leave. A rule lives next to the Agent, so an Assignee picking up a Task inherits it, a reviewer can see which tier produced the diff they are reading, and changing your mind next quarter is one edit instead of six conversations.

A routing rule is not a ranking of models. It does not claim a winner, and it does not need one. It says where each kind of work goes by default, and what happens when that default turns out to be wrong.

What the three tiers cost

Claude Opus 5.5 GPT-6 Sol GPT-6 Luna
API id claude-opus-5-5 gpt-6-sol gpt-6-luna
Input / output per 1M $4 / $20 $2 / $10 $0.10 / $0.50
Cached input $0.20 read, $5 write 90% off reads 90% off reads
Context 1M 872k 1M

For reference: Claude Opus 5 at $5 / $25, and GPT-6 Astra at $10 / $50. OpenAI describes Sol and Luna as 50% cheaper than GPT-5.6 promotional pricing. The word promotional is OpenAI’s own and it changes the meaning: the baseline for that headline is a discounted rate, not a list rate. OpenAI also says Astra “continues to be our best model across the board”, which is a useful reminder that Sol is the value tier of its own family rather than its flagship.

One practical check before you write a tier into a rule: OpenAI lists Sol and Luna in ChatGPT Work and in Codex for Plus, Pro, Business, Enterprise and Edu, with Free and Go getting Luna in the desktop app, and says neither is in Chat yet. Availability shapes a routing rule as much as price does, because a tier half your Crew cannot reach is not a tier.

Route on cost per completed task, not on the token rate

Token price is a proxy for the number you actually pay. A model at a fifth of the rate that needs four attempts and a human rescue is not cheaper. OpenAI published cost per task alongside score for this launch, which is unusually direct and is the right shape of evidence for a routing decision.

Benchmark, as OpenAI published it Result Cost per task, relative
AutomationBench 1.0.6, Sol at extra-high 33.2% $0.27 per task
AutomationBench 1.0.6, GPT-6 Astra at low 30.3% 3.9x Sol
AutomationBench 1.0.6, Claude Opus 5 at max 26.9% 11.1x Sol
AutomationBench 1.0.6, Fable 5.1 with Opus 5 fallback at max 31.4% more than 8.9x Sol
Agents’ Last Exam V1, Sol at max 56.4%, above Opus 5’s best 60% lower than Opus 5
OSWorld 2.0 offline, Sol at extra-high 60.5% against Opus 5 at medium on 60.3% about 80% cheaper
DeepSWE 1.1, Luna at max 66.6% 93% cheaper than Opus 5
DeepSWE 1.1, Sol at max 68.8%, within 1.1pp of Fable 5 at extra-high about 80% cheaper

Two things are true of that table at once. The margins are large, and they are measured against Claude Opus 5, not Opus 5.5, which did not exist when OpenAI’s post was written and shipped hours later at 40% less to run than Opus 5 by Anthropic’s own account. Nobody has re-run the comparison. Treat the rows as the last published numbers, and the ranking they imply as a hypothesis to test on your own repository.

What survives that caveat is the shape, and the shape is what a routing rule needs. The cheap tiers hold their own on bounded, repeatable work at a small fraction of the cost. Anthropic’s own claims for Opus 5.5 are not about price per task at all: 40% less to run than Opus 5, output about 30% faster, 128k maximum output, and staying on task for 18 hours or more. Those describe a model you point at long, ugly work, not one you point at a hundred small jobs.

The rule, written out

Here is the artifact. It is deliberately boring, and it belongs where the team reads it, not in a private config file on one laptop.

ROUTING RULE  (owner: platform, review: end of Sprint)

Tier 1  Runtime: Codex on gpt-6-luna
  Triage, labelling, changelog and release notes, first-pass review
  comments, test scaffolding, dependency bumps, doc edits.
  Rule of thumb: if a wrong answer costs a reviewer one minute, route here.

Tier 2  Runtime: Codex on gpt-6-sol           <- default
  Bounded feature work, bug fixes with a reproduction, refactors inside
  one module, anything repeated many times a day and cheap to retry.

Tier 3  Runtime: Claude Code on claude-opus-5-5
  Migrations, cross-cutting refactors, work that must hold context for
  hours, anything where stopping halfway costs more than the tokens.

ESCALATION
  A Tier 1 or Tier 2 Task that fails twice returns to the Task with its
  log and does NOT silently re-run on Tier 3. A person moves it.

EXCEPTIONS
  Named on the Task, with a reason, by the Assignee. Not in a DM.

Three lines do the real work. The default tier is named, so a Task nobody thought about lands somewhere sensible. Escalation is explicit, so a cheap failure cannot cascade into an expensive one unnoticed. Exceptions are recorded on the Task, which is the difference between a rule and a suggestion.

Escalation is the half teams skip

The savings in the table above are real per attempt. They evaporate if a cheap attempt routinely becomes an expensive retry plus a human rescue, because you then pay for both runs and the interruption. This is the failure mode of automatic fallback chains: they hide the cost of the cheap tier inside the bill of the expensive one, and nobody can see the exchange rate.

So make the escalation manual and visible. A failed run returns to the Task with its log, its diff so far, and the question it got stuck on. A person decides whether to re-route it up a tier, split it, or clarify it. Agents research, execute, test, and report. People set direction, grant authority, and accept the result. If your escalation path is a retry loop, you have automated exactly the decision that needed a human.

One more thing worth watching before you route latency-sensitive work down a tier. Artificial Analysis, a third party rather than either vendor, measured the “max” reasoning variants at 115.2 output tokens per second for Sol with 102.15 seconds to first token, and 153.9 tokens per second for Luna with 124.23 seconds. Confirm those against the current listing before you rely on them. [VERIFY] The shape is the point: at maximum reasoning effort the cheap models are not the fast-to-start models, and an Agent that is thinking for a hundred seconds looks exactly like an Agent that has hung. Whatever tier you route to, how the Runtime signals progress and failure decides whether anyone notices the difference.

When not to write a routing rule

If you run one agent on one repository, do not do any of this. Pick the tier that fits your work and move on. Routing costs a document to maintain and an argument to have, and it pays back only when several Agents run at once and their cost or failure pattern differs.

The threshold is roughly this: write the rule when two or more people can start a run, or when one budget covers several concurrent Agents. Below that, the rule is overhead. Above it, its absence is the reason nobody can explain last month’s invoice.

The number that does not move

Cost per completed task fell hard on both sides of this launch. For scale, OpenAI reports its median researcher spending over $600 a day on coding agents, with the 90th percentile at $7,000. Teams below that line just got several times more finished agent work for the same budget, which is worth tracking per Runtime rather than per invoice.

They got zero additional reviewer hours. Routing is a cost lever and a reliability lever, and it is not a throughput lever, because the queue of work waiting for a person to accept it is unchanged. What routing can do is make that queue legible: which tier produced this diff, what it cost, how many attempts it took, and whether the Task that failed twice is sitting in somebody’s Inbox or nowhere at all. Keep the goal, the run, the cost, the blockers and the review on the Task from request to release, and the routing rule becomes a setting the team edits rather than a habit six people hold privately.

Sharkly does not run the model and does not pick it for you. It is the shared Task, Computer, context and review layer around the tools that do, which is where a routing rule has to live if it is going to outlast the person who wrote it. If you are already running several coding agents, which model gets which task and which agent gets it are the same decision seen from two sides. Connect one Computer, create one Agent per tier, and give the cheap one something boring.

Explore more

Sonnet 5.5 or Opus 5.5: Which Agent Gets the Task?

Sonnet 5.5 or Opus 5.5: Which Agent Gets the Task?

Sonnet 5.5 costs half of Opus 5.5, but effort decides cost per task. A written routing rule for coding agents: which model, what effort, when to escalate.

29 September 2026

Model Handoffs Lose the Reasoning: What Claude Sonnet 5.5 Changes for Multi-Agent Work

Model Handoffs Lose the Reasoning: What Claude Sonnet 5.5 Changes for Multi-Agent Work

Claude Sonnet 5.5 binds thinking blocks to the model, conversation and account, so reasoning never survives an agent handoff. What the Task must carry instead.

29 September 2026

Claude Sonnet 5.5 vs GPT-6 Sol: Two Runtimes at the Same Price

Claude Sonnet 5.5 vs GPT-6 Sol: Two Runtimes at the Same Price

Claude Sonnet 5.5 and GPT-6 Sol both list at $2/$10. What the shared benchmarks show, why cost per task flips with effort, and how to test both on your code.

29 September 2026