Sonnet 5.5 or Opus 5.5: Which Agent Gets the Task?

Sonnet 5.5 costs half of Opus 5.5, but effort decides cost per task. A written routing rule for coding agents: which model, what effort, when to escalate.

Sophia Carter

Sophia Carter

29 September 2026

Sonnet 5.5 or Opus 5.5: Which Agent Gets the Task?

Anthropic shipped Claude Sonnet 5.5 on September 28, six days after Opus 5.5. At $2 per million input tokens and $10 per million output it costs half of Opus 5.5’s $4 and $20, and on the launch page it lands within a few points of Opus 5.5 on most rows and above it on Terminal-Bench 4.0, a benchmark Anthropic ran itself. For a team running several coding agents, the tempting move is to put every Agent on Sonnet and bank the difference.

The per-effort charts on the same page argue against that. The Terminal-Bench headline was scored at Sonnet’s top effort, at $12.54 an attempt. At settings you would run all day, the order flips: Opus 5.5 at medium scores 57.6% for $2.94, Sonnet 5.5 at high scores 43.0% for $1.94, and Sonnet only passes Opus-at-medium at xhigh, with 61.5% for $5.30.

TL;DR

Default well-scoped work to Sonnet 5.5 at medium. Send long, ambiguous or cross-cutting work to Opus 5.5 rather than turning Sonnet up to xhigh: on all three per-effort coding charts Anthropic published, Opus at high outscores Sonnet at xhigh for about the same cost or less. Write it down as a rule on the Agent, not a preference in someone’s head: one Agent per model-and-effort tier, named escalation conditions, a definition of done. Sharkly doesn’t pick the model or the effort; it keeps the record of which Agent did what on the Task.

What the launch numbers support

The launch table, with the effort behind each cell and who ran it:

Benchmark Reported by Sonnet 5.5 Opus 5.5
Terminal-Bench 4.0 Anthropic 70.6% (max) 66.4% (xhigh)
FrontierCode 1.1 Cognition 52.1% (xhigh), 46.2% (max) 54.4% (max)
CursorBench 4.0 Cursor 55.5% (max) 57.8% (max)
SWE-Bench Pro Anthropic system card, max effort 81.3 89.9
HLE, with tools Anthropic 64.5% 67.7%
OSWorld 2.1 Anthropic 80.1% 81.8%

Two caveats on the Terminal-Bench row. The system card says flagged requests affected 10% of Opus 5.5’s trials against 1.5% of Sonnet’s, and Artificial Analysis’s own run measured Sonnet 5.5 at 63.6%, not 70.6%. SWE-Bench Pro, meanwhile, shows an 8.6-point gap in Opus’s favor. Anthropic’s positioning agrees: Sonnet 5.5 for well-scoped everyday tasks and bug fixes, while Opus 5.5 “remains clearly stronger at complex, open-ended work.”

The headline cells also hide what a routing rule needs: the three coding benchmarks at efforts you would actually assign, as score and cost per task.

Benchmark Sonnet 5.5 medium Sonnet 5.5 high Sonnet 5.5 xhigh Opus 5.5 medium Opus 5.5 high
Terminal-Bench 4.0 28.8%, $0.83 43.0%, $1.94 61.5%, $5.30 57.6%, $2.94 64.2%, $3.88
CursorBench 4.0 39.2%, $0.70 47.8%, $1.67 53.1%, $3.88 52.5%, $2.91 56.0%, $3.97
FrontierCode 1.1 36.5%, $0.24 49.4%, $0.42 52.1%, $1.59 54.6%, $0.80 54.0%, $1.09

Sonnet 5.5 is the cheapest attempt at medium and high, where its scores are also the lowest. Past high, the lines cross: Sonnet at xhigh costs more than Opus at medium on all three rows, and loses to Opus at high on all three for about the same money or less. The cheap model stops being cheap exactly when you ask it to act like the expensive one. Sonnet’s CursorBench costs are Anthropic’s estimate, and none of these tasks are your repository, so treat the crossover as a hypothesis to confirm.

The routing rule, written down

Earlier posts argued for matching models to Agent roles, for deciding which agent gets a task by its shape, and for writing tiers down as a rule. Sonnet 5.5 adds a column. On Terminal-Bench the same model at medium and at max is about 15 times apart per attempt, $0.83 against $12.54, so “route this to Sonnet” is not a rule until it names an effort.

Task shape Agent (runtime model) Effort Why
Bug fix with a reproduction and a named test Sonnet 5.5 medium Anthropic’s prompting guide suggests medium for agentic coding on well-specified tasks
Test scaffolding, docs, changelog, dependency bumps Sonnet 5.5 medium Scoped, cheap to check, under a dollar a task on all three charts
Bounded feature inside one module, clear acceptance Sonnet 5.5 high Still cheaper than Opus at medium on every chart; the top of Sonnet’s useful range
Overnight, high-volume triage Sonnet 5.5 low, plus a verification step Cheapest per run, but at low it can report done without running tests
Long or ambiguous refactor, migration, cross-cutting change Opus 5.5 medium or high At high, beats Sonnet at xhigh on all three charts for similar or lower cost
Planning: turning a vague request into Tasks Opus 5.5 high Complex, open-ended work is where Anthropic says Opus stays ahead
Review of another Agent’s output Opus 5.5, then a person high A second model checks the evidence; a person accepts the result

The escalation line is where the rule earns its keep. When a Sonnet Agent struggles, the reflex is to raise its effort. The charts say change the model: a Task that needs Sonnet at xhigh is, on this evidence, a Task for Opus at high. Set your own conditions; these show the shape:

  • The same check failed twice. A person moves the Task to the Opus Agent with the log attached; nothing silently retries on a bigger tier.
  • The change needs more files than a worker is allowed (say eight), or touches a path the Task doesn’t name.
  • The Agent asks a design question, not a question about the Task’s wording.
  • A class of Task only passes with Sonnet above high. Move the class to the Opus Agent next Sprint.

It works in reverse: Opus Tasks that keep finishing in one attempt with small diffs belong one tier down.

What each model does to the review queue

A rule that names only model and effort is half a rule. The other half is what “done” looks like, because each tier hands a reviewer something different, and review was already the bottleneck before either launch. Cheaper runs add diffs, not reviewer hours.

Anthropic’s prompting guide for Sonnet 5.5 is direct about what a reviewer will see:

  • At low and medium, it checks in early on long agentic tasks: on a board, a Task waiting for a human reply at 2 a.m. One Hacker News commenter blamed a 7.4% score on a custom benchmark (Sonnet 5 got 17.8%) on exactly that.
  • At low, it can skip verification and report code as done without running tests. A triage Agent’s “done” means nothing without the check output.
  • At every effort, it adds tests, docs and small files nobody asked for. Some is welcome, and all of it is more diff to read.
  • At xhigh and max, it adds review rounds or reviewer subagents. The launch page says that is why Sonnet 5.5 scored lower on FrontierCode at max than at xhigh: it more often ran Claude Code’s code-review skill, which fans out to many subagents, and in two cases Cognition examined this caused a timeout or out-of-scope edits. Those are the costliest thing to hand a reviewer, because someone has to find them line by line.

Opus 5.5 costs the queue something else. Anthropic says it stays on task 18 hours or more with up to 128k tokens of output, so its burden is size, not noise, and the handoff packet a reviewer needs applies to it too.

So the definition of done belongs in the worker Agent’s instructions, which Sharkly includes in every run:

Agent: sonnet-worker  (Claude Code, claude-sonnet-5-5, effort medium)

Done means:
- The checks named in the Task ran. Paste each command and its exit
  code. "Tests pass" without output is not done.
- Changes stay inside the paths the Task names. List every file you
  created that the Task did not ask for, and why.
- New tests are welcome. New docs files only if the Task asks.

Keep working on long Tasks. Ask only when two readings of the Task
would produce different code; otherwise proceed and record the
assumption in your report.

Stop and return to the Task, without retrying, when a check has failed
twice, the change needs more than 8 files, or the question is about
design rather than the Task's wording.

Running both tiers in Sharkly

Sharkly doesn’t choose the model, and it doesn’t set effort. The Agents documentation is explicit: “The Agent follows the Runtime’s default model. It does not promise or display a specific model name. Change models in the Runtime or Computer tool configuration, not on the Agent.” There is no thinking-depth control on the Agent either. A tier is a Claude Code configuration on a Computer; the Agent is how the team assigns work to it.

So: one Agent per tier. sonnet-worker runs on a Computer whose Claude Code is set to claude-sonnet-5-5, opus-lead on one set to claude-opus-5-5. A Computer can be a container, so a second tier is not a second laptop. Three details from Claude Code’s model configuration docs decide whether a tier is what its name says:

  • Pin the full id. The sonnet alias means Sonnet 5.5 only on the Anthropic API (Bedrock and Google Cloud still resolve it to Sonnet 4.5), and Sonnet 5.5 needs Claude Code v2.1.284 or later.
  • Claude Code’s default model is Opus 5.5 on Pro, Max, Team, Enterprise and the API, so a Computer nobody configured is already an Opus tier.
  • Sonnet 5.5’s default effort in Claude Code is medium, the worker tier above. Changing it happens on the Computer, not on the Agent.

Sharkly cannot reliably read which model a Runtime is using, so put the tier in the Agent’s name and change name and configuration together. The Sonnet 5.5 setup guide has the steps.

For work that needs both tiers on one Task, use a Crew. A Crew is a reusable group of People and Agents coordinated by one leader Agent. With opus-lead as leader and Sonnet workers as members, the leader reads the Task, handles it or mentions a member for a piece, and members add their results to the same Task. That is the Opus-orchestrator shape Hacker News commenters suggested, with the handoffs where a person can read them. For small, well-defined work, skip the Crew and assign sonnet-worker directly.

Claude Code’s opusplan alias does a version of this in one session: Opus in plan mode, Sonnet for execution. Its Sonnet half means Sonnet 5.5 only on the Anthropic API; on Bedrock, Google Cloud or Foundry it resolves to an older Sonnet. Use it when one person runs and reviews the session; use two Agents when someone else needs to see which one planned and which executed.

One thing never crosses between tiers: reasoning. Sonnet 5.5’s thinking blocks are bound to the model and the conversation. No other model reads them, and Sonnet 5.5 cannot read Opus 5.5’s. Crew directory reuse shares files, not the previous Agent’s session. What the leader decided reaches the worker only if it is written on the Task, which is the subject of its own post.

None of this changes who accepts the work. Agents research, execute, test, and report. People set direction, grant authority, and accept the result. The costs above land on your provider bill, since model usage continues through the subscriptions or API keys configured in those tools. The assignment, run log, questions, result and human review return to the Task whichever model leads next quarter, and the logic holds across vendors, as the Sonnet 5.5 and GPT-6 Sol comparison shows. If you want to keep using multiple coding agents rather than committing to one ecosystem, that’s the problem Sharkly is designed around.

Explore more

Model Handoffs Lose the Reasoning: What Claude Sonnet 5.5 Changes for Multi-Agent Work

Model Handoffs Lose the Reasoning: What Claude Sonnet 5.5 Changes for Multi-Agent Work

Claude Sonnet 5.5 binds thinking blocks to the model, conversation and account, so reasoning never survives an agent handoff. What the Task must carry instead.

29 September 2026

Claude Sonnet 5.5 vs GPT-6 Sol: Two Runtimes at the Same Price

Claude Sonnet 5.5 vs GPT-6 Sol: Two Runtimes at the Same Price

Claude Sonnet 5.5 and GPT-6 Sol both list at $2/$10. What the shared benchmarks show, why cost per task flips with effort, and how to test both on your code.

29 September 2026

Can You Run Coding Agents on Claude Sonnet 5.5 for Free?

Can You Run Coding Agents on Claude Sonnet 5.5 for Free?

Sonnet 5.5 is free to chat with on Claude.ai, but Claude Code is not on Free and a chat seat cannot start a run. The real free routes and the cheapest setup.

29 September 2026