Anthropic shipped Claude Sonnet 5.5 on September 28 at $2 per million input tokens and $10 per million output. OpenAI shipped GPT-6 Sol on September 22 at the same $2 and $10. Six days apart, identical list price, and each sits one step below its lab’s flagship. For a team choosing the default model for coding agents, the price line settles nothing. The real choice is Claude Code vs Codex as the Runtime around each model, and what each pair costs per finished task.
An honest trial means two Runtimes, two configurations, two sets of results, and one reviewer trying to remember which terminal produced which diff. Picking the model is a one-line setting. Comparing the two is agent chaos in miniature.
TL;DR
Claude Sonnet 5.5 and GPT-6 Sol cost the same per token and differ in the Runtime a team usually runs them in: Claude Code for Sonnet 5.5, Codex for Sol. Anthropic’s launch table has four rows with a GPT-6 Sol column. Sonnet 5.5 leads three clearly; the fourth, FrontierCode, depends on effort. Cost per task flips by benchmark: on FrontierCode Sonnet 5.5 matches Sol’s best for about a fifth of the cost; on AA-Briefcase Sol’s best costs $2.67 a task, Sonnet 5.5’s $29.19. It is one vendor’s table. Settle Sonnet 5.5 vs GPT-6 Sol by running the same Task through both Runtimes on your own code and counting review minutes per accepted change. Keeping both is a legitimate answer if the Task stays the record.
What the published numbers actually let you compare
Four rows. The Claude Sonnet 5.5 launch post compares Sonnet 5.5, Sonnet 5, Opus 5.5 and GPT-6 Sol, and only these rows carry a Sol number:
| Row on Anthropic’s table | Run by | Sonnet 5.5 | GPT-6 Sol |
|---|---|---|---|
| FrontierCode 1.1 (Main) | Cognition: Claude models in Claude Code, GPT models in Codex CLI | 46.2% at max, 52.1% at xhigh | 49.3% |
| GDPval-AA v2.1 | Artificial Analysis | 1844 | 1487 |
| AA-Briefcase v1.1 | Artificial Analysis | 1811 | 1483 |
| Chartography, no tools | Surge AI’s benchmark; Sonnet 5.5 graded by Anthropic, Sol as Surge AI reported it | 61.6% | 53.6% |
This is one vendor’s table. OpenAI’s Sol post went up six days before Sonnet 5.5 existed, so the only vendor table with both models is Anthropic’s. Both columns carry a bug caveat: Sonnet 5.5’s Artificial Analysis runs used a pre-release deployment with a structured-output bug, since fixed, and OpenAI recently fixed an image-understanding bug in Sol that the Artificial Analysis and Surge AI scores may not reflect.
The FrontierCode row is the one a coding team should read, because it is already a Runtime comparison: 52.1% against 49.3% is Claude Code on Sonnet 5.5 against Codex CLI on Sol, the pair you would actually deploy. The footnote on the max score matters as much. FrontierCode penalizes out-of-scope changes, and at max effort Sonnet 5.5 more often ran Claude Code’s code-review skill, which fans out to many subagents. In two cases Cognition examined, that caused a timeout or edits outside the task’s scope. More effort, wider diff, more review.
Nothing else on the page has a Sol number. The Terminal-Bench 4.0 and CursorBench 4.0 effort charts plot GPT-5.6 Sol, because no GPT-6 Sol figures were published.
One third-party aggregate exists. Artificial Analysis puts Sonnet 5.5’s max-effort endpoint at 56 on its Intelligence Index, third of 216, as of September 29. Sol stood at 48 on September 23, when we last recorded it. Two dates, a moving index, and neither number measures your repository.
Cost per finished task, not per token
Identical token prices say little about the bill, because tokens per task depend on model, effort and Runtime. Anthropic’s launch page publishes score and cost per task by effort for two of the shared rows:
| Effort | FrontierCode: Sonnet 5.5 | FrontierCode: GPT-6 Sol | AA-Briefcase: Sonnet 5.5 | AA-Briefcase: GPT-6 Sol |
|---|---|---|---|---|
| low | 29.3%, $0.19 | 37.3%, $0.43 | 1264, $0.87 | 905, $0.12 |
| medium | 36.5%, $0.24 | 45.9%, $0.77 | 1461, $1.64 | 1142, $0.34 |
| high | 49.4%, $0.42 | 47.7%, $1.04 | 1634, $3.95 | 1289, $0.63 |
| xhigh | 52.1%, $1.59 | 48.4%, $1.32 | 1746, $9.63 | 1364, $1.19 |
| max | 46.2%, $20.78 | 49.3%, $2.07 | 1811, $29.19 | 1483, $2.67 |
Read FrontierCode first. Sonnet 5.5 at high scores 49.4% for $0.42 a task, level with Sol’s best of 49.3% at $2.07; Anthropic’s summary is “about a fifth of the cost per task”. The detail to act on sits lower: at medium, Sonnet 5.5 scores 36.5%, below Sol even at low, and medium is Claude Code’s default effort for Sonnet 5.5. Anthropic’s prompting guide suggests starting at high, or medium for agentic coding on well-specified tasks. The setting nobody touches can decide the comparison.
AA-Briefcase, a knowledge-work benchmark, runs the other way. Sol’s best costs $2.67 a task; Sonnet 5.5’s best costs $29.19, about eleven times as much, for a higher score. The Sonnet 5.5 setting nearest Sol’s max in cost, medium at $1.64, scores 1461 against 1483. Same two models, opposite answer on which is cheaper for the result.
Caching narrows the gap and moves it. Sonnet 5.5 reads cached input at $0.20 per million, with 5-minute cache writes at $2.50 (Anthropic pricing). GPT-6 takes 90% off cached input reads, also $0.20 on a $2 rate. The difference is what breaks the cache: OpenAI says changing reasoning effort or tool availability no longer invalidates it, while Anthropic’s Sonnet 5.5 prompting guide says changing top-level effort between requests does, and points to a beta per-message effort setting. Agents that change effort mid-task will see that on the bill.
None of these figures is your cost per accepted change. A cheap run that needs two reruns and a long review is the expensive one.
How to decide on your own codebase
Run the same Task through both Runtimes and count what a person spends accepting the result. The criteria are in how to compare coding agents on your own codebase; here is the setup for this pair.
Sharkly does not pick models. The Agents documentation is explicit: “The Agent follows the Runtime’s default model. It does not promise or display a specific model name. Change models in the Runtime or Computer tool configuration, not on the Agent.” So the model work happens on the Computer:
- Claude Code on
claude-sonnet-5-5. It needs Claude Code v2.1.284 or later. Claude Code’sdefaultmodel is Opus 5.5 on Pro, Max, Team, Enterprise and the Anthropic API, so a Computer runs Opus until someone changes it, and thesonnetalias means Sonnet 5.5 only on the Anthropic API. Pin the full id and confirm withclaude --model claude-sonnet-5-5on that host. Running Sonnet 5.5 agents in Sharkly has the full path. - Codex on
gpt-6-sol. Set the model in Codex’s configuration on that Computer; running GPT-6 Sol agents in Sharkly has the steps. - Temporary working directory on both Agents. Each run gets an isolated directory and repository-backed runs can prepare a fresh worktree, so the two attempts never share a checkout.
Then write the trial down before either run starts:
Task: SH-312, one real Task, same description and starting commit for both runs
Run A: Claude Code Agent, claude-sonnet-5-5, effort: medium (Claude Code default) or high
Run B: Codex Agent, gpt-6-sol, effort: whatever Codex on that Computer is set to, written here
Accept: the same acceptance test for both
Record: accepted y/n, human review minutes, reruns, files touched outside scope,
provider cost from your own Anthropic and OpenAI bills
Decide: after several Tasks of each kind you actually assign, not one
Assign the Task to one Agent and start the other by mentioning it in a comment; Sharkly lets an explicit mention start an Agent that is not the Assignee. Both runs return to the same Task, so two diffs, two test runs and any questions sit in one timeline instead of two scrollbacks. The single-bug version, with a scoring rubric, is Claude Code vs Codex on the same bug.
Weight review minutes most. Anthropic’s prompting guide says Sonnet 5.5 adds tests, docs and small files at every effort level, checks in early on long tasks at low and medium, and starts extra review rounds or reviewer subagents at xhigh and max. Whatever Codex on Sol does in the same spots, the trial will show it. All of it lands in review time, and none of it appears in a benchmark score.
Why “both” is a legitimate answer
The published evidence does not produce a winner for everyone, and your trial may not either. A split result is a finding, not a failure.
Pick one Runtime when the trial shows no difference that survives the review-minute count, or when a small team would rather hold one subscription and one configuration. If you only use Claude Code and like a terminal-first workflow, you may not need a management layer at all. Keep both when results split by task type, or when you want to change the default next month without rebuilding the workflow. Next month is not hypothetical: Opus 5.5, Sol and Luna shipped on one day, and Sonnet 5.5 six days later.
Keeping both works only if the Task, not the Runtime, is the record. No other model reads Sonnet 5.5’s thinking blocks, so when a Task moves from a Claude Code Agent to a Codex Agent, the reasoning stays behind and only what is written on the Task travels. Decisions, assumptions and open questions belong in the description or a comment. Running Claude Code and Codex together on one codebase covers the daily version, and Sonnet 5.5 or Opus 5.5 covers routing on the Claude side.
Agents research, execute, test, and report. People set direction, grant authority, and accept the result. The Runtime can change every quarter; context, progress, blockers, results and human review should stay visible in one place from request to release.
Last week’s version of this question, Opus 5.5 vs GPT-6 Sol, had the mirror-image evidence problem: OpenAI’s post benchmarked Sol against a Claude model superseded hours later. Write the default down, measure it on your code, and keep the Runtime a setting you can edit. If you want to keep using multiple coding agents rather than committing to one ecosystem, that’s the problem Sharkly is designed around. Connect one Computer, create one Agent per Runtime, and give both the same real Task.



