GPT-6 Sol at $0.27 a Task: What Your Fixed Agent Budget Now Buys

OpenAI put GPT-6 Sol at $0.27 per AutomationBench task against Claude Opus 5 at 11.1x that. What a fixed monthly agent budget now buys, and why the new ceiling is your reviewers.

Mia Parker

Mia Parker

23 September 2026

GPT-6 Sol at $0.27 a Task: What Your Fixed Agent Budget Now Buys

Most teams running several coding agents set a monthly ceiling on model spend at some point, usually after a bad week. Pick a number, cap concurrency until the invoice stays under it, revisit next quarter. On 2026-09-22 that number quietly stopped meaning what it meant.

OpenAI shipped GPT-6 Sol at $2 per million input tokens and $10 per million output, API id gpt-6-sol, and published something more useful than a token rate: cost per task. On AutomationBench 1.0.6, Sol at extra-high reasoning effort scores 33.2% at $0.27 per task. Claude Opus 5 at max effort scores 26.9% at 11.1 times Sol’s cost per task on the same harness.

So work out what your cap buys now, before you raise it. The answer for most teams is that it already buys more finished runs than anyone can read. Which means the cap was never protecting the budget. It was protecting your reviewers, by accident, and it is about to stop.

TL;DR

Cost per accepted change is what your team spends on model inference for every diff a person actually reads and accepts, and it is the only agent cost figure that connects to shipped work. At published rates, a four-reviewer team’s entire monthly reviewable throughput costs roughly $136 of GPT-6 Sol. The binding constraint moved from the invoice to the acceptance step, and no model release touches that. Agents research, execute, test, and report. People set direction, grant authority, and accept the result.

The numbers OpenAI actually published

OpenAI reports score alongside cost per task, which is unusually honest and much more useful than a price list.

Model, at the effort level OpenAI used AutomationBench 1.0.6 Cost per task
GPT-6 Sol, extra-high 33.2% $0.27
GPT-6 Astra, low 30.3% 3.9x Sol
Claude Opus 5, max 26.9% 11.1x Sol
Fable 5.1 with Opus 5 fallback, max 31.4% more than 8.9x Sol

OpenAI’s own summary of the first and third rows: Sol at extra-high beats Opus 5 at max for about 9% of Opus 5’s cost per task. The same pattern shows up elsewhere in its post. On Agents’ Last Exam V1, Sol at max scores 56.4%, above Opus 5’s best, at 60% lower cost per task. On DeepSWE 1.1, Sol at max scores 68.8%, within 1.1 percentage points of Fable 5 at extra-high (69.9%), about 80% cheaper per task.

Two qualifiers belong with those figures and rarely travel with them. OpenAI’s headline is that Sol and Luna are each 50% cheaper than GPT-5.6 promotional pricing, and “promotional” is OpenAI’s word: the comparison is against a discounted rate. And OpenAI benchmarked against Claude Opus 5, because Claude Opus 5.5 shipped hours later at $4 and $20 per million against Opus 5’s $5 and $25. The Opus 5 column is the last published comparison, not the current one.

Converting a budget into runs

Multiply $0.27 by 11.1 and Opus 5 at max lands near $3.00 per task on that harness. That arithmetic is ours, from OpenAI’s two figures. Divide a monthly cap by each and you get the table that changes how the cap reads.

Monthly model budget Tasks on Sol, extra-high Tasks on Astra, low Tasks on Opus 5, max
$250 about 925 about 240 about 83
$1,000 about 3,700 about 950 about 330
$5,000 about 18,500 about 4,750 about 1,660

A benchmark’s cost per task is not your cost per task. AutomationBench tasks are benchmark shaped: bounded, short, and nothing like a real ticket in a repository with years of history. Your own tasks may burn ten or a hundred times the tokens. What travels between the benchmark and your repository is the ratio between models, not the absolute dollar. Treat the table as the shape of the change, then measure your own cost per run across agents and rescale it.

There is a second correction, and it cuts the other way. Cost per task is per attempt, not per success. At the benchmark’s own 33.2%, roughly two attempts in three do not land, so Sol costs about $0.81 per benchmark success and Opus 5 at max about $11.15. Ours again, and rough, because a benchmark pass rate is not your team’s acceptance rate. The direction is what matters: the cost of trying something twice is now small enough that retry strategy is a design choice rather than a budget decision.

Where the ceiling sits now

Put the review side on the same page. A reviewer who properly reads a diff, understands what the Agent chose, and takes responsibility for the result gets through some number of changes in a working day. Call it six, and adjust it to your team.

Reviewers Accepted changes per month Sol cost at $0.27 Opus 5 max cost at $3.00
1 about 126 $34 $378
2 about 252 $68 $756
4 about 504 $136 $1,512
8 about 1,008 $272 $3,024

Six accepted changes per reviewer per day over 21 working days, our assumption and our arithmetic. Put in your own numbers and the conclusion survives: everything a normal engineering team can actually accept in a month costs a few hundred dollars of inference, even at the expensive end. A $1,000 cap on a four-person team is not a constraint. It is about seven times more agent output than the team can absorb.

That is the whole point of the price cut for a team that reviews its work. It did not raise your output. It removed the last excuse for the queue being short, and made visible a bottleneck that was always the human review step.

OpenAI’s internal spend is the useful counter-example, not a contradiction. Its median researcher runs more than $600 a day of coding-agent inference at API prices and the 90th percentile runs $7,000 a day. Those are real numbers and they dwarf the table above, because that work has a different shape: enormous fan-out, automated scoring, and no step where a named person accepts each result into a production branch. If your process has that step, your model bill will stay small and your queue will not.

What to buy with the headroom

Once the budget is not binding, the question stops being how many runs to buy and becomes what to spend the difference on. Two answers reduce the work arriving at a reviewer instead of increasing it.

Buy effort, not volume. Reasoning effort is a cost dial, and OpenAI’s own table uses different levels for different benchmarks. The $0.27 figure is the extra-high row, so the premium for a more careful run is already inside a number that is 9% of the alternative. Paying it on the same number of tasks costs almost nothing and produces diffs that take less time to check.

Buy cache, not tokens. GPT-6 discounts cached input reads by 90%, keeps hit rates higher by default, and no longer invalidates the cache when you change reasoning effort or tool availability. Explicit breakpoints let you decide where a cached prefix ends. OpenAI cites GitHub seeing more than 50% fewer prompt tokens needing fresh processing across billions of requests. For a Crew this is the difference between the first Agent on a repository and the tenth: they re-send nearly the same prefix of layout, house rules, and conventions, and that prefix is exactly what a cache is for. The token cost of repository context is where a fleet’s bill actually lives.

The honest decision pair: raise the number of concurrent runs when the extra runs converge on one result a person accepts, such as three attempts at one migration where a test suite picks the winner. Do not raise it when each run produces its own diff, its own summary, and its own acceptance decision, however cheap the runs have become. In the second case the invoice is the only thing that scaled well.

Make the new ceiling countable

The budget was a bad proxy for capacity, but it had one virtue: it was a number somebody watched. Replace it with the number that now binds, which is completed runs waiting for a person.

That is only measurable if runs stop living in private terminals. Sharkly is not a replacement for Codex, Claude Code, or the other tools running these models. It is the layer around them, where context, progress, blockers, results, and human review stay visible from request to release. Execution and blockers return to the Task, so “waiting for a human” becomes a state you can count on a board rather than a feeling you get on Friday afternoon. The review-capacity arithmetic only works if you can see the queue.

If you are already running several coding agents, the move this week is not to raise the cap. It is to find out how many finished runs your team accepts in a normal week, set concurrency against that, and spend the leftover budget on runs that are easier to accept.

Explore more

What It Actually Costs to Run an Agent Crew on Free and Cheap Tiers

What It Actually Costs to Run an Agent Crew on Free and Cheap Tiers

Which models a Crew can run for free after the 2026-09-22 launches, why a free desktop entitlement is not an API key, and what the cheap API path really costs per week.

23 September 2026

When a Cheaper Model Costs More: Price Agent Runs Per Accepted Change

When a Cheaper Model Costs More: Price Agent Runs Per Accepted Change

GPT-6 Luna at $0.10 per million tokens lowers the price of an attempt and raises attempts, output and decisions per shipped change. Price agent runs per accepted change, not per token.

23 September 2026

GPT-6 Luna for the Boring 80 Percent: What to Hand a Cheap Runtime

GPT-6 Luna for the Boring 80 Percent: What to Hand a Cheap Runtime

GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output. Which triage, labelling, first-pass review and test scaffolding belongs on a cheap runtime, and which does not.

23 September 2026