GPT-6 Luna for the Boring 80 Percent: What to Hand a Cheap Runtime

GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output. Which triage, labelling, first-pass review and test scaffolding belongs on a cheap runtime, and which does not.

Lucas Hayes

Lucas Hayes

23 September 2026

GPT-6 Luna for the Boring 80 Percent: What to Hand a Cheap Runtime

Look at what your crew produced last week. A small part of it was the feature discussed in standup. Most of it was the surround: bug triage, labels, a changelog, first-pass comments on eleven pull requests, test scaffolding for a module nobody wants to own, dependency bumps, a docs page that went stale when a response shape changed. High volume, low glory, and until this week it cost the same per token as the interesting work, because it ran on the same runtime.

OpenAI shipped GPT-6 Luna on September 22 at $0.10 per million input tokens and $0.50 per million output, API id gpt-6-luna, 1M context window. That is the tier the boring 80 percent belongs in. The useful question is not whether Luna is good, but which jobs move down there, which must not, and how to write a Task a cheap runtime can finish.

TL;DR

Route work to the cheap tier on checkability, not difficulty. If a result can be verified without redoing it, and a wrong answer costs a reviewer a minute rather than an afternoon, it belongs on Luna. Keep judgement calls, anything the board automates on, and anything nobody will check on the expensive runtime. Write acceptance criteria a machine can evaluate, keep effort modest for short jobs, and let every result return to the Task where a person accepts it.

What Luna is priced at, and what it is priced against

OpenAI’s published Luna comparisons are cost-per-task rather than leaderboard positions, the right frame for this work.

Published result Luna Compared against
DeepSWE 1.1 66.6% at max effort Comparable to Claude Opus 5 and Fable 5 at medium, at 93% lower cost per task than Opus 5, 96% lower than Fable 5
AutomationBench 1.0.6 +5.4pp over its predecessor at high effort 58% lower cost per task
OSWorld 2.0 offline Beats GPT-5.6 Sol at medium effort Roughly a tenth the cost
Factuality Matches GPT-5.6 Sol at higher effort Roughly a hundredth the cost

Two qualifiers travel with those numbers. OpenAI describes Sol and Luna as 50% cheaper than GPT-5.6 promotional pricing. Promotional is OpenAI’s word and it matters: the baseline is a discount, not list. Against the GPT-5.6 Luna list price our own launch coverage documented, $1 per million input and $6 per million output, the drop to $0.10 and $0.50 is nearer 90% on input and 92% on output. The real cut is larger than the headline, and the headline is measured against a sale.

And every Anthropic comparison in that table is against Claude Opus 5. Anthropic replaced Opus 5 with Opus 5.5 the same day OpenAI published, at $4 per million input and $20 per million output against Opus 5’s $5 and $25. Read those margins as measured against the previous generation, because they were. We claim no winner between the two current flagships, and neither vendor has published that head-to-head.

Artificial Analysis, a third party rather than either vendor, places Luna at 37 on its intelligence index against 58 for Opus 5.5. Luna is not a frontier model and is not sold as one. It is a competent worker priced like a utility, which is the point.

The property that decides what moves

Teams sort agent work by difficulty. That is the wrong axis: difficulty is a property of the problem, and your exposure is a property of the review loop. Sort by checkability instead.

Work belongs on the cheap tier when both hold. The result can be verified without being redone, by a test, a diff, a lint rule, or ten seconds of reading. And a wrong result costs a reviewer a minute, not an afternoon, because the blast radius is one label, one comment, or one file.

That pair is why test scaffolding qualifies and a security review does not, though both are text generation over the same repository. A generated test either runs or it does not, and you find out in the same minute. A missed authorization hole in a review comment looks exactly like a clean one.

What to hand over

Job What the Agent returns How it is checked Blast radius if wrong
Triage of incoming bug reports Summary, suspected component, reproduction attempted, priority suggestion Four lines read when picking up the Task One misrouted Task
Labelling and metadata Labels and Task type, naming the rule applied Visible on the board One wrong label
Changelog and release notes Draft notes from a Sprint’s merged diffs Diff against the commits Editorial, caught before release
First-pass review comments Observations as comments, explicitly not an approval The reviewer reads or ignores them Noise, if nothing depends on it
Test scaffolding Test files that compile and run, with obvious gaps marked The test run Red build
Dependency bumps One bump per Task, with changelog excerpt and test result CI One revert
Documentation follow-ups after a merge Doc edits citing the source diff Diff review A stale page stays stale

The unifying shape: each returns an artifact somebody else evaluates in seconds, on evidence the Agent attached. None is the last word on anything, which is the test to apply when you extend the list to whatever your crew picks up next.

What not to hand over

Three boundary lines are worth stating, because a cheap tier invites these mistakes.

First-pass review is not review. An Agent posting observations on a diff generates reading material. It is not an approval, and the moment a team treats a clean first pass as a signal that a change is safe, the cheap tier has been promoted into the acceptance decision. Be precise about what a review agent can and cannot verify, and label the comment as a first pass in the comment itself.

A label the board automates on is a decision, not a note. If moving a Task to a status triggers a run, or priority drives what the crew picks up next, a labelling Agent is setting direction. Either keep the automation off those fields or keep labelling on the tier you trust with direction.

Work nobody will check does not get cheaper, it gets invisible. A cheap wrong answer is cheap only when somebody notices it. If a job has no reader, the fix is deleting the job, not a cheaper runtime.

The decision pair, then. Use Luna when work is high volume, bounded, and verified by something other than trust. Use Sol or Opus 5.5 when a Task must hold context across hours, crosses several modules, or costs more to stop halfway than the tokens would. If you run one Agent on one repository, do not split tiers at all: pick the runtime that fits your hardest work.

Writing a Task a cheap runtime can finish

An Agent in Sharkly is a saved working configuration rather than a one-off prompt. Its instructions, Runtime, Skills, repositories, environment, and run settings shape how it handles work, so the right unit is one Agent per boring job type, not one clever Agent doing all seven. A team can keep separate Agents for code changes, review, testing, documentation, or operational triage without rewriting instructions for every Task.

A cheap runtime needs a narrow scope and a mechanical finish line:

SH-312  Triage: 9 unassigned bug reports from this week
Assignee: Triage Agent (Codex Runtime, gpt-6-luna)

For each report, return on the Task:
  - one-sentence restatement of the observed behavior
  - the suspected component, with the file or module that led you there
  - whether you could reproduce it locally, and the exact command you ran
  - a priority suggestion with one reason

Do not change files. Do not close or merge anything.
Finish with any question that must be answered before implementation.

Three things in that Task do the work. The scope is one pass over nine items, so a failure is cheap to discard. Every claim carries its evidence, so checking is reading rather than re-investigating. And the last two lines set the boundary: this Agent reports, it does not act.

One setting is worth deciding on purpose. Luna’s published wins sit at high and max reasoning effort, and effort is a cost and latency decision rather than a quality dial you leave at the top. Artificial Analysis, again a third party, measured the max reasoning variant at 153.9 output tokens per second for Luna with about 124 seconds to first token. [VERIFY] Whatever the current figure, the shape holds: at maximum effort the cheap models are not the fast-to-start ones. A labelling job that thinks for two minutes looks hung on a board. Run short jobs at modest effort and save the high settings for Tasks that earn them.

Two notes on access, since this tier is where teams onboard. Luna runs in Codex and ChatGPT Work for Plus, Pro, Business, Enterprise, and Edu plans, and Free and Go accounts get it in the desktop app. It is not in Chat yet. In Sharkly the model is set in the Runtime or Computer tool configuration, never on the Agent, and Sharkly will not display the model name back to you.

The part a price cut does not fix

Nine triage summaries at Luna prices cost close to nothing. They still produce nine things a person must read. A crew that moves its whole surround onto a cheap tier has not reduced the decisions waiting on a human, only made them cheaper to generate, which is not the same as cheaper to accept. That is why the review queue stays the ceiling no matter what the token rate does.

What you can change is how legible that queue is. Keep the goal, the run, the evidence, the blockers, and the human review on the Task from request to release, so a first-pass comment is visibly a first pass, a triage summary carries the command that made it, and the items actually waiting on you are separable from the ones that are merely noisy. Agents research, execute, test, and report. People set direction, grant authority, and accept the result.

If you run several agents and the cheap ones are about to multiply, Sharkly gives you one place to assign their Tasks, see what each run produced, and decide what ships. Connect one Computer, create an Agent for your most boring job, and assign it one real Task.

Explore more

What It Actually Costs to Run an Agent Crew on Free and Cheap Tiers

What It Actually Costs to Run an Agent Crew on Free and Cheap Tiers

Which models a Crew can run for free after the 2026-09-22 launches, why a free desktop entitlement is not an API key, and what the cheap API path really costs per week.

23 September 2026

When a Cheaper Model Costs More: Price Agent Runs Per Accepted Change

When a Cheaper Model Costs More: Price Agent Runs Per Accepted Change

GPT-6 Luna at $0.10 per million tokens lowers the price of an attempt and raises attempts, output and decisions per shipped change. Price agent runs per accepted change, not per token.

23 September 2026

GPT-6 Sol at $0.27 a Task: What Your Fixed Agent Budget Now Buys

GPT-6 Sol at $0.27 a Task: What Your Fixed Agent Budget Now Buys

OpenAI put GPT-6 Sol at $0.27 per AutomationBench task against Claude Opus 5 at 11.1x that. What a fixed monthly agent budget now buys, and why the new ceiling is your reviewers.

23 September 2026