What It Actually Costs to Run an Agent Crew on Free and Cheap Tiers

Which models a Crew can run for free after the 2026-09-22 launches, why a free desktop entitlement is not an API key, and what the cheap API path really costs per week.

Ashley Innocent

Ashley Innocent

24 September 2026

What It Actually Costs to Run an Agent Crew on Free and Cheap Tiers

The answer splits in two, because the bill does. The platform line can genuinely be zero: Sharkly is free for organizations of up to 10 people and does not sell model tokens. The model line cannot be zero for a Crew that keeps working while you do something else. Of the three models the labs shipped on 2026-09-22, exactly one has a free route, that route is GPT-6 Luna for Free and Go users in the ChatGPT desktop app, and a desktop app is a seat for a person rather than a credential a run can use. Free to coordinate, cents to execute, and nothing about either is free to review.

A Crew’s cost has two lines that people merge into one: the seat line for the platform where work is assigned, tracked and accepted, and the token line for the Runtime that executes. They move independently. The seat line is a published price you can read today. The token line is variable, it belongs to your own provider accounts, and on 2026-09-22 it fell hard.

TL;DR

Sharkly’s seat line is $0 up to 10 people and $7 per user per month billed yearly above that, with model usage bring-your-own-key. On the model line, GPT-6 Luna is free for Free and Go users in the ChatGPT desktop app, GPT-6 Sol is not free, and Claude Opus 5.5 is not available on Claude Free at all. A desktop entitlement cannot start an unattended run. The genuinely cheap path is the API: Luna at $0.10 and $0.50 per million tokens, Sol at $2 and $10, cached input reads 90% off. None of that touches the one line that never had a free tier, which is the person who accepts the work.

The seat line: what free means on the platform

Sharkly’s pricing has one variable, the number of people. Free is $0 for organizations of up to 10 people. Team is $7 per user per month billed yearly for organizations larger than that. Self-hosted is priced case by case. Free and Team are the same product, so the plan does not gate features.

The part that matters for a Crew: guests, Agents, Crews, Computers and Runtimes do not occupy seats. A seat is a person in the organization. Adding a sixth Agent, a second Computer or a third Runtime does not change the seat line. An eleventh colleague does.

Two honest caveats. Sharkly is free to use today and the Team figure is the published seat price, so currently free is the accurate phrasing, not free forever. And Sharkly does not sell tokens: a run bills against the Claude Code, Codex or API plan already attached to the Runtime that executes it. That is why the two lines here stay separate, and why none of the model economics below depends on which platform you route through. Plan mechanics: Sharkly pricing explained.

The token line: which new models you can reach for nothing

Model Free route What that route is not Paid entry
GPT-6 Luna Free and Go users, in the ChatGPT desktop app Not in Chat yet, and not an API key API gpt-6-luna, $0.10 / $0.50 per 1M
GPT-6 Sol None Plus, Pro, Business, Enterprise and Edu only, in ChatGPT Work and Codex API gpt-6-sol, $2 / $10 per 1M
Claude Opus 5.5 None. Claude Free carries Sonnet and Haiku Claude Code is included in paid plans, not Free Pro at $17/month on annual ($200 up front) or $20 monthly; Max from $100/month; API claude-opus-5-5, $4 / $20 per 1M

Two of the three answers are no, and the third has a condition attached that decides everything.

A desktop entitlement is not an API key

This is the distinction most free-tier roundups skip, and it is the one that decides whether a free model is useful to a Crew.

Sharkly separates the roles: the Agent defines how work should be handled, the Computer supplies the host, the Runtime performs the actual session, and the Task stays the shared record. A run begins when a Task with a ready Agent reaches a Computer with an available Runtime, and that Runtime authenticates using the subscription or API key configured in the tool on that Computer.

A free ChatGPT desktop entitlement is not one of those credentials. It is a person signed into an app on their own laptop, typing. Nothing in it attaches to a Runtime, and nothing in it starts while its owner is asleep. A free model in a desktop app gives you a model to evaluate. It does not give a Crew something to run on.

The free route is still worth using, so here is the decision pair.

Use the free desktop entitlement when a person is present and the question is judgement: how a model reasons about your codebase, whether an approach holds up before you commit a Sprint to it, whether the model belongs in your routing rule at all. That is a real afternoon of evaluation at no cost.

Use a paid subscription or an API key when the run has to start without you. Queued Tasks, overnight assignments, anything a teammate triggers while you are in a meeting. Runs that happen while you are away need a credential, not an entitlement.

One planning note: OpenAI states these models are not yet available in Chat. A free route that lives on one surface today can move next month, which is an argument for writing the routing rule down rather than shaping a workflow around whichever tier happened to be free this quarter.

The cheap path that does power a Crew

Runtime model Input / output per 1M Cache lever Other levers
gpt-6-luna $0.10 / $0.50 90% off cached input reads Higher default hit rates, caching dashboard, diagnostics tool
gpt-6-sol $2 / $10 90% off cached input reads Explicit cache breakpoints
claude-opus-5-5 $4 / $20 Cached reads $0.20, cache writes $5 Batch API 50% off; fast mode is $8 / $40, so avoid it

OpenAI describes Sol and Luna as 50% cheaper than GPT-5.6 promotional pricing. The qualifier is OpenAI’s own word and it matters: the comparison runs against a promotional rate, not a list rate.

Caching is the larger lever anyway. GPT-6 raises default hit rates, discounts cached input reads by 90%, ships a caching dashboard and a diagnostics tool, and no longer invalidates the cache when you change reasoning effort or tool availability. GitHub reports more than 50% fewer prompt tokens needing fresh processing across billions of requests. Opus 5.5 reads cached input at $0.20 against $4 standard input. For a Crew that re-reads the same repository all day, the bill is decided there, not on the headline rate.

What a week actually costs

OpenAI published one figure that converts into a budget: Sol at extra high reasoning scores 33.2% on AutomationBench 1.0.6 at $0.27 per task. By our own arithmetic, a Crew completing 40 such tasks in a week spends about $11 on the model line. Your repository is not AutomationBench and your number will differ, but at a small team’s volume the model line is now a rounding error against one engineer hour.

The comparisons behind it: Sol at extra high beats Claude Opus 5 at max reasoning for 9% of Opus 5’s cost per task, and Luna at max scores 66.6% on DeepSWE 1.1 at 93% less per task than Opus 5. One correction applies to all of it. OpenAI’s baseline throughout is Opus 5, because Anthropic shipped Opus 5.5 at $4 and $20 hours after that post went up, 20% under Opus 5’s $5 and $25 and 40% less to run. Neither lab benchmarked the other’s current model, so treat both tables as directional and pick on what each Runtime is for, not on a leaderboard.

For the other end of the range: OpenAI reports its own researchers’ median coding agent spend above $600 a day, with a 90th percentile of $7,000. Nothing prevents a Crew from spending like that except a written routing rule and an agreed cap, which is a management decision rather than a tier you shop for. If spend is already spread across several tools, cost tracking across agents comes before any price comparison.

Where the free plan stops helping

Add the two lines up at small scale and the total is genuinely small: $0 of seats up to 10 people, tens of dollars a week of tokens on the cheap tiers. Neither line contains the expensive step.

Agents research, execute, test, and report. People set direction, grant authority, and accept the result. Accepting has no free tier, no cheap tier and no price cut, and it scales with people rather than with dollars. Human review is the constraint whatever the Runtime charges. Free for 10 people and $0.10 per million input tokens make it cheap to start ten runs this afternoon; they make it no cheaper to read ten results this evening. What the management layer changes is not the price of a run but where it lands: execution, blockers, results and follow-up discussion return to the Task, so context, progress, blockers, results and human review stay visible from request to release, and the work waiting on a person is a number instead of a feeling you get on Friday.

The line that was never free

Free and cheap are two different questions now. The platform question has a published answer: $0 up to 10 people, $7 per user per month billed yearly after that, keys and subscriptions your own. The model question has a smaller answer than the headlines suggest, because one free route runs through a desktop app no unattended run can use and the other two do not exist.

The useful conclusion is not which tier to pick. It is that the model bill stopped being what limits how much your Crew ships, and nothing either lab shipped on 2026-09-22 reads a diff for you.

If you are already running several coding agents, Sharkly gives you one place to see which Runtime produced which result, what evidence it attached, and how much of it is sitting in front of a person right now. Connect one Computer, create one Agent, and assign it something boring. At ten people or fewer, that part costs nothing.

Explore more

When a Cheaper Model Costs More: Price Agent Runs Per Accepted Change

When a Cheaper Model Costs More: Price Agent Runs Per Accepted Change

GPT-6 Luna at $0.10 per million tokens lowers the price of an attempt and raises attempts, output and decisions per shipped change. Price agent runs per accepted change, not per token.

23 September 2026

GPT-6 Luna for the Boring 80 Percent: What to Hand a Cheap Runtime

GPT-6 Luna for the Boring 80 Percent: What to Hand a Cheap Runtime

GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output. Which triage, labelling, first-pass review and test scaffolding belongs on a cheap runtime, and which does not.

23 September 2026

GPT-6 Sol at $0.27 a Task: What Your Fixed Agent Budget Now Buys

GPT-6 Sol at $0.27 a Task: What Your Fixed Agent Budget Now Buys

OpenAI put GPT-6 Sol at $0.27 per AutomationBench task against Claude Opus 5 at 11.1x that. What a fixed monthly agent budget now buys, and why the new ceiling is your reviewers.

23 September 2026