Why an Agent Platform Should Not Sell You Tokens

Sharkly does not sell tokens. Runs bill against the Claude Code, Codex, or API plan attached to the Runtime, not a resold token balance. What that separation costs you, and what it saves.

Mia Parker

Mia Parker

15 September 2026

Why an Agent Platform Should Not Sell You Tokens

Most agent platforms put themselves between you and the model provider, then charge a markup on every token that crosses. Sharkly does not. Model usage is bring-your-own-key: a run bills against the Claude Code, Codex, or API plan already attached to the Runtime that executes it, not against a token balance Sharkly sells you. That’s a smaller claim than it sounds, and it’s worth being honest about what it costs you as well as what it saves you before you decide it matters.

What a reseller model does to your incentives

Start with the platform that does sell you tokens, because that’s the default in this category and it’s worth naming plainly rather than dancing around it.

When a platform resells model access, its revenue is tied to how many tokens your Agents burn. That’s not a hidden flaw, it’s the business model working as designed, and it produces a predictable incentive: the platform has no reason to make a run cheaper. A more efficient prompt, a smaller context window, a Runtime that finishes in fewer tokens, all of that reduces the platform’s revenue on your account. Nothing forces a token reseller to pass along price cuts the underlying model vendor makes, and nothing rewards it for helping you spend less. You’re paying a markup for coordination, but the markup is denominated in the currency the platform wants you to consume more of.

The second cost is quieter and shows up later. If your team already has a Claude Code or Codex plan, or committed API spend with a provider, you likely negotiated pricing for that: a subscription tier, a volume discount, a committed-use rate. Route that usage through a reseller instead, and you’re paying twice for the same leverage: once to the reseller for the privilege of access, and you lose the benefit of the rate you already secured directly, because the reseller’s rate is the one that applies now. The pricing relationship you built with the model vendor doesn’t travel with you into a platform that resells tokens under its own account.

Put a number to it, roughly. A team running a handful of Agents through a reseller at even a modest per-token markup, on top of committed spend they already have with the model vendor directly, is paying for the same tokens twice over: once at the negotiated rate they can no longer use, once at whatever rate the reseller decided to charge. Scale that across a Sprint’s worth of runs, not a single Task, and the gap compounds in exactly the direction the reseller wants it to.

None of this makes token resale dishonest. It makes it a model where the platform’s interests and your cost interests point in different directions, and that’s worth knowing before you hand a vendor your runs.

What BYOK actually costs you

The separation isn’t free on Sharkly’s side either, and an article that only lists the advantages of bring-your-own-key would read like marketing rather than an honest account of the tradeoff. Say the costs plainly.

You manage the keys and the quotas yourself. Nobody at the coordination layer is watching your Claude Code plan’s rate limit or your API account’s remaining credit on your behalf. If your team runs five Agents against one Anthropic API key and that key hits a rate limit, that’s a constraint you’re responsible for noticing and planning around, the same way you would if a person on your team were making those calls directly from a terminal.

A run can fail for a reason that has nothing to do with Sharkly. If the Claude Code plan attached to a Runtime is rate-limited, or the API account backing it runs out of credit, the run stops, and the failure surfaces as a token or auth error from the model provider, not a Sharkly outage. Debugging that means checking two systems instead of one: is the Task stuck because of something in Sharkly, or because the underlying plan ran dry. That’s a real seam, and it’s the direct price of not having a platform-controlled token pool smoothing it over.

Spend visibility lives in two places. Sharkly’s usage dashboard shows consumption per Agent, but the bill itself shows up on your Claude Code, Codex, or API provider invoice, not on a single Sharkly line item. If your finance team wants one number for “what AI coding cost us this month,” they’re reconciling two sources instead of reading one summary, and that reconciliation is manual work someone has to actually do.

There’s an onboarding cost too, and it’s worth naming rather than assuming away. Every Runtime you connect needs its own key or plan configured, someone on the team has to own renewing and rotating those credentials, and a new hire who wants to run Agents needs access to the underlying provider account, not just an invite to Sharkly. A reseller platform can hide all of that behind a single sign-up. BYOK doesn’t, by design, and a team evaluating Sharkly should budget a real hour or two of setup per Runtime rather than expecting it to be zero.

None of that is a reason to avoid BYOK. It’s the honest list of what you take on in exchange for not paying a token markup, and a team should walk in knowing it rather than discovering it during an incident.

Where the management layer still earns its keep

Given that Sharkly isn’t selling you tokens or smoothing over provider rate limits, it’s fair to ask what the coordination layer is actually doing. Three things, concretely.

It decides which Computer a run executes on. An Agent’s Runtime needs a host, local or cloud, with the repositories and resources a Task requires, and that assignment is Sharkly’s job regardless of who’s paying for the tokens the run consumes. Token billing and execution placement are separate concerns, and BYOK only removes the platform from the first one.

It keeps a human in the approval path. Agents research, execute, test, and report; people set direction, grant authority, and accept the result. That review step, who signs off on a Task before it merges or closes, doesn’t depend on who’s paying for the model call underneath it. A cheaper reseller markup doesn’t buy you a better review process, and a BYOK model doesn’t cost you one either.

And it gives you a usage dashboard that shows consumption per Agent even though the tokens themselves bill elsewhere. You can see which Agent burned the most in a Sprint, which Task types are expensive to automate, and where a Runtime is running hot, all without that visibility requiring Sharkly to be the one charging you for the tokens it’s reporting on. The dashboard’s usefulness doesn’t depend on Sharkly owning the bill.

That’s the actual shape of the separation: Sharkly manages where work runs, who reviews it, and what it’s costing you to run, without needing to also be the vendor selling you the thing being measured.

Deciding whether that tradeoff is worth it

If your team is small, has one Runtime, and rarely bumps into a rate limit, the reseller-versus-BYOK question probably doesn’t change your day much either way. The tradeoff gets real once you’re running several Agents against the same underlying plan, coordinating Crews across more than one Computer, or trying to answer “what did AI coding actually cost us this Sprint” for more than one team. At that point, a platform with a hidden incentive to keep your token spend high is a cost you can’t fully see, and a platform that just tells you what you spent, without also being the one selling it to you, is worth the two-dashboards inconvenience.

Sharkly’s bet is that the coordination problem, not the metering problem, is the one worth charging for. It’s the layer above the agents you already run, not a second vendor between you and the tokens they spend. If you’re already paying for Claude Code or Codex directly and want the runs assignable, reviewable, and visible without a markup sitting on top of the tokens, that’s the layer Sharkly is built to be. See the current pricing for exactly what the seat side of that looks like.

Explore more

What It Actually Costs to Run an Agent Crew on Free and Cheap Tiers

What It Actually Costs to Run an Agent Crew on Free and Cheap Tiers

Which models a Crew can run for free after the 2026-09-22 launches, why a free desktop entitlement is not an API key, and what the cheap API path really costs per week.

23 September 2026

When a Cheaper Model Costs More: Price Agent Runs Per Accepted Change

When a Cheaper Model Costs More: Price Agent Runs Per Accepted Change

GPT-6 Luna at $0.10 per million tokens lowers the price of an attempt and raises attempts, output and decisions per shipped change. Price agent runs per accepted change, not per token.

23 September 2026

GPT-6 Luna for the Boring 80 Percent: What to Hand a Cheap Runtime

GPT-6 Luna for the Boring 80 Percent: What to Hand a Cheap Runtime

GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output. Which triage, labelling, first-pass review and test scaffolding belongs on a cheap runtime, and which does not.

23 September 2026