You dispatch three runs before lunch. Ninety seconds later, two of the three terminals have printed nothing at all. One of them is reasoning through the repository. One of them hit a 504 on the way out and is never coming back. From where you are sitting, they are the same rectangle of scrollback with a cursor in it.
This is not a new problem, but the launches of 2026-09-22 made it worse in a specific, measurable way. The highest reasoning settings on the new models buy accuracy by thinking for a long time before they say anything. Artificial Analysis clocked the “max” reasoning variant of GPT-6 Sol at 102.15 seconds to first token, and GPT-6 Luna at 124.23 seconds. Those are third-party measurements on a specific configuration, not vendor numbers, and that caveat matters, but the direction is not in doubt: the silence at the start of a run is now long enough to be indistinguishable from a failure.
TL;DR
Time to first token is the delay between sending a request and the first character of the reply arriving, and on a high-effort reasoning model that delay is where most of the thinking happens. A terminal cannot tell you the difference between a run that is thinking, a run that is waiting for a human, and a run that died two minutes ago, because all three print nothing. A status can. Agents research, execute, test, and report. People set direction, grant authority, and accept the result. What that contract needs to work is a run state that is visible without a person watching for it: Sharkly keeps context, progress, blockers, results, and human review on the Task from request to release, so a long thought and a dead run stop looking alike.
The wait got longer, and the cheap model is not the fast one
Here is what the third-party measurements say about the two new GPT-6 models at their maximum reasoning setting.
| Model | API id | Output speed | Time to first token | Input / output per 1M tokens |
|---|---|---|---|---|
| GPT-6 Sol (max) | gpt-6-sol |
115.2 tokens/s | 102.15s | $2 / $10 |
| GPT-6 Luna (max) | gpt-6-luna |
153.9 tokens/s | 124.23s | $0.10 / $0.50 |
Read the last two columns together. Luna costs a twentieth of Sol per token and makes you wait longer before it speaks. Cost and latency are separate axes, and the intuition that a cheaper model is a snappier model is simply wrong here.
Two honest qualifications before anyone builds a dashboard on those numbers. First, they come from Artificial Analysis rather than from OpenAI, and they describe the “max” reasoning variants specifically. The lower effort settings are a different animal, and your own numbers will move with prompt size, tool availability, and where your traffic lands. Second, time to first token is not the same measurement as output speed. Anthropic says Claude Opus 5.5 produces output about 30% faster than Opus 5, which is a statement about the tokens after the first one. A model can stream quickly and still leave you staring at nothing for a minute and a half.
A quiet terminal is not a status
Silence in a terminal has at least four causes, and the terminal renders all of them identically:
- The Runtime is working. Reasoning is happening, no output yet.
- The run finished a while ago and is waiting for a human to answer a question.
- The request failed. The Sharkly docs are blunt about how easy this one is to misread: “Timeout / 408 / 504 means the request was slow, canceled, or delayed by a gateway. That is not the same as the Computer being offline.”
- The Computer really is gone, or the directory the run needed was never free.
One agent, one terminal, and you resolve this by waiting another minute. Five agents across three repositories and two Computers, and you resolve it by tabbing between windows and guessing. That is the ordinary shape of agent chaos: not a dramatic failure, just a team spending its attention on the question “is this one alive?” several times an hour.
The reflex people reach for is to kill it and re-run. At high reasoning effort that reflex is expensive twice over. You pay the whole first-token wait again, and you throw away the thinking that was about to arrive.
What a status has to separate
Sharkly’s own model is worth copying whether or not you use Sharkly, because it separates two things a terminal fuses. It is the same split that makes run status, logs, and failure signals legible at all. A run has an execution state: Queued, Dispatched, Waiting for local directory, Running, Completed, Failed, or Canceled. A Task has a lifecycle status, and an Agent has a working state on top of it: Working, Waiting for human reply, Waiting for human review, or Error. The docs state the boundary plainly: these run states “do not automatically imply that the Task has the same workflow status.”
That split is the whole answer to the stuck-or-thinking question.
| What the terminal shows | What it may actually be | What tells them apart | Who acts next |
|---|---|---|---|
| Nothing, 100s in | Running, pre-first-token | Run state is Running, log has a dispatch timestamp | Nobody. Leave it. |
| Nothing, 10 min in | Waiting for human reply | Working state changed, Task raises an attention item in Inbox | The Assignee answers in context |
| Nothing, forever | Failed | Run state is Failed, execution log carries the title and the original model-service response | The Assignee reads the log before retrying |
| Nothing, never started | Queued or Waiting for local directory | Run state, plus which directory pool is busy | Free a directory or raise parallel capacity |
None of those four rows requires a person to be watching at the moment it changes. That is the point. The state is recorded, and the question “is it thinking or is it stuck?” becomes a thing you look up rather than a thing you sit through.
The per-Task timeout you set in March
An Agent in Sharkly carries a maximum number of parallel running Tasks and a per-Task timeout. Most teams set that timeout once, early, against how long they were willing to stare at a terminal. If that number is two minutes and you have since moved an Agent to a max reasoning setting, you have configured a policy that cancels runs during the thinking phase and reports them as failures.

Go and look at yours this week. The floor is not “how long am I willing to wait”, it is “how long does this Runtime take at the effort level I actually use, plus the work itself”. For a run whose first token alone can take over a hundred seconds, a timeout under five minutes is a coin flip dressed as a policy.
Two related settings worth the same pass. Maximum parallel running Tasks is your real concurrency limit, and when every configured directory is busy, new runs sit in Waiting for local directory rather than failing, which is correct behavior that looks like a hang if you are only watching a terminal.
Does caching fix it?
Partly, and not the part you want. OpenAI’s prompt caching for GPT-6 gives a 90% discount on cached input reads with higher hit rates by default, adds explicit breakpoints so you choose where a cached prefix ends, and no longer breaks the cache when you change reasoning effort or tool availability. GitHub reports more than 50% fewer prompt tokens needing fresh processing across billions of requests. That is a real cost and latency win on the prompt, and the effort-change fix is genuinely useful for agents that escalate a task from a cheap setting to an expensive one mid-flow.
It does not remove the reasoning pass. A warm cache means less of your prompt is processed from scratch. It does not mean the model has already thought about your bug. Budget for the wait anyway.
When to pay the wait, and when not to
Use a maximum reasoning setting when the task is one you do not want to run twice: a migration, a cross-cutting refactor, anything where a wrong answer costs more than ninety seconds. OpenAI reports Sol at its xhigh setting scoring 33.2% on AutomationBench 1.0.6 at $0.27 per task, which is the shape of the trade. Accuracy is cheap in money now and expensive in wall-clock.
Use a lower effort setting when a person is sitting there waiting on the answer. Interactive work, a quick question about a file, a small scoped fix: the silence problem mostly disappears at lower effort, because the first token shows up while you still remember what you asked.
And the honest boundary: if you are running one agent on one Task, you do not need a board for this. Tail the log. The state problem starts at the point where the number of runs exceeds the number of terminals you can meaningfully watch, which for most people is about three.
What this does not fix
A status tells you a run is Working. It does not tell you the run is doing something useful, and a confidently wrong agent shows the same green state as a good one for as long as the run lasts, which on Opus 5.5 can be 18 hours or more. Distinguishing “running” from “running well” is still a review problem, and review is still a person’s job. Automation stops where team judgment is required.
If you are already running several agents at once, the smallest useful change this week is not a new model. It is making every run’s state readable without opening a terminal, so the next ninety seconds of silence is information instead of a question.



