OpenAI published a number this week that deserves more attention than the dollar figure next to it. As of mid-August, its research organization uses 3.1 agent-workdays of effort for every workday of human labor. The line that got quoted was the spend: the median researcher runs more than $600 of inference a day at API prices. The line that matters for anyone running a team is the other one.
Three agent days per human day means the human day is now the constraint. Every result an agent produces has to be read, judged, and accepted by someone, and that someone has one workday, not three. This post does the arithmetic and shows where the queue forms.
TL;DR
OpenAI’s “Research acceleration” post puts agent effort at 3.1 workdays per human workday, notes a trend toward researchers running four or more agents at once, and reports that over half of successful 4-8 hour tasks needed at least one human intervention. Put those together and a working day contains several moments where an agent has stopped and is waiting for a person, plus a steady stream of finished work waiting to be accepted. The queue that fills is “results waiting for a person,” and most tooling does not show it. Sharkly makes that queue explicit: waiting states on the Task, an Inbox that routes them to the responsible person, and acceptance as a recorded status change.
The numbers
Everything below is from OpenAI’s post unless marked otherwise.
| Figure | Value | Period |
|---|---|---|
| Agent-workdays per human workday, research org | 3.1 | Mid-August 2026 |
| Median researcher daily inference at API prices | More than $600 | Mid-August 2026 |
| 90th percentile researcher daily token spend | More than $7,000 | Current |
| Successful 4-8 hour tasks with 1+ human interventions | Over half | January to July 2026 |
| Concurrency trend | 4 or more agents simultaneously | Increasing |
Simon Willison’s read of the same post adds the shape of the curve: per-researcher spend went from roughly zero in February to about $600 by late August, with a steep climb between July and August that he guesses coincides with internal access to the model later released as GPT-6 Astra.
None of these figures are about a coding team shipping product, and OpenAI’s researchers are not a typical population. But the ratio travels. Any team that has moved from one agent in one terminal to several agents in flight is on the same curve, a few months behind.
The arithmetic
Take the ratio at face value. In an eight-hour human day, agents produce roughly twenty-five hours of effort. Not all of it needs a person. Some runs fail, some are exploratory, some feed other agents. But every run that ends in a change someone will ship needs a human decision, and the intervention figure says the decision often arrives before the run is over.
If more than half of long successful tasks needed at least one intervention, then with four agents running, a person should expect to be interrupted by two or three of them over the course of a task, at times the person did not choose. Those interruptions are not failures. An agent that stops and asks beats one that guesses and ships. The cost is in where the question lands.
A question that lands in a terminal is answered when someone happens to look at that terminal. A result that lands in a terminal is reviewed when someone remembers it exists. Multiply by four agents, and by three agent-days per human day, and the terminal is where work goes to wait.
The queue nobody dashboards
Every team running agents has two queues. The first is work waiting for an agent: the backlog, visible in whatever tracker the team uses. The second is work waiting for a person: results to accept, questions to answer, failures to triage. The second queue is the one that grows at 3.1 to 1, and it is the one most tooling does not show.
Sharkly treats that queue as the primary object. A Task is the main unit of work, and Task status and Agent working state are kept separate: a Started Task can show Working, Waiting for human reply, Waiting for human review, or Error. The waiting states are first-class, not inferred from a quiet log.
Those states route somewhere. The Inbox is a personal queue of task activity, and its Primary section holds tasks with an active signal that needs attention: waiting for your reply, waiting for your review, failed or blocked work, direct mentions, assignment to you. It groups activity by task, so four agents reporting on one Task are one row, not four notifications. Slack is an optional outbound channel; muting a group there does not remove the Inbox record.
Replying is how work continues. A comment on a Task with a ready assigned Agent can start a follow-up run with the thread as context, so answering the question resumes the work rather than restarting it.
What acceptance looks like at this ratio
At one agent per person, review can be informal. At three agent-days per human day it cannot, because the person reviewing will not be the person who started the run, and the run may have ended hours ago.
Acceptance in Sharkly is a person changing the Task status after reading the evidence. The Agent’s final response is stored on the Task. The execution log records why the run was queued, when it started and ended, tool calls, Runtime events, and failure details, in an Executions tab separate from the comment timeline. Read the summary, then the log, then the diff. Merging a linked pull request does not change the Task status, so a merge is never mistaken for a decision. Our human-in-the-loop review workflow walks through the five gates in order.
The Hacker News discussion of OpenAI’s post fixed on a related worry: results scored by an agentic classifier, with the thing measured and the thing grading it coming from the same house. That is the research version of a product team’s problem. When agents produce more than people can read, the temptation is to let an agent decide what people should read. A visible queue with a human owner is the alternative.
Where to spend the human day
Use the human day for direction and acceptance: shaping the ticket, granting authority, reading the evidence, deciding merge and release. Leave execution, testing, and reporting to the agents. Agents research, execute, test, and report. People set direction, grant authority, and accept the result. The ratio in OpenAI’s post is what that contract looks like when it works.
Use the human day for supervision only when the work genuinely needs a person watching in real time: a production change with no rollback, a run in a directory that other work depends on. Sharkly is not a replacement for Claude Code, Codex, or the Runtime doing the work. It is the layer that keeps context, progress, blockers, results, and human review visible from request to release, so the human day is spent on decisions and not on finding out which terminal has a question in it.
Frequently asked questions
Where does the 3.1 figure come from? OpenAI’s post “Research acceleration: the view inside OpenAI,” which states that as of mid-August the research organization uses 3.1 agent-workdays of effort for every workday of human labor.
Does this apply to product engineering teams? The specific numbers are from a research org. The ratio’s direction applies to any team moving from one agent to several: agent output grows faster than the capacity to review it.
What is the intervention figure? Over half of successful 4-8 hour tasks between January and July involved one or more human interventions. Long runs succeed with a person stepping in, not without one.
How does Sharkly show the review queue? Waiting for human reply and waiting for human review are Agent working states on the Task, routed to the Inbox Primary section of the responsible person. Failed and blocked work appears there too.
What counts as acceptance? A person changing the Task status after reading the Agent’s response, the execution log, and the diff. A merged pull request does not change the status on its own.
The short version
OpenAI’s ratio says agents already produce three days of effort for every human day, and its intervention figure says a person is in the loop for most long tasks. The scarce resource is the human day, and the queue that fills is results waiting for a person. Make that queue visible, route it to an owner, and record the decision. Connect one Computer, create one Agent, and assign it something boring. The docs cover each concept in more depth: docs.sharkly.ai.



