OpenAI Says It Has Reached the Automated Research Intern Milestone
OpenAI reports 3.1 agent-workdays for every human workday, but multi-hour success still needs frequent intervention. The build decision is to invest in orchestration, not wait for full autonomy.
What changed
On September 6, OpenAI said its internal systems had reached its definition of an automated research intern: completing well-defined, human-directed research tasks that would take a skilled researcher several days.
By mid-August, the organization was running 3.1 agent-workdays for every human workday. That is a measure of execution time, not validated productivity or savings.
The hidden bill
The same disclosure punctures the unattended-autonomy story. More than half of successful tasks estimated at four to eight hours still required at least one human intervention, while high-level planning remained largely human.
With median usage above $600 per researcher per day at API prices, token price is an incomplete metric. Teams need cost per accepted result, including review, retries and rework.
What we would ship
Virtual Arc would begin with bounded queues such as bug reproduction, patch preparation, log analysis and verifiable experiments. Humans would retain prioritization, acceptance criteria and final approval.
We would impose per-task budgets, checkpoints before risky tool calls, complete execution traces and model portability. We would not scale until the total economics of a completed task beat the existing alternative.
OpenAI’s milestone matters, but not because teams can now replace researchers. The more useful signal is that a frontier lab has normalized running several times more agent-hours than human-hours, with the median researcher consuming more than $600 a day of inference at API list prices. Yet more than half of successful four-to-eight-hour tasks still required human intervention, and high-level planning remained a minimal share of agent output. Our view is blunt: this is not autonomous labor; it is a new operating model for supervised, parallel execution. Virtual Arc would not wait for a public automated researcher, and we would not rebuild a product around OpenAI’s internal claim. We would put two to four agents on bounded engineering and analysis queues now, add hard spend limits, preserved traces, explicit checkpoints and a model-switch layer, then judge them by cost per accepted result: a merged change, closed incident, reproduced finding or validated experiment. If that metric beats the human-only baseline after review and rework, scale it. If it does not, more agent runtime is just a larger bill.