30 September 2026

GPT-6.1 Sol resets the default model for production agents

OpenAI claims near-Astra capability at one-fifth of its standard token price. The right move is to reroute evaluations now, not rebuild the stack.

Virtual Arc · Editorial image

What changed

OpenAI released GPT-6.1 Sol on September 29 through its API, ChatGPT Work and Codex. Standard pricing is $2 per million input tokens, $10 per million output tokens and $0.10 per million cached input tokens.

OpenAI reports that Sol approaches Astra on selected coding, computer-use and professional-work evaluations at substantially lower task cost. These remain vendor-reported results, not evidence from a customer’s production workload.

Why it matters

The important change is not another benchmark lead. It is a new cost-capability curve: Sol’s standard input and output rates are one-fifth of Astra’s, while cached input is half the price of the earlier GPT-6 Sol.

That matters for agents repeatedly carrying tool definitions, repository context and operating instructions. Teams already using the Responses API can evaluate it without adopting a new orchestration stack.

What we would do

Virtual Arc would add Sol to the existing model route immediately and replay representative production tasks. We would measure end-to-end completion, latency, retries, tool calls, human corrections and the cache-hit rate—not tokens in isolation.

If reliability holds, Sol becomes the default and Astra becomes the escalation path. We would not accept a fivefold task-cost claim until our own completion data and invoices demonstrate it.

Our take

GPT-6.1 Sol should replace Astra as the first model we test for production agent work, not because OpenAI has proved parity, but because the price gap is now too large to justify flagship-first routing. Virtual Arc would put Sol behind the existing model gateway this week, replay real production traces, and compare cost per successfully completed task, p95 latency, retries, human intervention and safety failures. We would keep Astra only as an escalation route for the smaller set of tasks where any lower completion rate from Sol costs more than the token savings. We would not redesign around Sol-specific behavior or count on cached-input savings until production telemetry shows that our prompts actually hit the cache. The build decision is yes: test now, route selectively, and make the premium model earn every call rather than paying for maximum capability whether a workflow needs it or not.

Sources
  1. Introducing GPT-6.1 Sol
  2. GPT-6.1 Sol Model — OpenAI API
  3. The 5 biggest announcements from OpenAI's blockbuster AI conference
  4. OpenAI launches GPT-6.1 Sol at one-fifth Astra's token price
  5. AutomationBench AI benchmark leaderboard

← All posts