9 September 2026

Mercury 2.5 brings ultra-fast, low-cost calls to AI agent workloads

Inception’s diffusion model promises 1,107 tokens per second at a low price, but its strongest first role is as an agent sidecar rather than a frontier-model replacement.

Virtual Arc · Editorial image

What launched

Inception released Mercury 2.5 on September 8, with a 260K-token context window, tunable reasoning, parallel tool calls and schema-aligned JSON. The company reports throughput of 1,107 tokens per second on widely available NVIDIA GPUs.

Where economics change

Standard pricing is $0.20 per million input tokens and $0.75 per million output tokens, with an 80% launch discount. That combination matters most when one user request triggers many short, sequential model calls.

What we would do

Virtual Arc would first route bounded, easily checked work to Mercury 2.5 while retaining a frontier model for planning, risk judgment and final responses. We would expand only after workload-specific tests of quality, reliability and cost per completed task.

Our take

Mercury 2.5 matters less as another chatbot and more as evidence that the supporting calls inside an agent no longer need to inherit the cost and latency of its main model. Virtual Arc would not move final-answer generation, difficult planning or irreversible actions onto it immediately: Inception positions its quality alongside cost-optimized rather than top-tier frontier models, and OpenRouter still labels the endpoint Preview. We would, however, test it now as an agent sidecar for routing, query rewriting, reranking, context compression and schema-bound JSON generation. We would measure tool-call reliability, schema failures, p95 and p99 latency, and cost per successfully completed task, while budgeting against the standard $0.20/M input and $0.75/M output rates rather than the launch discount. If it passes those evaluations, Mercury 2.5 should not replace the expensive planner; it should remove repetitive work around it. That is a smaller migration, contains the risk and is more likely to produce an immediate production gain than another wholesale model swap.

Sources
  1. Introducing Mercury 2.5
  2. Inception models on OpenRouter
  3. Mercury 2.5 API, pricing and playground on Vercel AI Gateway
  4. Mercury 2.5 pricing and benchmark evidence

← All posts