Mercury 2.5 brings ultra-fast, low-cost calls to AI agent workloads
Inception’s diffusion model promises 1,107 tokens per second at a low price, but its strongest first role is as an agent sidecar rather than a frontier-model replacement.
What launched
Inception released Mercury 2.5 on September 8, with a 260K-token context window, tunable reasoning, parallel tool calls and schema-aligned JSON. The company reports throughput of 1,107 tokens per second on widely available NVIDIA GPUs.
Where economics change
Standard pricing is $0.20 per million input tokens and $0.75 per million output tokens, with an 80% launch discount. That combination matters most when one user request triggers many short, sequential model calls.
What we would do
Virtual Arc would first route bounded, easily checked work to Mercury 2.5 while retaining a frontier model for planning, risk judgment and final responses. We would expand only after workload-specific tests of quality, reliability and cost per completed task.
Mercury 2.5 matters less as another chatbot and more as evidence that the supporting calls inside an agent no longer need to inherit the cost and latency of its main model. Virtual Arc would not move final-answer generation, difficult planning or irreversible actions onto it immediately: Inception positions its quality alongside cost-optimized rather than top-tier frontier models, and OpenRouter still labels the endpoint Preview. We would, however, test it now as an agent sidecar for routing, query rewriting, reranking, context compression and schema-bound JSON generation. We would measure tool-call reliability, schema failures, p95 and p99 latency, and cost per successfully completed task, while budgeting against the standard $0.20/M input and $0.75/M output rates rather than the launch discount. If it passes those evaluations, Mercury 2.5 should not replace the expensive planner; it should remove repetitive work around it. That is a smaller migration, contains the risk and is more likely to produce an immediate production gain than another wholesale model swap.