Gemini 3.8 Flash: Test task cost before upgrading
Google's new stable Flash model improves agentic performance without raising its introductory token price, but independent measurements show that high reasoning costs more per completed task.
What shipped
Google released Gemini 3.8 Flash on September 2 as a stable API model ready for production. It supports a one-million-token context window, up to 64,000 output tokens, built-in tools and low, medium or high reasoning effort.
Its introductory global price matches 3.7 at $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Both rates double on January 1, 2027.
The hidden meter
Google explicitly says 3.8 takes more reasoning steps, calls tools iteratively and verifies its work on complex jobs. Artificial Analysis measured a higher score than 3.7, but also 30% more output tokens at high reasoning.
That lifted estimated cost per task by about 40%, from $0.40 to $0.58, while average task time increased from 2.2 to 2.5 minutes. The unchanged token price therefore hides a different operating profile.
What we would do
We would not replace a production default globally. We would first route a representative sample of long-horizon coding and agent jobs to 3.8, using the same tools, limits and acceptance tests as the current system.
We would measure cost per successful task, p95 latency, tool calls, retries and human repair time. Version 3.8 earns default status only on routes where its extra reasoning lowers the total cost of completion.
Gemini 3.8 Flash is worth testing now, but we would not make it the default just because Google kept its introductory token price unchanged from 3.7. For an operator, tokens are an input; completed tasks are the product. Independent testing shows that high reasoning raises the Intelligence Index from 56 on 3.7 to 59, yet average output grew 30%, cost per task rose from $0.40 to $0.58, and task time moved from 2.2 to 2.5 minutes. That can still be a bargain if better first-pass completion eliminates retries, human review or failed tool loops, but a vendor benchmark cannot prove that for our workload. At Virtual Arc, we would put 3.8 behind a router, shadow it on long-horizon coding and agent jobs, default to medium or low effort, and retain 3.7 as an immediate fallback. We would promote it only where completion rate and repair time beat the extra spend, and we would budget against the January 1, 2027 standard price rather than the launch discount. The build decision is to test now, migrate selectively and refuse to confuse a flat token tariff with a flat production bill.