Mistral Large 4 public preview: test now, migrate later
Mistral's new multimodal model combines aggressive API pricing with a promise of open weights, but production teams should wait for evidence on real operating cost and reliability.
A credible challenger
On October 6, Mistral opened a public-preview API for Mistral Large 4, a natively multimodal mixture-of-experts model with 1.05 trillion total parameters and 49 billion active parameters per token. Mistral lists a one-million-token context window and plans to release the weights on October 31.
Launch pricing is $0.68 per million input tokens, $0.07 for cached input and $2.09 for output. That is a 50% discount lasting two weeks, not a durable basis for production budgeting.
Control is the product
The potential product is not another benchmark score. It is the option to run a capable model under your own policies, nearer your data and without making one hosted API the permanent control point for your system. That matters when refusals, policy changes or model withdrawals can become reliability incidents.
But a model of this scale will require serious infrastructure, and the published evidence is still largely preliminary or vendor-led. Open weights do not automatically mean a lower cost per completed task.
What we would do
Virtual Arc would place ML4 behind a portable routing layer immediately and run the same real workload against incumbent production models. We would measure the whole outcome: correctness, latency, retries, refusals, failed calls and the total cost of successful work.
We would keep core traffic where it is until the weights, license, infrastructure requirements and wider independent evaluations are available. This is the moment for a controlled challenger test, not a migration programme.
Mistral Large 4 matters because it makes an open-weight model a credible candidate for demanding production systems again, but Virtual Arc would not move core traffic to it yet. The API is available, the launch price is aggressive and the reported results are strong enough to justify immediate testing. The weights, however, are not due until October 31, independent evaluation remains thin, and the economics of self-hosting a model this large are still unknown. We would add ML4 behind a provider-neutral interface now, replay a representative production workload and measure cost per successfully completed task, latency, retries, refusals and structured-output reliability. We would not budget around a two-week discount or begin a migration before seeing the license, real memory and throughput requirements, and broader independent results. Our position is deliberately conservative: test now, preserve portability and make the migration decision only when open weights become an operating fact rather than a roadmap promise.