3 October 2026

Microsoft MAI-Transcribe-2-Streaming: benchmark leader, not production-ready

Microsoft’s new model leads a public streaming-transcription test, but its preview status and missing SLA make this a pilot decision, not a migration decision.

Virtual Arc · Editorial image

What launched

On October 1, Microsoft released MAI-Transcribe-2-Streaming for real-time transcription in 60 languages with automatic, continuous language detection. It returns partial text while someone speaks and then confirms final segments; introductory pricing is $0.54 per audio hour through the end of 2026.

A benchmark, not a contract

Microsoft says the model leads Artificial Analysis for both partial and final transcript accuracy. The public leaderboard currently covers 38 streaming models, but Azure’s documentation also labels the service a public preview without an SLA and says it is not recommended for production workloads.

What we would do

Virtual Arc would pilot it now through a provider-neutral interface and evaluate real calls, domain terms, accents and degraded connections. We would not move production traffic until general availability, durable pricing and our own tests show a lower cost per successfully completed conversation.

Our take

Microsoft’s MAI-Transcribe-2-Streaming is the strongest reason this week to reopen a voice-stack evaluation, but not a reason to migrate production. It supports live transcription in 60 languages, carries an introductory price of $0.54 per audio hour, and Microsoft says it leads Artificial Analysis for both partial and final transcript accuracy. The operational catch matters more than the leaderboard: Azure still labels it a public preview without an SLA and explicitly does not recommend it for production. At Virtual Arc, we would put it behind the same provider interface as the incumbent, replay a representative call corpus, and shadow a small slice of live traffic. We would judge it on cost per successfully completed conversation, correction rate, endpointing behavior, fallback frequency and p95 latency—not headline word-error rate alone. Until it reaches general availability, survives our languages and domain vocabulary, and has durable pricing, we would keep the current provider as the production default. The right move is to test now so migration stays cheap later, while refusing to make preview infrastructure part of the reliability budget.

Sources
  1. Microsoft AI: Our first streaming transcription model debuts at no. 1 on Artificial Analysis
  2. Microsoft Learn: MAI-Transcribe-2-Streaming overview
  3. Artificial Analysis: Speech to Text Providers Leaderboard and Comparison
  4. Vercel AI Gateway: MAI-Transcribe-2-Streaming API and Pricing

← All posts