Microsoft unveils new transcription, voice models
Article excerpt
Highlighted: the sentence this signal was extracted from
Microsoft MSFT unveiled three new artificial intelligence models on Thursday, aimed at transcription and voice. The MAI-Transcribe-2-Streaming model is aimed at transcription, turning live speech into text as it arrives, Microsoft said in a blog post. "Instead of waiting for someone to finish speaking before returning text, MAI-Transcribe-2-Streaming transcribes continuously across 60 languages, with automatic language detection," Microsoft wrote in the post. "It produces its first hypotheses, known as partials, within the low hundreds of milliseconds of receiving audio, then refines them as more context arrives and commits a stable transcript once the utterance ends. That distinction matters when an application needs to act while someone is speaking. A customer-service agent can begin identifying a caller's request before the sentence is complete. A voice assistant can start reasoning or preparing a tool call sooner. A live transcription experience can surface words almost as quickly as they are spoken." MAI-Transcribe-2-Streaming is the top transcription model for accuracy on both partial and final transcripts, according to Artificial Analysis. In most cases, words appear as early as 320 milliseconds after they're spoken, compared to more than 500 milliseconds for the competition. MAI-Transcribe-2-Streaming is available in 60 languages and costs $0.54 per hour on an...
Keep reading with a free account
The rest of this article, and every signal for Microsoft, is in your free account.
