Skip to main content
MicrosoftLaunch

Microsoft unveils new transcription, voice models

What happened

Microsoft has launched MAI-Transcribe-2-Streaming, a new artificial intelligence model for live speech-to-text transcription available in 60 languages.

Source

Article excerpt

Highlighted: the sentence this signal was extracted from

Microsoft MSFT unveiled three new artificial intelligence models on Thursday, aimed at transcription and voice. The MAI-Transcribe-2-Streaming model is aimed at transcription, turning live speech into text as it arrives, Microsoft said in a blog post. "Instead of waiting for someone to finish speaking before returning text, MAI-Transcribe-2-Streaming transcribes continuously across 60 languages, with automatic language detection," Microsoft wrote in the post. "It produces its first hypotheses, known as partials, within the low hundreds of milliseconds of receiving audio, then refines them as more context arrives and commits a stable transcript once the utterance ends. That distinction matters when an application needs to act while someone is speaking. A customer-service agent can begin identifying a caller's request before the sentence is complete. A voice assistant can start reasoning or preparing a tool call sooner. A live transcription experience can surface words almost as quickly as they are spoken." MAI-Transcribe-2-Streaming is the top transcription model for accuracy on both partial and final transcripts, according to Artificial Analysis. In most cases, words appear as early as 320 milliseconds after they're spoken, compared to more than 500 milliseconds for the competition. MAI-Transcribe-2-Streaming is available in 60 languages and costs $0.54 per hour on an...

Keep reading with a free account

The rest of this article, and every signal for Microsoft, is in your free account.

Extracted by Autobound

From the Signal API record
Event
Launch

What this signalsA launch often needs new go-to-market and support spend.

Product
MAI-Transcribe-2-Streaming

The full record

From the Signal API record

Details

Release type
Model

Topics and mentions

Product tags

  • online technology
  • general technology
  • future tech
  • audio

Extraction

Confidence
90%
Detected
Oct 1, 2026
signal_type
news
signal_subtype
launches

Use this data

Get every launch signal for Microsoft and the companies you sell to, in the tools you already use.

  1. Ask Claude about it

    Connect Autobound to Claude, Claude Code or Cursor with MCP. Then ask: “What changed at Microsoft this week?”

  2. Send it to your own tools

    The Signal API returns launch signals for any list of companies as JSON, for your CRM, warehouse or app.

  3. Try it free

    Sign up and spend your free credits on the companies you sell to.

    Start Free1,000 free credits

The API returns more than this page shows

This page shows a preview. The full news record in the Signal API and MCP can also have these 8 fields. Some fields are empty for some signals.

Company

  • linkedin_urlValue in the API
  • industriesValue in the API
  • employee_count_lowValue in the API
  • employee_count_highValue in the API
  • revenueValue in the API
  • descriptionValue in the API

Signal

  • signal_nameValue in the API
  • associationValue in the API
Show the full JSONThe record on this page and the API request

GET /v1/signals/199f3cd4-3c65-cf51-bef6-8a8e57c0204c returns this record as JSON. POST /v1/companies/enrich returns every signal for microsoft.com.

{
  "signal_id": "199f3cd4-3c65-cf51-bef6-8a8e57c0204c",
  "signal_type": "news",
  "signal_subtype": "launches",
  "detected_at": "2026-10-01T16:44:01+00:00",
  "company": {
    "name": "Microsoft",
    "domain": "microsoft.com"
  },
  "data": {
    "url": "https://www.tradingview.com/news/seekingalpha:f7f26d436094b:0-microsoft-unveils-new-transcription-voice-models/",
    "title": "Microsoft unveils new transcription, voice models - TradingView",
    "excerpt": "Microsoft MSFT unveiled three new artificial intelligence models on Thursday, aimed at transcription and voice. The MAI-Transcribe-2-Streaming model is aimed at transcription, turning live speech into text as it arrives, Microsoft said in a blog post. \"Instead of waiting for someone to finish speaking before returning text, MAI-Transcribe-2-Streaming transcribes continuously across 60 languages, with automatic language detection,\" Microsoft wrote in the post. \"It produces its first hypotheses, known as partials, within the low hundreds of milliseconds of receiving audio, then refines them as more context arrives and commits a stable transcript once the utterance ends. That distinction matters when an application needs to act while someone is speaking. A customer-service agent can begin identifying a caller’s request before the sentence is complete. A voice assistant can start reasoning or preparing a tool call sooner. A live transcription experience can surface words almost as quickly as they are spoken.\" MAI-Transcribe-2-Streaming is the top transcription model for accuracy on both partial and final transcripts, according to Artificial Analysis. In most cases, words appear as early as 320 milliseconds after they're spoken, compared to more than 500 milliseconds for the competition. MAI-Transcribe-2-Streaming is available in 60 languages and costs $0.54 per hour on an...",
    "product": "MAI-Transcribe-2-Streaming",
    "summary": "Microsoft has launched MAI-Transcribe-2-Streaming, a new artificial intelligence model for live speech-to-text transcription available in 60 languages.",
    "planning": false,
    "image_url": "https://s.tradingview.com/static/images/illustrations/news-story.jpg",
    "confidence": 0.9,
    "product_data": {
      "name": "MAI-Transcribe-2-Streaming",
      "full_text": "The MAI-Transcribe-2-Streaming model",
      "fuzzy_match": false,
      "release_type": "model"
    },
    "product_tags": [
      "online_technology",
      "general_technology",
      "future_tech",
      "audio"
    ],
    "published_at": "2026-10-01T16:44:01Z",
    "article_sentence": "The MAI-Transcribe-2-Streaming model is aimed at transcription, turning live speech into text as it arrives, Microsoft said in a blog post."
  }
}

Long text fields are shortened on this page.

Looking up one signal by its id is free. Enrich costs 2 credits per signal returned; a call with no results is free.