Microsoft AI announced last week MAI-Transcribe-1.5. The company has developed a second generation of its own speech-totext technology. Models are designed for accuracy in 42 languages and environments with noise. Microsoft positions the model for transcription production workloads.
What is MAI Transcribe 1.5
This model is an ASR (automatic speech recognition) model. This model uses audio to input text. Microsoft developed it internally, and not using a base from a third party. It can handle 43 languages using a single model. The model is optimised for accents and dialects as well as real world acoustic situations.
Microsoft has integrated it into Copilot Teams, GitHub and Dynamics 365. Microsoft’s Foundry model platform also offers it.
Accuracy Case
WER is a measure of accuracy. A lower WER indicates fewer errors per word. Microsoft has reported the best WER in 43 languages using FLEURS. FLEURS, a multilingual standard benchmark for transcription is the industry-standard.
The model has a WER score of 2,4% on the Artificial Analysis leaderboard. It is ranked third amongst a set of open benchmarks. The picture looks split. Microsoft Team claims first place in FLEURS but third in Artificial Analyses.
Another story of accuracy is language expansion. From 25 languages, the coverage has grown to 43. It was possible to add 18 additional languages without losing accuracy. Ten are South Asian languages, such as Bengali, Tamil and Telugu. The other eight are European languages, including Greek, Ukrainian and Catalan.
The speed of the vehicle
The Artificial Analysis leaderboard shows that MAI-Transcribe 1.5 is the fastest and most accurate of all models. The model can run up to five times faster than other models with comparable accuracy. This effect is most noticeable on audio files that are long. This model is capable of trancribing an hour’s worth of audio files in less than 15 seconds.
Microsoft claims that it can speed up long audio by as much as 5 times compared to Gemini 3.1. Scribe v2 and GPT-4o Transcribe. The Azure card claims up to 5.7x speedier long-form analysis compared with the previous MAI-Transcribe-1. The latency gap can quickly increase for large batch processing pipelines.
This is the feature worth knowing about: Keyword Biasing
Generic transcriptions are often unable to capture words that belong in a specific domain. This includes people, products, medical terms and internal acronyms. These words are often the most important to users in enterprise environments.
In MAI-Transcribe 1.5, keyword biasing is also known as entity biasing. The list is domain-specific. Azure supports 200 keywords. This list is used to bias the model’s predictions. It does not force blindly matches. The shared context is used to determine when the biasing feature should be applied. Microsoft reported a 30% reduction in WER on FLEURS using biasing.
The effect is demonstrated in a simple example. Names are rendered as they appear without bias. “Sean,” “Oif,” The following are some examples of how to get started: “Societal.” The model can be recovered with a name list. “Shaun,” “Aoife,” The following are some examples of how to get started: “Xochitl.” It is useful for healthcare, call centers, meetings and niche vocabulary.
Use Cases
The Azure model card lists concrete production scenarios. Each card corresponds to a typical engineering workload.
- Videos captioned Media and Content Platforms
- Accessibility tools You can rely on captions that are accurate.
- Meeting transcription Teams collaboration tools are a great way to collaborate.
- Call analysis Contact centers and analytics support.
- Content creation workflows You need a fast draft of transcripts.
- Vocal agents It is better to convert your speech to text than reasoning.
The automatic identification of the language helps to identify input languages that are unknown. It detects without manual input the language spoken.
The difference between MAI-Transcribe-1 and MAI-Transcribe 1.5
This table compares two generations by stating only the facts.
| Attribute | MAI-Transcribe-1 | MAI-Transcribe-1.5 |
|---|---|---|
| Covered languages | 25 | 43 |
| Biasing by keyword/entity | No listings | You can use up to 200 Keywords |
| The speed of long-form inference | Baseline | Faster than 5.7 times |
| Artificial Analysis WER | Other than that, | 2.4% (ranked #3) |
| FLEURS position (per Microsoft) | State-of-the-art | The best in class across 43 languages |
| Automated language recognition | Other than that, | You can say that. |
| Life Cycle | Prior release | General availability (GA) |
| Input / output | Audio / Text | Audio / Text |
Strengths and limits
Strengths:
- A single model can cover 43 languages instead of 25.
- The WER can be reduced by up to 30% when using FLEURS.
- Transcript an audio hour in less than 15 seconds.
- Azure AI Foundry offers a wide range of services.
- Microsoft says that the software is robust on real world audio with noisy noise.
Limitations:
- No diarization yet, so speaker labels are unavailable.
- There is no native API for streaming, which limits the real-time usage.
- First-party claims have been made for accuracy, cost, and speed.
- Ranking third in Artificial Analysis behind two competitors
Sources

