Close Menu
  • AI
  • Content Creation
  • Tech
  • Robotics
AI-trends.todayAI-trends.today
  • AI
  • Content Creation
  • Tech
  • Robotics
Trending
  • Meta says it’s going to run advertisements for ‘Musk’ documentary in spite of everything
  • Finish-to-Finish Multimodal Information Augmentation and Adversarial Robustness Benchmark with AugLy for Pictures, Textual content, Audio, and PyTorch
  • Meta’s Muse Is Adults-Solely. Why Does It Look Like a Children’ Toy?
  • Exa Launches Agent Extremely: A Subagent Swarm Deep Analysis API Constructed for Exhaustive Checklist Constructing
  • Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Imaginative and prescient-Language Fashions With As much as 3.13x Sooner Decoding
  • Thieves Stole ‘Nvidia’ Trailers. They Bought 20 Tons of Sand
  • Perplexity Trains Its Pc Agent on Actual Errors With Trace-Guided Self-Distillation
  • Appeals Court docket Lets the Pentagon Designate Anthropic a Provide-Chain Threat
AI-trends.todayAI-trends.today
Home»Tech»Meta Superintelligence Labs releases Muse voice Transcribe, a real-time model for streaming ASR, dialization and endpointing

Meta Superintelligence Labs releases Muse voice Transcribe, a real-time model for streaming ASR, dialization and endpointing

Tech By Gavin Wallace02/09/20266 Mins Read
Facebook Twitter LinkedIn Email
Meta AI Introduces Multi-SpatialMLLM: A Multi-Frame Spatial Understanding with Multi-modal
Meta AI Introduces Multi-SpatialMLLM: A Multi-Frame Spatial Understanding with Multi-modal
Share
Facebook Twitter LinkedIn Email

Three systems are usually used to create most production voice stacks. A model is used to transcribe, another separates the speakers and finally a detector determines whether or not the speaker has stopped speaking. With each handoff, latency is added and there’s a new mode of failure.

Muse Voice TranscribeMeta Superintelligence Labs announced this week that it has combined these three tasks into one auto-regressive model. Meta claims it is its first model of real-time sound perception. The model performs ASR streaming, speaker diarization of 20+ speakers and endpointing all in one go, without any post-processing.

Is it deployable? The API is hosted. It’s live on the Meta Model API As well as muse-voice-transcribe-1.0 Meta AI Mac and PC already has dictation powered by this technology, which costs $3.00 for 1,000 audio minutes (0.18% per hour). Muse Code. There are no weights available, therefore there is not a self-hosted route.

The foundation for streaming ASR

Muse Voice Transcribe comes from Muse Spark’s multimodal family. Audio comes in 80ms pieces at 12.5. Each chunk becomes a soft token.

It makes a binary selection after every chunk. Either it predicts or it doesn’t. The model can either emit a sound token or keep listening. If the model is able to predict . That token is then replaced by the audio data in the input. The stream will end when the audio is finished. When the token is placed, it flushes out all of the remaining audio without needing to request more.

The alignment of listening and writing is done by the same decoder, therefore there’s no need for a separate stage.

Adaptive delay with RL

It also determines how much context audio is behind each phrase, because the model has control over when it listens. Meta calls that gap ‘delay.’ The longer the delay, the more accurate and latency-increasing transcript will be.

Meta trains the trade-off instead of fixing it. Reinforcement Learning combines word error rates and delay rewards multiplicatively to produce a policy which varies the delay per word based on difficulty. Meta reports that the new model, based on time taken to complete the transcription, is ahead of previous systems such as Soniox Cartesia and ElevenLabs.

More tokens for Diarization, Endpointing and other Tokens

Meta has not added a second speaker-attribution model. It added tokens for special purposes to the stream.

For diarization, a A token indicates a speaker that could be switched, as well as a The tag identifies a speaker. Turn token is fired as soon as switchable, but speaker tag will be delayed until the end of chunk. One speaker’s audio can be broken up into multiple segments which resolve all to the same tag.

The endpointing of the arrow Marks the beginning of speech The user can mark the exact point that they finished. The ASR stream is used to train both tasks, with extra rewards on top.

Capabilities

Model was tested on more than 70 languages. 25 of them were extensively checked and are recommended for launch. The model is able to switch code both in a sentence as well as between sentences. This can be important for multilingual speakers that mix up languages during a clause. Language, keyword and context biasing can improve accuracy.

A practical difference is the ability to manage long contexts. Meta claims that the model supports native audio inputs exceeding an hour, and up to 20 speakers without any post-processing required.

Benchmarks

Meta is ranked first in Artificial Analysis of streaming speech-to text and public diarization benchmarks as of September 1, 2020.

The following are some of the ways to get in touch with us. Artificial Analysis AA-WER StreamingMuse Voice Transcribe records a final-transcripted WER of 3.1% at 0.16s from the end of the speech. Cartesia Ink-2 is 3.4% with semantic ends at 0.43s. ElevenLabs Scribe Realtime has a 3.6% accuracy at 0.14s. Cartesia Ink-2, with external endpoints, is the fastest (0.07s) but also the least accurate (4.0%). Muse Voice Transcribe, on the first partial transcript it records at 0.13s, 3.6% of WER.

Meta reported a diarization average error rate of 17.5% across AMI IHM, AMI SDM and VoxConverse. The five other systems on the chart are between 21,1% and 28,6%.

The price is also a major factor. With a price of $3.00 per 1000 minutes, this product is cheaper than Cartesia Ink-2 ($4.00) and less expensive than ElevenLabs Scribe Realtime or Deepgram Flux ($6.50).

Interactive explainer

intel intelligence met meta Streaming
Share. Facebook Twitter LinkedIn Email
Avatar
Gavin Wallace

Related Posts

Finish-to-Finish Multimodal Information Augmentation and Adversarial Robustness Benchmark with AugLy for Pictures, Textual content, Audio, and PyTorch

26/09/2026

Exa Launches Agent Extremely: A Subagent Swarm Deep Analysis API Constructed for Exhaustive Checklist Constructing

26/09/2026

Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Imaginative and prescient-Language Fashions With As much as 3.13x Sooner Decoding

26/09/2026

Perplexity Trains Its Pc Agent on Actual Errors With Trace-Guided Self-Distillation

25/09/2026
Top News

US Tech Giants race to spend Billions in UK Artificial Intelligence Push

Redditors use AI to lower outrageous World Cup ticket costs

Google Pixel 11 smartphones: New camera features

The Enigma of Enforcing GDPR on LLMs • AI Blog

Power Play: The Great Big Power Play

Load More
AI-Trends.Today

Your daily source of AI news and trends. Stay up to date with everything AI and automation!

X (Twitter) Instagram
Top Insights

Newbie’s Information to Social Media Promoting (+ Prices on Every Platform)

25/08/2025

The Actual Demon Inside ChatGPT

29/07/2025
Latest News

Meta says it’s going to run advertisements for ‘Musk’ documentary in spite of everything

26/09/2026

Finish-to-Finish Multimodal Information Augmentation and Adversarial Robustness Benchmark with AugLy for Pictures, Textual content, Audio, and PyTorch

26/09/2026
X (Twitter) Instagram
  • Privacy Policy
  • Contact Us
  • Terms and Conditions
© 2026 AI-Trends.Today

Type above and press Enter to search. Press Esc to cancel.