News
AI Summary
27 Apr 202610 Dhuʻl-Qiʻdah 1447 AH
microsoft/VibeVoice

microsoft/VibeVoice

Microsoft launched the VibeVoice speech-to-text model on January 21, 2026, featuring MIT licensing and integrated speaker diarization. This model was tested on a Mac using mlx-audio tools, demonstrating its effectiveness with various audio files. VibeVoice allows users to convert audio files into text with high accuracy, evidenced by its application in a podcast episode with Lenny Rachitsky. The results showed a processing time of 524.79 seconds, generating 20,248 tokens at a rate of 38.585 tokens per second, indicating the model's efficiency with lengthy audio.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In