Microsoft's open-source VibeVoice-ASR: one pass over a 60-minute meeting yields speakers, timestamps and content; 54k GitHub stars

Microsoft's open VibeVoice voice family (54k+ GitHub stars, MIT license) includes VibeVoice-ASR (7B, open-sourced Jan 2026), a unified speech-to-text model that transcribes up to 60 minutes of meeting audio in a **single forward pass** — outputting structured transcripts with who (speaker ID), when (timestamps) and what (content); one model replaces the usual transcription + diarization + alignment stack. It natively supports 50+ languages with automatic detection, mid-utterance code-switching (e.g. English-Mandarin), and custom hotwords/context for names and domain terms; runs fully locally so sensitive recordings never leave the device. The foundation is the family's shared 7.5 Hz ultra-low-framerate continuous speech tokenizer (vs 50-75 Hz typical; ~3200x compression — 90 minutes of audio fits in ~40k tokens of LLM context) plus an LLM backbone and diffusion head. The family also ships TTS-1.5B (90-min, 4-speaker long-form generation, ICLR 2026 Oral), Realtime-0.5B (~300 ms first-audio streaming TTS), ASR-Streaming (live who-said-what, Sep 2026), and ASR-BitNet (heterogeneous quantization to 1.58 GB with real-time CPU-only inference via VibeASR.cpp, no GPU needed).




