VibeVoice is Microsoft Research's family of open voice models. It became famous in August 2025 for VibeVoice-TTS, which generated up to 90 minutes of podcast-style conversation between four speakers and was accepted as an oral paper at ICLR 2026. Within two weeks Microsoft found it being misused and removed the TTS code from the repository; the 1.5B weights remain on Hugging Face, but the hosted demo is disabled.
What the repository offers today is mostly speech recognition. VibeVoice-ASR transcribes up to 60 minutes of audio in one pass and labels who spoke and when, in more than 50 languages, with custom hotwords for names and jargon; it has a streaming variant, vLLM serving, fine-tuning code, a Hugging Face Transformers integration and a CPU-only BitNet build. On the speech side, VibeVoice-Realtime-0.5B is a small streaming English voice with about 300 ms to first audio and preset speakers only, deliberately without voice cloning. The code is MIT-licensed and the repository has about 55,000 stars.
- Repository: github.com/microsoft/VibeVoice
- Licence: MIT (MIT License)
- Language: Python. Stars: 54.6K. Forks: 6,141. Last push: Sep 3, 2026.
- Scan: safe, Sep 3, 2026, commit 1541f59
Who it is for
Developers and researchers who need long-form transcription with speaker labels (meetings, interviews, podcasts), and builders who want a light, fast English voice for an assistant.
Getting started
1. Clone and install (an NVIDIA GPU and Microsoft's suggested PyTorch container are recommended)
git clone https://github.com/microsoft/VibeVoice.git && cd VibeVoice && pip install -e .2. Transcribe in a local web demo
python demo/vibevoice_asr_gradio_demo.py --model_path microsoft/VibeVoice-ASR3. For realtime speech, install the extra and run its demo
pip install -e .[streamingtts] && python demo/vibevoice_realtime_demo.py --model_path microsoft/VibeVoice-Realtime-0.5BVibeVoice-ASR is a 7B model and wants a CUDA GPU with plenty of memory; the separate VibeASR.cpp project runs a BitNet version on a CPU. The README's demo command adds --share, which publishes a public Gradio link; it is left out above. Microsoft says the models are for research and not recommended for commercial use without further testing.
Safety scan
We cloned microsoft/VibeVoice at commit 1541f59 on Sep 3, 2026 and ran the checks described on the GitHub Tools page: credential patterns, decode-and-execute code, install-time scripts, committed binaries, risky CI workflows, every host the code talks to, known vulnerabilities in pinned dependencies, and project hygiene. A person read every hit. This is what we found.
- No secrets, no pattern hits, no committed binaries, no bare-IP URLs and no install hooks across 110 files and about 24,000 lines of Python. The code talks to Hugging Face for model weights and to nothing else at run time.
- The removed long-form TTS really is gone from the scanned commit: what remains is the ASR model, its streaming and vLLM variants, fine-tuning code and the Realtime-0.5B voice, whose speakers ship as embedded voice prompts rather than a cloning path.
- pyproject.toml lists dependencies without pinned versions and there is no lockfile, so the scan had nothing to compare with advisory databases; you get whatever versions pip resolves on the day. Microsoft's docs suggest installing inside an NVIDIA PyTorch container, which also keeps it away from your system Python.
- The ASR Gradio demo command in the docs includes --share, which publishes a temporary public link to your machine; leave it off for local use.
- No workflows. Security policy, licence and contributing guide present; no Dependabot or CodeQL, and no tagged releases.
What the scanner counted
| Check | Result |
|---|---|
| Secrets | None found. |
| Suspicious code | None found. |
| Install-time code | None: nothing runs at install beyond the package manager itself. |
| Committed binaries | None. |
| CI workflows | No GitHub Actions workflows. |
| Network hosts | 6 distinct hosts referenced from source; most often github.com, arxiv.org, huggingface.co, colab.research.google.com. No URLs to bare IP addresses. |
| Known vulnerabilities | No lockfile to check: dependencies are declared as ranges, so what gets installed is whatever is current on the day. |
| Project hygiene | Has security policy, licence file, contributing guide. Missing automated dependency updates, CodeQL. |
| OpenSSF Scorecard | Not scored: the project is not in Scorecard's weekly index. |
By the numbers
| Stars | 54.6K |
|---|---|
| Forks | 6,141 |
| Contributors | 23 |
| Commits | 159 |
| Open issues | 129 |
| Open pull requests | 70 |
| Releases | 0 |
| Latest release | none tagged |
| Licence | MIT |
| Main language | Python |
| Project age | 1 year |
| Last push | Sep 3, 2026 |
| Tracked files | 110 |
| Lines of code | 23.7K |
| Checkout size | 150 MB |
Lines by language: Python 20.5K, Markdown 1,513, HTML 1,018, JSON 379, Jupyter 217, TOML 57.
Questions
Is VibeVoice free?
Yes. The code is MIT-licensed and the model weights are free to download from Hugging Face. Microsoft describes the models as research releases and advises against commercial or real-world deployment without more testing. VibeVoice-ASR can also be tried in Azure AI Foundry Labs.
Can I still use VibeVoice to make AI podcasts?
Not from this repository. Microsoft removed the long-form multi-speaker TTS code in September 2025 after misuse, and its demo is disabled. The remaining Realtime-0.5B voice is single-speaker, English-focused, and uses only preset voices.
How is VibeVoice-ASR different from Whisper?
Whisper transcribes in 30-second windows and needs a separate tool to say who spoke. VibeVoice-ASR takes up to an hour of audio at once and outputs speaker, timestamps and text together, which keeps speaker labels consistent over a long recording. It is larger and needs more GPU memory than most Whisper sizes.
This post is part of GitHub Tools, where every repository is cloned and scanned before it is written up. The scan is a snapshot of one commit on one day; the repository has moved on since, so check it before you install.
