Instructions to use microsoft/VibeVoice-1.5B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use microsoft/VibeVoice-1.5B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="microsoft/VibeVoice-1.5B")# Load model directly from transformers import AutoModelForSeq2SeqLM model = AutoModelForSeq2SeqLM.from_pretrained("microsoft/VibeVoice-1.5B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
[audio.cpp] C++/GGML VibeVoice 1.5B/7B/LORA: 93.9-minute podcast in 18.2 mins (10 steps, No quantization) 5.15x faster than real time.
#53
by audio-cpp - opened
Tested on RTX 5090. Check the repo here: https://github.com/0xShug0/audio.cpp
Also a demo of VibeVoice 1.5B on iPhone. 1.28x faster than realtime and ~2.2GB VRAM.https://www.reddit.com/r/StableDiffusion/s/SCvgmszwdp