Instructions to use ResembleAI/chatterbox with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Chatterbox
How to use ResembleAI/chatterbox with Chatterbox:
# pip install chatterbox-tts import torchaudio as ta from chatterbox.tts import ChatterboxTTS model = ChatterboxTTS.from_pretrained(device="cuda") text = "Ezreal and Jinx teamed up with Ahri, Yasuo, and Teemo to take down the enemy's Nexus in an epic late-game pentakill." wav = model.generate(text) ta.save("test-1.wav", wav, model.sr) # If you want to synthesize with a different voice, specify the audio prompt AUDIO_PROMPT_PATH="YOUR_FILE.wav" wav = model.generate(text, audio_prompt_path=AUDIO_PROMPT_PATH) ta.save("test-2.wav", wav, model.sr) - Inference
- Notebooks
- Google Colab
- Kaggle
Works great in ComfyUI
I found this tutorial online and found it very easy to get Chattterbox running in ComfyUI:
https://www.nextdiffusion.ai/tutorials/chatterbox-in-comfyui-tts-voice-cloning-conversion
I was able to clone my voice easily and had fun uploading and playing with others. What a great model. Thank you!
what even supported languages it can generate ?
On their website, their product supports 150 languages (not sure that applies to Chatterbox, which is opensourced.) More info here: https://www.resemble.ai/pricing/?trail=Pricing
Worth separating the two numbers: the 150 languages figure covers Resemble's hosted product, and open-sourced Chatterbox supports a smaller set. Outside that set the usual result is the right words in the wrong phonology, which lands as an accent rather than an obvious error. Good tutorial find though, ComfyUI makes the voice-conversion side much easier to experiment with.