๐Ÿ‡ฒ๐Ÿ‡ฒ F5-TTS Burmese v2 (1,025,604 Updates Foundation Model)

The official v2 release of the F5-TTS (Flow-Matching Diffusion Transformer) Burmese speech foundation model, trained on 794.5+ hours of Burmese speech across 36 full epochs (1,025,604 updates).

"They are fighting for people's freedom. I am fighting for the language's freedom.
I just want to preserve their beautiful, lovely, and brave voices embedded in AI to last forever โ€” marking the first time in history for a massive Burmese open-source TTS foundation model."


๐ŸŒŸ What's New in v2

  • Over 1 Million Updates: Pushed training from 569,780 steps (v1) to 1,025,604 updates (36 epochs) for significantly tighter phonetic alignment and tone clarity.
  • 50% Smaller Model (FP16): Pruned weights down from 1.35 GB to 674 MB for 2x faster downloads and lower GPU VRAM consumption.
  • Production-Grade Zero-Shot Voice Cloning: Natural conversational cadence, improved tone handling on conjuncts and Pฤแธทi loanwords, and zero cut-offs.
  • Raw Checkpoints Available: Intermediate checkpoints (model_1020000.pt, model_1025000.pt, and model_last.pt) are included under checkpoints/ for developers wanting to continue fine-tuning.

โšก Quickstart: Python Package (f5-myanmar-tts)

The easiest way to use this model is with the official PyPI package:

pip install --upgrade f5-myanmar-tts

Python (Just 3 lines):

from f5_myanmar_tts import MyanmarTTS

# Auto-downloads lightweight v2 FP16 model (~674MB) on first run
tts = MyanmarTTS()

# Generate Burmese speech
tts.speak(
    "แ€œแ€ฐแ€žแ€ฌแ€ธแ€แ€ฝแ€ฑ แ€กแ€ฌแ€ธแ€œแ€ฏแ€ถแ€ธแ€€แ€ญแ€ฏ แ€กแ€šแ€ฏแ€แ€บแ€กแ€œแ€แ€บแ€กแ€™แ€ผแ€แ€บแ€™แ€›แ€ฝแ€ฑแ€ธ แ€แ€ปแ€…แ€บแ€แ€„แ€บแ€œแ€ฑแ€ธแ€…แ€ฌแ€ธแ€•แ€ซ", 
    output_file="speech.wav"
)

Voice Cloning (Clone any voice in 3โ€“5 seconds):

tts.speak(
    text="แ€’แ€ซแ€€แ€ผแ€ฑแ€ฌแ€„แ€ทแ€บ แ€กแ€ฏแ€ถแ€ทแ€™แ€พแ€ญแ€ฏแ€„แ€บแ€ธแ€”แ€ฑแ€แ€ฒแ€ท แ€€แ€ฑแ€ฌแ€„แ€บแ€ธแ€€แ€„แ€บแ€€แ€ญแ€ฏ แ€กแ€™แ€ญแ€”แ€ทแ€บแ€•แ€ฑแ€ธแ€•แ€ผแ€ฎแ€ธ แ€™แ€ญแ€ฏแ€ธแ€€แ€ฑแ€ฌแ€„แ€บแ€ธแ€€แ€„แ€บ แ€แ€ถแ€แ€ซแ€ธแ€แ€ฝแ€ฑแ€€แ€ญแ€ฏ แ€–แ€ฝแ€„แ€ทแ€บแ€œแ€ญแ€ฏแ€€แ€บแ€แ€šแ€บ",
    ref_audio="my_voice.wav",
    ref_text="แ€กแ€•แ€ผแ€„แ€บ แ€™แ€žแ€ฝแ€ฌแ€ธแ€›แ€œแ€ญแ€ฏแ€ท แ€…แ€ญแ€แ€บแ€Šแ€…แ€บแ€”แ€ฑแ€•แ€ซแ€แ€šแ€บ แ€™แ€ญแ€ฏแ€ธแ€แ€ฝแ€ฑ แ€แ€กแ€ฌแ€ธ แ€›แ€ฝแ€ฌแ€”แ€ฑแ€•แ€ซแ€แ€šแ€บ",
    output_file="cloned_speech.wav"
)

๐Ÿ“Š Model Specifications

Parameter Specification
Architecture Diffusion Transformer (DiT Base)
Parameters 337,138,310 (~337M)
Layers / Heads / Dim 22 layers, 16 heads, dim=1024, text_dim=512
Training Steps 36 Epochs (1,025,604 updates)
Training Audio 794.5+ Hours Burmese Speech
Sampling Rate 24,000 Hz
Vocoder Vocos (24kHz Mel)
Vocabulary 2,626 Burmese & Pฤแธทi Unicode Tokens
Format Safetensors (674 MB FP16 Pruned EMA weights)
License Creative Commons Attribution-NonCommercial 4.0 (CC BY-NC 4.0)

๐Ÿ•Š๏ธ Dedication & Acknowledgements

  1. GEMINI AI (Google): Co-engineering partner through every line of code, architecture debugging, memory optimizations, and fine-tuning pipelines.
  2. F5-TTS Research Team: Yushen Chen and the creators of F5-TTS for open-sourcing this world-class flow-matching speech architecture.
  3. The Brave Voices of Freedom: National Unity Government (NUG), PVTV broadcasters, independent journalists, and creators whose voices form the backbone of this heritage preservation project.

๐Ÿ“œ License

Released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. Dedicated to free public research, language preservation, education, and open-source innovation.

Downloads last month
370
Safetensors
Model size
0.3B params
Tensor type
F16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using freococo/F5-Myanmar-TTS 1