https://arxiv.org/abs/2408.16532
jishengpeng
novateur
AI & ML interests
speech language model, discrete codec, text to speech
Recent Activity
authored a paper 1 day ago
Omni Interaction Agent Technical Report authored a paper 1 day ago
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for
Zero-Shot Speech Synthesis authored a paper 1 day ago
WavChat: A Survey of Spoken Dialogue Models