ChiReSSD
Speaker-preserving reconstruction of disordered speech.
Given a target transcript and a short reference recording of a speaker, ChiReSSD synthesizes that transcript with canonical pronunciation while keeping the speaker's voice and prosody. Pronunciation enters through the text pathway; identity and prosody come from the style pathway. That separation is the point: ordinary style-preserving TTS treats disordered articulation as part of the speaker's style and so reproduces the mispronunciation it was meant to correct.
- Code: https://github.com/Lab-MSP/ChiReSSD
- Paper: Generative Reconstruction of Pediatric Disordered Speech for Automated Clinical Evaluation
Model description
StyleTTS2 fine-tuned on child disordered speech. Two 128-dimensional style vectors are
extracted from the reference recording โ acoustic (timbre) and prosodic โ and interpolated
with a style sampled from the adapted diffusion prior. alpha weights the acoustic side and
beta the prosodic; 1.0 is fully diffusion-sampled, 0.0 fully reference-driven.
It works on unseen speakers from a reference as short as a few seconds. No per-speaker model is trained, and no paired typical/atypical recordings are required.
Base model and license chain
Fine-tuned from the StyleTTS2 LibriTTS second-stage checkpoint by yl4579 (MIT). That base checkpoint is not redistributed here โ obtain it from the upstream release. The frozen helper models (ASR text aligner, JDCNet pitch extractor, PL-BERT) likewise ship inside the upstream repository and are not redistributed.
ChiReSSD builds upon StyleTTS2, with its modifications published as patch files in the code repository.
Intended use
Research on speech reconstruction and on automated clinical evaluation of speech sound disorders.
Out of scope
- Not a medical device. No diagnostic or treatment decision should rest on its output.
- Not a replacement for assessment by a licensed speech-language pathologist.
Training data
Child speech from UltraSuite/CLP-derived recordings, obtained under their own data use
agreements. The corpus is not released here and is not redistributable: it is identifiable
child clinical speech. To fine-tune your own model,
https://huggingface.co/datasets/changelinglab/ultrasuite-benchmark is a suitable starting
point; see DATA.md in the code repository.
Training configuration
| Epochs | 4 |
| Batch size | 4 |
| Max length | 600 frames |
| Learning rate | 1e-5 (lr, bert_lr, ft_lr) |
lambda_F0 |
20 (upstream: 1) |
lambda_mel |
5 |
| Style diffusion from epoch | 2 |
| Joint SLM-adversarial from epoch | 3 |
| Decoder | HiFi-GAN, multispeaker |
| Sample rate | 24 kHz |
| LR schedule | OneCycleLR, pct_start=0.1 (upstream: 0) |
The heavy pitch weighting is deliberate: the pitch extractor was pretrained on adult voices, and children's F0 is both higher and more variable, so it needs the strongest adaptation of any component. Conversely, only four epochs โ longer schedules start fitting the disordered articulation itself.
Trained on 2ร 48 GB GPUs. At batch 4 and max_len 600 the recipe needs more than 48 GB, so a
single smaller card requires lowering both.
Inference
Two presets ship with the code:
| Preset | alpha | beta | steps | Purpose |
|---|---|---|---|---|
default |
0.8 | 0.5 | 10 | The released operating point |
torgo |
1.0 | 0.5 | 15 | Adult dysarthric speech |
alpha is high on purpose. A low alpha leans on the reference acoustics, which is exactly
where the disordered articulation lives.
Synthesis is stochastic. The initial style latent is zeros rather than a Gaussian draw,
but the ADPM2 sampler is ancestral and adds fresh noise at every step, so repeated calls
differ in waveform and in duration. Pass a seed for reproducible output.
Usage
from chiressd.model import ChiReSSD
from chiressd.config import load_preset
model = ChiReSSD.from_pretrained() # downloads this checkpoint
style = model.compute_style("speaker_reference.wav")
wav = model.synthesize(
"butterfly butterfly butterfly",
style,
seed=1234,
**load_preset("default").as_kwargs(),
)
Run chiressd-setup first: it clones and patches the upstream StyleTTS2 checkout that
supplies the architecture and frozen helper models.
Checkpoint provenance
Derived from the fine-tuning run's final checkpoint by keeping state['net'] only, removing
the module. prefix that DataParallel added to ten of the thirteen submodules, and making
every tensor detached, CPU-resident and contiguous. Precision is unchanged (float32; no fp16
cast, which would alter outputs).
Citation
Rosero, Yeo, Mortensen, Van't Slot, Hallac, and Busso. Generative Reconstruction of Pediatric Disordered Speech for Automated Clinical Evaluation. In Proceedings of the IEEE Spoken Language Technology Workshop 2026 (SLT '26), Palermo, Italy, 2026.