🗣️ WavSLM

A single-stream speech language model based on WavLM distillation and FocalCodec.

This repository contains the checkpoint with a codebook size of 65536 trained on Libri-Light, as described in the paper.


▶️ Quickstart

See the readme at: https://github.com/lucadellalib/wavslm


@ Citing

@inproceedings{dellalibera2026wavslm,
    title     = {{WavSLM}: Single-Stream Speech Language Modeling via {WavLM} Distillation},
    author    = {Luca {Della Libera} and Cem Subakan and Mirco Ravanelli},
    booktitle = {Interspeech},
    year      = {2026},
}

📧 Contact

luca.dellalib@gmail.com


Downloads last month
15
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lucadellalib/wavslm_65k

Finetuned
(1)
this model

Collection including lucadellalib/wavslm_65k

Paper for lucadellalib/wavslm_65k