-
A Simple Exploration of Discrete Diffusion Distillation for TTS
I had a quick exploration on how autoregressive speech synthesis could leverage from block-causality applied to the latest discrete text-to-speech (TTS) models. Despite their low-confidence outputs, with only a simple distillation, it shows a trade-off point of the inference speed up to x2~3 times compared to the original without any further optimization with minimal degradation on speech quality and zero-shot capability.