用离散隐变量联合优化语音编码器与语言模型,提升语音生成与识别效果。
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
- 在梅尔频谱上构建离散隐变量模型,联合训练编码器与自回归模型。
- 零样本语音合成与语音识别任务上超越基线,静音过长和漏词问题减少。
- 适合关注语音建模与自回归生成优化的研究者。
近期的语音语言模型依赖于与自回归模型独立优化的编码器。由于这些编码器不了解下游任务目标,提取的表征可能并非最优。为解决这一问题,我们提出一种基于梅尔频谱的离散隐变量模型,联合优化编码器与语音语言模型。联合优化不仅在零样本文本到语音(TTS)和语音到文本(STT)任务上优于基于编解码器和其他梅尔频谱基线模型,还有效缓解了自回归梅尔频谱建模中的常见问题,如生成过长静音和单词遗漏。
原文摘要 · Abstract (English)
Recent speech language models rely on encoders that are optimized separately from autoregressive models. Since these encoders are unaware of the downstream objectives, the extracted representations may not be optimal for downstream tasks. To address this limitation, we introduce a discrete latent variable model on mel spectrograms that jointly optimizes the encoder and the speech language model. Joint optimization not only brings improvements over codec-based and other mel-spectrogram-based baselines on zero-shot Text-to-Speech (TTS) and Speech-to-Text (STT) tasks, but also effectively alleviates common issues in autoregressive mel spectrogram modeling, such as prolonged silence generation and word omissions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。