通过多尺度编码与生成,解决语音合成中对粗粒度信息关注不足的问题。
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
- 用多尺度离散编码将语音分层表示,提升对长时结构的建模能力。
- 在零样本语音合成中,自然度和说话人相似度显著优于单尺度基线。
- 适合追求高质量语音合成的开发者与研究者,尤其关注生成稳定性与细节还原。
神经编解码语言模型(CLM)在文本到语音(TTS)合成中表现出色,但受制于‘近期偏差’,对更高时间尺度上的粗粒度信息关注不足,常导致语音不自然甚至不可懂。本文提出CoFi-Speech,一种自粗到细的CLM-TTS方法,采用多尺度语音编码与生成来解决此问题。我们训练了一个多尺度神经编解码器CoFi-Codec,将语音编码为具有不同时间分辨率的多层级离散表示,包含多个不同时间粒度的词元序列。随后,提出CoFi-LM,支持两种生成模式:基于单一语言模型的逐级生成与基于多语言模型的堆叠式生成。实验表明,CoFi-Speech在零样本TTS中显著优于单尺度基线系统,在自然度和说话人相似度上表现更优。多尺度编码分析验证了CoFi-Codec在学习多尺度离散语音表示方面的有效性,同时保持高质量语音重建。特别是堆叠式生成策略,被证实是构建高质量神经编解码语言模型的关键路径。
原文摘要 · Abstract (English)
The neural codec language model (CLM) has demonstrated remarkable performance in text-to-speech (TTS) synthesis. However, troubled by ``recency bias", CLM lacks sufficient attention to coarse-grained information at a higher temporal scale, often producing unnatural or even unintelligible speech. This work proposes CoFi-Speech, a coarse-to-fine CLM-TTS approach, employing multi-scale speech coding and generation to address this issue. We train a multi-scale neural codec, CoFi-Codec, to encode speech into a multi-scale discrete representation, comprising multiple token sequences with different time resolutions. Then, we propose CoFi-LM that can generate this representation in two modes: the single-LM-based chain-of-scale generation and the multiple-LM-based stack-of-scale generation. In experiments, CoFi-Speech significantly outperforms single-scale baseline systems on naturalness and speaker similarity in zero-shot TTS. The analysis of multi-scale coding demonstrates the effectiveness of CoFi-Codec in learning multi-scale discrete speech representations while keeping high-quality speech reconstruction. The coarse-to-fine multi-scale generation, especially for the stack-of-scale approach, is also validated as a crucial approach in pursuing a high-quality neural codec language model for TTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。