arXiv:2409.04016cs.SDeess.AS2024-09中稿 · SLT-2024被引 10

探究音频编码器对语音大模型生成效果的影响,发现重建质量不等于生成质量。

Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation

  • 统一数据和损失函数,对比多个编码器在语音大模型中的表现。
  • 高保真解码器更利于自然语音生成,量化机制影响语音可懂度。
  • 为语音生成用的音频编码器设计提供关键指导,适合语音合成研究者。

神经音频编码器的标记是基于语音语言模型(SLM)语音生成的基础构建块。然而,目前尚缺乏对编码系统如何影响SLM语音生成性能的系统性理解。本文在统一设置下,重新训练现有高性能神经编码模型,以比较其在相同数据集和损失函数下的表现,并将编码器标记集成到两种SLM系统中:基于掩码的并行语音生成系统和基于自回归(AR)与非自回归(NAR)模型的混合系统。研究发现,编码系统中更好的语音重建并不保证在SLM中获得更优的语音生成效果。高质量的编码器解码器对生成自然语音至关重要,而语音可懂度则更多取决于量化机制。本工作为高效编码器设计提供了重要洞见。

原文摘要 · Abstract (English)

Neural audio codec tokens serve as the fundamental building blocks for speech language model (SLM)-based speech generation. However, there is no systematic understanding on how the codec system affects the speech generation performance of the SLM. In this work, we examine codec tokens within SLM framework for speech generation to provide insights for effective codec design. We retrain existing high-performing neural codec models on the same data set and loss functions to compare their performance in a uniform setting. We integrate codec tokens into two SLM systems: masked-based parallel speech generation system and an auto-regressive (AR) plus non-auto-regressive (NAR) model-based system. Our findings indicate that better speech reconstruction in codec systems does not guarantee improved speech generation in SLM. A high-quality codec decoder is crucial for natural speech production in SLM, while speech intelligibility depends more on quantization mechanism.

语音生成音频编码语言模型量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。