arXiv:2409.00933cs.SDeess.AS2024-09中稿 · SLT 2024被引 17

用多流语义编码压缩语音序列,提升语言模型语音合成效率

SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis

  • 将语音转为有序多流离散语义序列,减少模型处理负担
  • 在240毫秒帧移下仍优于基线系统,压缩比达12倍
  • 适合追求高效语音合成的工业级应用

基于语言模型的语音合成面临长语音序列带来的建模复杂度与效率问题。本文提出SoCodec,一种语义有序的多流语音编码器,将语音压缩为更短的多流离散语义序列,每帧包含多个标记。同时引入有序产品量化,约束该序列为有序表示。可配合多流延迟语言模型,在时间和流维度上实现更优的自回归生成。实验表明,即使将语音帧移从20毫秒压缩至240毫秒(12倍),该方法仍显著优于基线系统。消融实验证明学习该有序多流语义表示对实现更短语音序列至关重要。

原文摘要 · Abstract (English)

The long speech sequence has been troubling language models (LM) based TTS approaches in terms of modeling complexity and efficiency. This work proposes SoCodec, a semantic-ordered multi-stream speech codec, to address this issue. It compresses speech into a shorter, multi-stream discrete semantic sequence with multiple tokens at each frame. Meanwhile, the ordered product quantization is proposed to constrain this sequence into an ordered representation. It can be applied with a multi-stream delayed LM to achieve better autoregressive generation along both time and stream axes in TTS. The experimental result strongly demonstrates the effectiveness of the proposed approach, achieving superior performance over baseline systems even if compressing the frameshift of speech from 20ms to 240ms (12x). The ablation studies further validate the importance of learning the proposed ordered multi-stream semantic representation in pursuing shorter speech sequences for efficient LM-based TTS.

语音合成多流编码语言模型高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。