5Hz超低帧率下实现高保真语音合成,速度提升3倍
U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
- 基于Transformer设计跨帧长程依赖模块,结合残差向量量化优化配置
- 在5Hz极低帧率下保持语音自然度与相似性,推理速度提升约3倍
- 适合追求高速生成的语音合成系统,尤其适配大语言模型驱动场景
我们提出U-Codec,一种超低帧率(5Hz)神经语音编解码器,在5帧/秒的极低帧率下实现高保真重建与快速语音生成。传统5Hz极端压缩常导致可懂度和频谱细节严重损失,为此我们引入基于Transformer的跨帧长程依赖模块,并系统探索残差向量量化(RVQ)深度与码本大小,确定最优配置。此外,我们将U-Codec应用于基于大语言模型(LLM)的自回归语音合成模型,利用多层分层架构有效捕捉跨层词元依赖关系。将原本在50Hz下的3层RVQ扩展至5Hz下的32层RVQ。实验表明,相比高帧率编解码器,U-Codec使LLM-based TTS推理速度提升约3倍,同时保持语音相似性与自然度。结果验证了使用高度压缩的5Hz离散词元进行快速高保真语音合成的可行性。
原文摘要 · Abstract (English)
We propose \textbf{U-Codec}, an \textbf{U}ltra low frame-rate neural speech \textbf{Codec} that achieves high-fidelity reconstruction and fast speech generation at an extremely low frame-rate of 5Hz (5 frames per second). Extreme compression at 5Hz typically leads to severe intelligibility and spectral detail loss, we introduce a Transformer-based inter-frame long-term dependency module and systematically explore residual vector quantization (RVQ) depth and codebook size to identify optimal configurations. Moreover, we apply U-Codec into a large language model (LLM)-based auto-regressive TTS model, which leverages global and local hierarchical architecture to effectively capture dependencies across multi-layer tokens. We extend LLM-based TTS from 3-layer RVQ at 50Hz to 32-layer RVQ at 5Hz. Experimental results demonstrate that U-Codec improves LLM-based TTS inference speed by around 3 $\times$ over high-frame-rate codecs while maintaining similarity and naturalness. These results validate the feasibility of using highly compressed 5Hz discrete tokens for fast and high-fidelity speech synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。