arXiv:2412.01053cs.SDeess.AS2024-12中稿 · Interspeech 2025被引 15

FreeCodec通过解耦语音特性,用更少的编码符实现更优重建效果。

FreeCodec: A disentangled neural speech codec with fewer tokens

  • 将语音分解为音色、语调和内容三部分独立建模
  • 在低码率下仍保持高重建质量,优于现有方法
  • 适合语音生成与编码场景,尤其关注高效性

神经语音编解码器因其使用离散编码符实现优异重建效果而备受关注,是语音编码和大语言模型等生成任务的关键组件。然而,基于残差向量量化的多数方法在编码符数量较少时性能下降,因难以有效建模复杂耦合信息。本文提出一种名为FreeCodec的新型神经语音编解码器,通过将语音内在属性分解为不同组件:1)提取全局向量作为音色信息;2)采用长步长结构的语调编码器建模语调信息;3)内容信息由内容编码器获取。结合差异化训练策略,FreeCodec在重建与解耦场景中均达到当前最优表现。主观与客观实验结果均表明其显著优于现有方法。

原文摘要 · Abstract (English)

Neural speech codecs have gained great attention for their outstanding reconstruction with discrete token representations. It is a crucial component in generative tasks such as speech coding and large language models (LLM). However, most works based on residual vector quantization perform worse with fewer tokens due to low coding efficiency for modeling complex coupled information. In this paper, we propose a neural speech codec named FreeCodec which employs a more effective encoding framework by decomposing intrinsic properties of speech into different components: 1) a global vector is extracted as the timbre information, 2) a prosody encoder with a long stride level is used to model the prosody information, 3) the content information is from a content encoder. Using different training strategies, FreeCodec achieves state-of-the-art performance in reconstruction and disentanglement scenarios. Results from subjective and objective experiments demonstrate that our framework outperforms existing methods.

语音编码解耦表示低码率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。