arXiv:2410.15764eess.AScs.AI2024-10中稿 · Interspeech 2025被引 27

低比特率下实现语音与说话人解耦的离散语音编解码器

LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec

  • 多阶段无监督训练+说话人扰动,构建离散解耦表征空间
  • 仅用单个码本,词汇量更小,音质和可懂度优于基线
  • 适合语音合成、语音转换等需解耦说话人的任务

尽管离散语音标记在基于语言模型的语音生成中展现出巨大潜力,但其高比特率和冗余的音色信息限制了模型发展。本文提出LSCodec,一种兼具低比特率和说话人解耦能力的离散语音编解码器。该方法采用多阶段无监督训练框架,并引入说话人扰动技术:首先建立连续信息瓶颈,再通过向量量化生成离散的说话人解耦空间,最后由离散标记声码器精细还原语音细节。重建评估显示,LSCodec仅使用单个码本和较小词汇量,即在可懂度和音频质量上优于基线。语音转换与说话人探测实验验证了其优异的说话人解耦性能,消融实验证明了所提训练框架的有效性。

原文摘要 · Abstract (English)

Although discrete speech tokens have exhibited strong potential for language model-based speech generation, their high bitrates and redundant timbre information restrict the development of such models. In this work, we propose LSCodec, a discrete speech codec that has both low bitrate and speaker decoupling ability. LSCodec adopts a multi-stage unsupervised training framework with a speaker perturbation technique. A continuous information bottleneck is first established, followed by vector quantization that produces a discrete speaker-decoupled space. A discrete token vocoder finally refines acoustic details from LSCodec. By reconstruction evaluations, LSCodec demonstrates superior intelligibility and audio quality with only a single codebook and smaller vocabulary size than baselines. Voice conversion and speaker probing experiments prove the excellent speaker disentanglement of LSCodec, and ablation study verifies the effectiveness of the proposed training framework.

语音编码离散表示说话人解耦低比特率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。