arXiv:2606.21888eess.AScs.SD2026-06

让语音编码更懂语调,提升语音转换自然度。

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion

  • 将语调建模为条件残差,分离内容与语调特征。
  • 用低频梅尔频带和同声者配对训练,减少音色泄露。
  • 适合需要保留原语音调的语音转换任务。

神经语音编码器能高效压缩语音,已成为语音生成的基础,但通常作为整体表示学习,将语言内容、说话人身份和语调混在一起。这种设计虽适用于零样本语音克隆,却阻碍了需保留或迁移语调的下游任务,如语音转换。为此,我们提出ProsoCodec,一种面向语调的语音编码器,将语调建模为条件残差而非独立流。具体而言,通过在编码器和解码器中以文本和说话人嵌入作为前缀标记进行条件化,离散瓶颈被鼓励捕捉未被内容和说话人解释的语调变化。为进一步保留语调,我们采用低频梅尔频带,并在同声者配对语句上训练模型。语音转换实验表明,该方法显著提升了语调保留效果并减少了源音色泄露。

原文摘要 · Abstract (English)

Neural speech codecs efficiently compress speech and have become a foundation for speech generation, but they are typically learned as holistic representations that intertwine linguistic content, speaker identity, and prosody. While this design is effective for zero-shot voice cloning, it hinders downstream tasks that require prosody preservation or transfer, such as voice conversion. To address this, we introduce ProsoCodec, a prosody-oriented speech codec that models prosody as a conditional residual rather than as a disentangled stream. Specifically, by conditioning both the encoder and decoder on text and speaker embeddings as prefix tokens, the discrete bottleneck is encouraged to capture prosodic variation not explained by content and speaker. To further preserve prosody, we use the low-frequency mel band and train the model on paired same-speaker utterances. Experiments on voice conversion show improved prosody preservation and reduced source-timbre leakage.

语音转换语音编码语调建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。