arXiv:2411.19770cs.SDcs.CL2024-11中稿 · APSIPA ASC 2025被引 1

让语音转换在嘈杂环境下仍能保持高精度。

Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning

  • 双分支编码器+抗噪对比损失,提升噪声下语音转换效果
  • 在干净与嘈杂场景中均优于基线系统
  • 可挖掘语音转换模型的隐含说话人表征能力

真实场景中,参考语音常因网络来源带有背景噪声,导致单次语音转换(One-shot VC)性能下降。为此,本文提出噪声鲁棒的单次语音转换系统Noro,包含专为噪声参考语音设计的双分支编码模块和抗噪对比损失函数。实验表明,Noro在干净与嘈杂环境下均优于基线系统,验证了其在实际应用中的有效性。此外,通过将基线系统的参考编码器重用于说话人编码,发现其在SUPERB基准下表现媲美多个先进的自监督学习模型,揭示了单次语音转换任务在推动说话人表征学习方面的潜力。

原文摘要 · Abstract (English)

The effectiveness of one-shot voice conversion (VC) decreases in real-world scenarios where reference speeches, which are often sourced from the internet, contain various disturbances like background noise. To address this issue, we introduce Noro, a noise-robust one-shot VC system. Noro features innovative components tailored for VC using noisy reference speeches, including a dual-branch reference encoding module and a noise-agnostic contrastive speaker loss. Experimental results demonstrate that Noro outperforms our baseline system in both clean and noisy scenarios, highlighting its efficacy for real-world applications. Additionally, we investigate the hidden speaker representation capabilities of our baseline system by repurposing its reference encoder as a speaker encoder. The results show that it is competitive with several advanced self-supervised learning models for speaker representation under the SUPERB settings, highlighting the potential for advancing speaker representation learning through one-shot VC tasks.

语音转换噪声鲁棒说话人表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。