arXiv:2510.20677cs.SDcs.AI2025-10被引 2

提升歌声转换在真实噪声环境下的鲁棒性与表现力

R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion

  • 通过模拟噪声和音高扰动增强模型抗干扰能力
  • 融合多源歌声数据提升音色保留与风格表现
  • 引入神经声源-滤波器模型改善音质自然度

在真实场景的歌声转换应用中,环境噪声和对表现力输出的需求构成重大挑战。传统方法通常基于纯净数据训练与推理,与实际部署场景不符,难以应对音乐分离带来的各类噪声与失真。为此,我们提出R2-SVC框架:首先,通过随机基频(F₀)扰动及混响、回声等音乐分离失真模拟,显著提升噪声条件下的性能;其次,利用领域特定的歌声数据——包括干净人声、DNSMOS过滤后的分离人声以及公开歌唱语料库——丰富说话人表征,实现音色保持与演唱风格细节的捕捉;第三,集成神经声源-滤波器(NSF)模型,显式建模谐波与噪声成分,提升转换后歌声的自然度与可控性。R2-SVC在多个基准测试中,无论在干净还是噪声条件下均达到当前最优效果。

原文摘要 · Abstract (English)

In real-world singing voice conversion (SVC) applications, environmental noise and the demand for expressive output pose significant challenges. Conventional methods, however, are typically designed without accounting for real deployment scenarios, as both training and inference usually rely on clean data. This mismatch hinders practical use, given the inevitable presence of diverse noise sources and artifacts from music separation. To tackle these issues, we propose R2-SVC, a robust and expressive SVC framework. First, we introduce simulation-based robustness enhancement through random fundamental frequency ($F_0$) perturbations and music separation artifact simulations (e.g., reverberation, echo), substantially improving performance under noisy conditions. Second, we enrich speaker representation using domain-specific singing data: alongside clean vocals, we incorporate DNSMOS-filtered separated vocals and public singing corpora, enabling the model to preserve speaker timbre while capturing singing style nuances. Third, we integrate the Neural Source-Filter (NSF) model to explicitly represent harmonic and noise components, enhancing the naturalness and controllability of converted singing. R2-SVC achieves state-of-the-art results on multiple SVC benchmarks under both clean and noisy conditions.

歌声转换鲁棒性零样本音质增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。