arXiv:2508.04996eess.AS2025-08被引 5

提出一种抗噪且表达丰富的零样本语音转换方法,兼顾速度与音色保真。

REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers

  • 用随机遮蔽缓解自监督特征冗余,提升抗噪能力与表达力。
  • 通过隐式对齐减少非必要特征重建,抑制音色泄漏。
  • 引入捷径模型将推理步数降至4步,支持语音与歌声统一建模。

真实场景中,源语音的环境噪声及用户对表达力的需求构成关键挑战。传统ASR方法虽具抗噪性但压制韵律丰富性,而基于自监督学习(SSL)的模型提升表达力却存在音色泄漏和噪声敏感问题。本文提出REF-VC,一种抗噪、表达丰富且快速的零样本语音转换系统。核心创新包括:(1) 采用随机遮蔽策略缓解SSL特征固有的信息冗余,增强抗噪性与表达力;(2) 受E2TTS启发,引入隐式对齐机制以抑制非必要特征重建;(3) 融合捷径模型加速流匹配推理,显著降低至4步。实验表明,REF-VC在噪声数据集上优于基线如Seed-VC,在干净数据集上表现相当。此外,该模型可兼容歌唱语音转换任务。

原文摘要 · Abstract (English)

In real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ensure noise robustness but suppress prosody richness, while SSL-based models improve expressiveness but suffer from timbre leakage and noise sensitivity. This paper proposes REF-VC, a noise-robust expressive voice conversion system. Key innovations include: (1) A random erasing strategy to mitigate the information redundancy inherent in SSL features, enhancing noise robustness and expressiveness; (2) Implicit alignment inspired by E2TTS to suppress non-essential feature reconstruction; (3) Integration of Shortcut Models to accelerate flow matching inference, significantly reducing to 4 steps. Experimental results demonstrate that REF-VC outperforms baselines such as Seed-VC in zero-shot scenarios on the noisy set, while also performing comparably to Seed-VC on the clean set. In addition, REF-VC can be compatible with singing voice conversion within one model.

语音转换扩散模型抗噪快速生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。