arXiv:2409.06237cs.SDeess.AS2024-09中稿 · ISCSLP 2024被引 4

用HuBERT提取旋律,对抗训练让歌声转换更抗噪、更自然

RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion

  • 用HuBERT替代传统特征提取,增强对噪声的鲁棒性
  • 三判别器对抗训练,减少自监督表示中的信息泄露
  • 在有无噪声情况下均优于基线,适合实际演唱场景

歌唱语音转换(SVC)因推理时音高和能量提取方法不鲁棒而受噪声影响。由于干净音频是SVC中源音频的关键,音乐源分离预处理可有效应对带背景音乐的演唱。然而,现有分离方法难以完全去除噪声或过度抑制信号成分,影响处理后音频的自然度和相似性。为此,本文提出RobustSVC,一种新型的任意到一的SVC框架,可将含噪人声转换为由目标歌手演唱的纯净人声。我们采用基于HuBERT的旋律提取器替代非鲁棒特征,并引入包含三个判别器的对抗训练机制,以减少自监督表示中的信息泄露。实验表明,RobustSVC在有噪和无噪条件下均表现出更强的抗噪能力,且在相似性和自然度上优于基线方法。

原文摘要 · Abstract (English)

Singing voice conversion (SVC) is hindered by noise sensitivity due to the use of non-robust methods for extracting pitch and energy during the inference. As clean signals are key for the source audio in SVC, music source separation preprocessing offers a viable solution for handling noisy audio, like singing with background music (BGM). However, current separating methods struggle to fully remove noise or excessively suppress signal components, affecting the naturalness and similarity of the processed audio. To tackle this, our study introduces RobustSVC, a novel any-to-one SVC framework that converts noisy vocals into clean vocals sung by the target singer. We replace the non-robust feature with a HuBERT-based melody extractor and use adversarial training mechanisms with three discriminators to reduce information leakage in self-supervised representations. Experimental results show that RobustSVC is noise-robust and achieves higher similarity and naturalness than baseline methods in both noisy and clean vocal conditions.

语音转换歌声合成鲁棒性HuBERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。