用自监督特征提升伴奏下歌声转换的旋律保真度
Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features
- 用自监督学习提取旋律特征,抗伴奏干扰更强
- 在有无伴奏环境下均显著提升音高准确率和自然度
- 无需分离声源,避免伪影,适合实际应用
旋律保真是歌声转换(SVC)的关键。但在多数场景中,音频常伴有背景音乐(BGM),导致音高提取困难,严重降低转换质量。现有方法或依赖鲁棒性有限的神经网络旋律提取器,或采用源分离预处理,但后者易引入伪影且操作成本高。为此,本文提出一种基于自监督表示的新型SVC方法,利用自监督学习(SSL)模型提取旋律特征,首次系统评估不同SSL模型在旋律提取中的有效性。实验表明,该方法在含复杂伴奏与纯净音频环境下,均显著优于基线模型,在主观与客观评价中均实现更高的音高准确性、相似度与自然度。
原文摘要 · Abstract (English)
Melody preservation is crucial in singing voice conversion (SVC). However, in many scenarios, audio is often accompanied with background music (BGM), which can cause audio distortion and interfere with the extraction of melody and other key features, significantly degrading SVC performance. Previous methods have attempted to address this by using more robust neural network-based melody extractors, but their performance drops sharply in the presence of complex accompaniment. Other approaches involve performing source separation before conversion, but this often introduces noticeable artifacts, leading to a significant drop in conversion quality and increasing the user's operational costs. To address these issues, we introduce a novel SVC method that uses self-supervised representation-based melody features to improve melody modeling accuracy in the presence of BGM. In our experiments, we compare the effectiveness of different self-supervised learning (SSL) models for melody extraction and explore for the first time how SSL benefits the task of melody extraction. The experimental results demonstrate that our proposed SVC model significantly outperforms existing baseline methods in terms of melody accuracy and shows higher similarity and naturalness in both subjective and objective evaluations across noisy and clean audio environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。