arXiv:2508.05978cs.SDcs.AI2025-08中稿 · INTERSPEECH 2025被引 2

用目标音色特征替换源音色,提升音色相似度与音频质量。

DAFMSVC: One-Shot Singing Voice Conversion with Dual Attention Mechanism and Flow Matching

  • 用目标音色特征替代源音色特征,防止音色泄露。
  • 双注意力融合音色、旋律与语言信息,生成更自然音频。
  • 适合需要高质量单次音色转换的研究者与开发者。

歌唱语音转换(Singing Voice Conversion, SVC)旨在将源演唱者的音色转移到目标演唱者,同时保持旋律和歌词不变。任意对任意的SVC面临的一个关键挑战是:在不降低音质的前提下,将未见过的演唱者音色适配到源音频。现有方法要么存在音色泄露,要么难以在生成音频中实现理想的音色相似度与质量。为此,我们提出DAFMSVC,通过用最相似的目标音色自监督学习(SSL)特征替换源音频的SSL特征,避免音色泄露。同时引入双交叉注意力机制,自适应融合说话人嵌入、旋律与语言内容。此外,设计了流匹配模块,从融合特征中生成高质量音频。实验结果表明,DAFMSVC在主观与客观评估中均显著提升音色相似度与自然度,优于当前最优方法。

原文摘要 · Abstract (English)

Singing Voice Conversion (SVC) transfers a source singer's timbre to a target while keeping melody and lyrics. The key challenge in any-to-any SVC is adapting unseen speaker timbres to source audio without quality degradation. Existing methods either face timbre leakage or fail to achieve satisfactory timbre similarity and quality in the generated audio. To address these challenges, we propose DAFMSVC, where the self-supervised learning (SSL) features from the source audio are replaced with the most similar SSL features from the target audio to prevent timbre leakage. It also incorporates a dual cross-attention mechanism for the adaptive fusion of speaker embeddings, melody, and linguistic content. Additionally, we introduce a flow matching module for high quality audio generation from the fused features. Experimental results show that DAFMSVC significantly enhances timbre similarity and naturalness, outperforming state-of-the-art methods in both subjective and objective evaluations.

语音转换音色迁移流匹配注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。