arXiv:2512.04793cs.SDcs.AI2025-12被引 3

提升真实歌曲中零样本歌声转换的鲁棒性,兼顾音高与旋律保真。

YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases

  • 融合预训练、微调与强化学习,构建端到端鲁棒框架
  • 在伴奏和声干扰下仍保持音色相似度与语音可懂度领先
  • 专为歌唱特性设计归纳偏置,适合实际音乐应用

歌声转换(SVC)旨在保留旋律与歌词的同时,呈现目标歌手的音色。然而,现有零样本SVC系统在真实歌曲中仍易受和声干扰、基频(F0)误差及缺乏歌唱特有归纳偏置的影响。本文提出YingMusic-SVC,一个统一连续预训练、鲁棒监督微调与Flow-GRPO强化学习的零样本框架。模型引入歌唱训练的RVC音色转换器实现音色-内容解耦,设计F0感知音色适配器以支持动态演唱表现,并采用能量平衡的修正流匹配损失提升高频保真度。在分级多轨基准测试中,YingMusic-SVC在音色相似度、可懂度与感知自然度上持续优于强开源基线,尤其在伴奏与和声污染条件下表现突出,验证了其在真实场景部署中的有效性。

原文摘要 · Abstract (English)

Singing voice conversion (SVC) aims to render the target singer's timbre while preserving melody and lyrics. However, existing zero-shot SVC systems remain fragile in real songs due to harmony interference, F0 errors, and the lack of inductive biases for singing. We propose YingMusic-SVC, a robust zero-shot framework that unifies continuous pre-training, robust supervised fine-tuning, and Flow-GRPO reinforcement learning. Our model introduces a singing-trained RVC timbre shifter for timbre-content disentanglement, an F0-aware timbre adaptor for dynamic vocal expression, and an energy-balanced rectified flow matching loss to enhance high-frequency fidelity. Experiments on a graded multi-track benchmark show that YingMusic-SVC achieves consistent improvements over strong open-source baselines in timbre similarity, intelligibility, and perceptual naturalness, especially under accompanied and harmony-contaminated conditions, demonstrating its effectiveness for real-world SVC deployment.

歌声转换零样本音频生成强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。