arXiv:2603.24589eess.AScs.SD2026-03被引 1

用扩散模型实现可控歌声合成,改词不跑调,无需手动对齐。

YingMusic-Singer: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance

  • 全扩散模型架构,输入音色参考、旋律片段和改写歌词即可生成
  • 比现有最强基线Vevo2在旋律保持和歌词贴合上更优
  • 首个专用于旋律保真歌词修改的评测基准,适合音乐生成研究者

在保持旋律一致性的前提下,对歌词进行灵活修改并重生成歌声仍具挑战,现有方法或控制力不足,或需大量人工对齐。我们提出YingMusic-Singer,一种完全基于扩散模型的歌声合成方法,支持旋律可控且可灵活修改歌词。该模型接收三个输入:可选的音色参考、提供旋律的歌声片段和修改后的歌词,无需人工对齐。通过课程学习与组相对策略优化训练,YingMusic-Singer在旋律保持和歌词遵循方面优于目前唯一支持无对齐旋律控制的最可比基线Vevo2。我们还提出了首个用于旋律保真歌词修改评估的基准——LyricEditBench。代码、权重、基准及演示均已公开于https://github.com/ASLP-lab/YingMusic-Singer-Plus。

原文摘要 · Abstract (English)

Regenerating singing voices with altered lyrics while preserving melody consistency remains challenging, as existing methods either offer limited controllability or require laborious manual alignment. We propose YingMusic-Singer, a fully diffusion-based model enabling melody-controllable singing voice synthesis with flexible lyric manipulation. The model takes three inputs: an optional timbre reference, a melody-providing singing clip, and modified lyrics, without manual alignment. Trained with curriculum learning and Group Relative Policy Optimization, YingMusic-Singer achieves stronger melody preservation and lyric adherence than Vevo2, the most comparable baseline supporting melody control without manual alignment. We also introduce LyricEditBench, the first benchmark for melody-preserving lyric modification evaluation. The code, weights, benchmark, and demos are publicly available at https://github.com/ASLP-lab/YingMusic-Singer-Plus.

歌声合成扩散模型歌词修改旋律控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。