arXiv:2512.04779cs.SDcs.AI2025-12被引 8

无需标注旋律和音素对齐,一键生成任意歌词的歌声

YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance

  • 用音频直接提取旋律,不依赖人工标注
  • 零样本生成效果优于现有方法,语音清晰度与旋律准确性双高
  • 适合音乐创作、自动化配乐等需快速生成歌声的场景

歌唱语音合成(SVS)因依赖精确的音素级对齐和手动标注的旋律轮廓,难以在实际中部署。为克服此问题,我们提出一种以旋律驱动的SVS框架,可基于任意参考旋律合成任意歌词,无需音素对齐。方法基于扩散变换器(DiT),引入专用旋律提取模块,直接从参考音频中解析旋律表征。通过教师模型引导优化旋律提取器,并设计隐式对齐机制,强化旋律分布一致性,提升旋律稳定性。同时,利用弱标注数据优化时长建模,采用多目标奖励函数的Flow-GRPO强化学习策略,联合提升发音清晰度与旋律保真度。实验表明,该模型在客观指标和主观听感测试中均优于现有方法,尤其在零样本与歌词适配场景表现突出,且无需人工标注即可保持高质量输出。代码与模型已开源。

原文摘要 · Abstract (English)

Singing Voice Synthesis (SVS) remains constrained in practical deployment due to its strong dependence on accurate phoneme-level alignment and manually annotated melody contours, requirements that are resource-intensive and hinder scalability. To overcome these limitations, we propose a melody-driven SVS framework capable of synthesizing arbitrary lyrics following any reference melody, without relying on phoneme-level alignment. Our method builds on a Diffusion Transformer (DiT) architecture, enhanced with a dedicated melody extraction module that derives melody representations directly from reference audio. To ensure robust melody encoding, we employ a teacher model to guide the optimization of the melody extractor, alongside an implicit alignment mechanism that enforces similarity distribution constraints for improved melodic stability and coherence. Additionally, we refine duration modeling using weakly annotated song data and introduce a Flow-GRPO reinforcement learning strategy with a multi-objective reward function to jointly enhance pronunciation clarity and melodic fidelity. Experiments show that our model achieves superior performance over existing approaches in both objective measures and subjective listening tests, especially in zero-shot and lyric adaptation settings, while maintaining high audio quality without manual annotation. This work offers a practical and scalable solution for advancing data-efficient singing voice synthesis. To support reproducibility, we release our inference code and model checkpoints.

歌声合成零样本扩散模型自动旋律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。