直接从乐谱生成歌声,无需预设时长或对齐音高。
VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition

- 用乐谱符号和歌词交织表示,端到端生成连续声学特征。
- 在2300小时数据上训练,自然度比最强基线高0.42分CMOS。
- 适合音乐创作场景,支持真实作曲流程中的灵活编排。
现有歌唱语音合成系统通常需要预定义时长、显式时长预测或时间对齐的声学引导,限制了其在实际作曲流程中的兼容性。我们提出VocalRender,一种直接从歌词、音高、符号化音符值和速度生成歌唱的乐谱原生系统。它采用交错的歌词-音符表示方式,结合自回归扩散模型生成连续声学潜变量,并同时预测输出长度,无需显式时长预测。在2300小时歌唱数据集上训练后,VocalRender在域内与域外基准测试中均展现出强可懂性、良好旋律控制力与高说话人相似度。尤其值得注意的是,其自然度在CMOS评分上比最强基线高出0.42分,验证了所提乐谱原生架构的有效性。
原文摘要 · Abstract (English)
Existing singing voice synthesis systems often require predefined durations, explicit duration prediction, or time-aligned acoustic guidance, which limits their compatibility with practical composition workflows. We propose VocalRender, a score-native system that directly synthesizes singing from lyrics, pitches, symbolic note values, and tempo. It uses an interleaved lyric--note representation and an autoregressive diffusion model to generate continuous acoustic latents while predicting the output length, eliminating the need for explicit duration prediction. Trained on a 2,300-hour singing dataset, VocalRender achieves strong intelligibility, strong melody control, and high speaker similarity across both in-domain and out-of-domain benchmarks. Notably, it outperforms the strongest baseline by $0.42$ points in naturalness CMOS, demonstrating the effectiveness of our proposed score-native architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。