arXiv:2510.09016cs.SDcs.AI2025-10被引 3

用扩散模型生成500小时高质量中文歌声,无需逐音素标注

DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment

  • 用大模型生成歌词配固定旋律,构建小型训练集
  • 训练出可扩展的扩散变换器,合成超500小时高保真歌声
  • 隐式对齐机制避免音素时长标注,提升嘈杂数据鲁棒性

基于扩散模型的歌声合成虽表现力强,但受限于数据稀缺与模型扩展性。我们提出两阶段流程:通过固定旋律搭配大语言模型生成的歌词,构建小型真人演唱数据集;基于此训练旋律特定模型,合成超过500小时高质量中文歌声数据。在此基础上,提出DiTSinger,一种采用RoPE和qk-norm的扩散变换器,系统地在深度、宽度和分辨率上进行扩展以提升保真度。此外,设计隐式对齐机制,通过限制音素到声学特征的注意力范围在字符级跨度内,避免依赖音素级时长标签,从而增强在噪声或不确定对齐情况下的鲁棒性。大量实验验证该方法实现了可扩展、免对齐、高保真的歌声合成。

原文摘要 · Abstract (English)

Recent progress in diffusion-based Singing Voice Synthesis (SVS) demonstrates strong expressiveness but remains limited by data scarcity and model scalability. We introduce a two-stage pipeline: a compact seed set of human-sung recordings is constructed by pairing fixed melodies with diverse LLM-generated lyrics, and melody-specific models are trained to synthesize over 500 hours of high-quality Chinese singing data. Building on this corpus, we propose DiTSinger, a Diffusion Transformer with RoPE and qk-norm, systematically scaled in depth, width, and resolution for enhanced fidelity. Furthermore, we design an implicit alignment mechanism that obviates phoneme-level duration labels by constraining phoneme-to-acoustic attention within character-level spans, thereby improving robustness under noisy or uncertain alignments. Extensive experiments validate that our approach enables scalable, alignment-free, and high-fidelity SVS.

歌声合成扩散模型隐式对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。