arXiv:2509.07038cs.SDcs.AI2025-09中稿 · ASRU 2025被引 1

通过音素级能量序列实现可控的歌声合成,提升动态表现力。

Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence

  • 用音素级能量序列显式控制歌声动态变化
  • 音素级输入下能量误差降低超50%
  • 无需复杂标注,适合音乐创作与个性化声线设计

可控歌声合成旨在生成反映用户意图的富有表现力的歌声。尽管现有系统音频质量高,但多依赖概率建模,难以精确控制如强弱变化等动态属性。本文聚焦动态控制——即时间上的响度变化,这是音乐表现力的关键。我们通过从真实频谱图中提取的能量序列,显式地条件化歌声合成模型,降低标注成本并提升可控性。同时提出音素级能量序列,便于用户操作。据我们所知,这是首个实现用户驱动的动态控制的歌声合成方法。实验表明,相比基线和能量预测模型,本方法在音素级输入下能量序列的均方绝对误差减少超过50%,且不牺牲合成质量。

原文摘要 · Abstract (English)

Controllable Singing Voice Synthesis (SVS) aims to generate expressive singing voices reflecting user intent. While recent SVS systems achieve high audio quality, most rely on probabilistic modeling, limiting precise control over attributes such as dynamics. We address this by focusing on dynamic control--temporal loudness variation essential for musical expressiveness--and explicitly condition the SVS model on energy sequences extracted from ground-truth spectrograms, reducing annotation costs and improving controllability. We also propose a phoneme-level energy sequence for user-friendly control. To the best of our knowledge, this is the first attempt enabling user-driven dynamics control in SVS. Experiments show our method achieves over 50% reduction in mean absolute error of energy sequences for phoneme-level inputs compared to baseline and energy-predictor models, without compromising synthesis quality.

歌声合成可控生成音素级控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。