arXiv:2506.21478cs.SDcs.AI2025-06被引 1

用扩散模型直接优化歌声,生成更自然的音高和发音。

SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture

  • 用参考音频引导去噪,统一生成高质量歌声
  • 在Opencpop数据集上主观评分领先,减少失真
  • 适合需要高保真歌声合成的研究与应用

歌声合成(SVS)旨在从乐谱生成富有表现力且高质量的演唱音频,需精确建模音高、时长与发音。尽管扩散模型在图像和视频生成中取得显著进展,但其在歌声合成中的应用仍受制于歌唱复杂的声学与音乐特征,常导致失真影响自然度。本文提出SmoothSinger,一种条件扩散模型,可直接在统一框架内优化低质量合成音频,避免传统两阶段流程中声码器引入的失真。模型采用参考引导的双分支架构,以任意基线系统生成的低质量音频为参考,指导去噪过程,提升表达力与上下文感知能力;同时在传统U-Net基础上增加并行低频上采样路径,更好捕捉音高轮廓与长期频谱依赖。为改善训练对齐,将参考音频替换为退化的真值音频,缓解参考与目标信号间的时序错位。在大规模中文歌声语料库Opencpop上的实验表明,SmoothSinger在客观与主观评估中均达到最先进水平,消融实验验证其有效降低失真、提升自然度。

原文摘要 · Abstract (English)

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image and video generation, their application to SVS remains challenging due to the complex acoustic and musical characteristics of singing, often resulting in artifacts that degrade naturalness. In this work, we propose SmoothSinger, a conditional diffusion model designed to synthesize high quality and natural singing voices. Unlike prior methods that depend on vocoders as a final stage and often introduce distortion, SmoothSinger refines low-quality synthesized audio directly in a unified framework, mitigating the degradation associated with two-stage pipelines. The model adopts a reference-guided dual-branch architecture, using low-quality audio from any baseline system as a reference to guide the denoising process, enabling more expressive and context-aware synthesis. Furthermore, it enhances the conventional U-Net with a parallel low-frequency upsampling path, allowing the model to better capture pitch contours and long term spectral dependencies. To improve alignment during training, we replace reference audio with degraded ground truth audio, addressing temporal mismatch between reference and target signals. Experiments on the Opencpop dataset, a large-scale Chinese singing corpus, demonstrate that SmoothSinger achieves state-of-the-art results in both objective and subjective evaluations. Extensive ablation studies confirm its effectiveness in reducing artifacts and improving the naturalness of synthesized voices.

歌声合成扩散模型语音生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。