arXiv:2607.13278cs.SD2026-07中稿 · International Conf…

用音乐生成模型做语音和歌唱转换,效果媲美专用系统。

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion

论文配图:Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
图 1 · 摘自论文原文
  • 将音乐扩散模型扩展至语音/歌唱转换,用音高和发音图作为条件。
  • 自然度和演唱者相似性超越专用系统,且保持精准音高控制。
  • 无需人工标注,可用现成提取器实现大规模自监督训练。

近期基于扩散的生成模型在语音、歌唱和乐器音乐合成等特定音频任务中表现优异,但通常专用于单一领域,难以泛化到混合或中间类型。本文将原本用于多乐器音乐合成的扩散模型,适配至语音与歌唱转换任务,构建统一框架。具体地,将基于乐谱的条件扩展为包含发音后验图(PPGs)和音高轮廓,并将音色条件重新诠释为通过特征线性调制表达说话人或演唱者身份。实验表明,该模型在自然度和演唱者相似性上达到甚至超过专用语音转换系统,同时在语音与歌唱中均保持精确音高控制。然而,引入乐器训练数据时,发音准确性下降,嗓音质量有所降低。此外,我们证明了即插即用的特征提取器可提供有效条件信号,支持无须人工标注的大规模自监督训练。结果表明,跨领域模型迁移具有潜力,可构建统一处理语音、歌唱与音乐的生成系统。定性样本见项目页:https://benadar293.github.io/voice-conversion

原文摘要 · Abstract (English)

Recent diffusion-based generative models have achieved strong results in domain-specific audio generation tasks such as speech, singing, and instrumental music synthesis. However, these models are typically specialized and do not generalize well to mixed or intermediate audio types. In this work, we adapt a diffusion-based model originally designed for multi-instrument music synthesis to voice conversion, covering both speech and singing within a unified framework. Specifically, we extend musical note-based conditioning to include phonetic posteriorgrams (PPGs) and pitch contours, and reinterpret timbre conditioning as speaker or singer identity via feature-wise linear modulation. Experiments show that the adapted model matches or surpasses a dedicated voice conversion system in terms of naturalness and performer similarity, while maintaining accurate pitch control across speech and singing. At the same time, we observe limitations in phonetic fidelity and a degradation in vocal quality when incorporating instrumental training data. Furthermore, we demonstrate that off-the-shelf feature extractors provide effective conditioning signals, enabling large-scale self-supervised training without manual annotations. These results highlight the potential of cross-domain model transfer towards unified audio generation systems capable of handling speech, singing, and music. Qualitative samples can be found on our project page: https://benadar293.github.io/voice-conversion

语音转换扩散模型跨域迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。