arXiv:2501.13870cs.SDeess.AS2025-01被引 9

用语音参考实现零样本歌声合成与转换,支持多维度控制。

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference

  • 统一框架整合歌声合成与转换,基于语音参考实现零样本学习。
  • 在音色相似度和音乐性上超越现有方法,提升显著。
  • 适合需要快速克隆歌声或跨风格迁移的创作者使用。

我们提出一种统一的歌声合成(SVS)与转换(SVC)框架,解决现有方法在跨领域任务中表现不佳、输出音乐性差及歌唱数据稀缺的问题。该框架可控制语言内容(基于歌词)、表演特征(基于乐谱)、演唱风格与发声技巧(基于选择器)以及声音身份(基于语音样本)。采用零样本学习范式,包含一个SVS模型和两个SVC模型,利用预训练内容嵌入与基于扩散的生成器。框架在混合歌唱与语音音频的数据集上训练,实现仅需语音参考即可进行歌声克隆。实验表明,在音色相似度和音乐性方面显著优于当前最优基线,为其他低数据量音乐任务(如乐器风格迁移)提供新思路。示例见:everyone-can-sing.github.io。

原文摘要 · Abstract (English)

We propose a unified framework for Singing Voice Synthesis (SVS) and Conversion (SVC), addressing the limitations of existing approaches in cross-domain SVS/SVC, poor output musicality, and scarcity of singing data. Our framework enables control over multiple aspects, including language content based on lyrics, performance attributes based on a musical score, singing style and vocal techniques based on a selector, and voice identity based on a speech sample. The proposed zero-shot learning paradigm consists of one SVS model and two SVC models, utilizing pre-trained content embeddings and a diffusion-based generator. The proposed framework is also trained on mixed datasets comprising both singing and speech audio, allowing singing voice cloning based on speech reference. Experiments show substantial improvements in timbre similarity and musicality over state-of-the-art baselines, providing insights into other low-data music tasks such as instrumental style transfer. Examples can be found at: everyone-can-sing.github.io.

歌声合成语音克隆零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。