arXiv:2606.07015cs.SDcs.AI2026-06

统一歌曲生成与歌声转换,实现伴奏协同生成。

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation

论文配图:Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation
图 1 · 摘自论文原文
  • 构建统一声线空间,实现跨任务音色精细控制。
  • 采用课程学习策略,解决多任务优化冲突问题。
  • 支持零样本歌手克隆,适合音乐创作与智能制谱。

尽管歌曲生成和歌声转换(SVC)已取得显著进展,但长期处于孤立发展状态:前者缺乏零样本歌手克隆能力,后者忽视人声与伴奏的协同关系。为此,我们提出UniSinger,首个端到端统一框架,融合歌手克隆歌曲生成与伴奏协同生成的歌声转换。基于多模态扩散变换器,构建统一的说话人嵌入空间,将SVC中的说话人表征迁移至歌曲生成,实现细粒度跨任务音色控制。为缓解多任务优化冲突,设计任务特定模态掩码的课程学习策略,引导模型逐步掌握语义内容、人声音色与伴奏间的生成机制。实验表明,在两项任务上均达到当前最优性能,且实现互补增益,为智能音乐生产提供新可能。

原文摘要 · Abstract (English)

While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-shot speaker cloning, while the latter overlooks vocal-accompaniment synergy. To bridge this gap, we propose UniSinger, the first end-to-end framework unifying speaker cloning song generation and accompaniment co-generation SVC. Building on the multimodal diffusion transformer, we construct a unified speaker embedding space transferring speaker representation from SVC to song generation, endowing fine-grained cross-task timbre control. To mitigate multi-task optimization conflicts, we design a curriculum learning strategy using task-specific modality masking to guide the model to gradually master the generative mechanisms among semantic content, vocal timbre, and accompaniment. Experiments show state-of-the-art performance on both tasks and realizes complementary benefits, offering new possibilities for intelligent music production.

歌曲生成歌声转换伴奏生成多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。