arXiv:2505.14910eess.AScs.CL2025-05ACL被引 18

TCSinger 2实现多语言零样本歌声合成,支持多样提示风格控制。

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

  • 通过模糊边界编码器提升音素与音符间过渡平滑性。
  • 对比学习提取语音、文本等多模态提示的对齐表征。
  • 基于流模型与提示驱动的风格控制,生成质量更优。

可定制的多语言零样本歌声合成(SVS)在音乐创作与短视频配音中具有广泛应用前景。然而,现有模型过度依赖音素与音符边界标注,导致零样本场景下鲁棒性差,音素与音符间转换不自然,且缺乏通过多样化提示实现多层级风格控制的能力。为此,我们提出TCSinger 2,一个基于多种提示进行风格迁移与控制的多任务多语言零样本SVS模型。其核心包含三个模块:1)模糊边界内容(BBC)编码器,预测时长、扩展内容嵌入并掩码边界,实现平滑过渡;2)自定义音频编码器,采用对比学习从歌声、语音和文本提示中提取对齐表征;3)基于流的自定义变换器,结合Cus-MOE与基频(F0)监督,提升合成质量与风格建模能力。实验表明,TCSinger 2在主观与客观指标上均优于基线模型,在多个任务中表现优异。歌声样例可在 https://aaronz345.github.io/TCSinger2Demo/ 查看。

原文摘要 · Abstract (English)

Customizable multilingual zero-shot singing voice synthesis (SVS) has various potential applications in music composition and short video dubbing. However, existing SVS models overly depend on phoneme and note boundary annotations, limiting their robustness in zero-shot scenarios and producing poor transitions between phonemes and notes. Moreover, they also lack effective multi-level style control via diverse prompts. To overcome these challenges, we introduce TCSinger 2, a multi-task multilingual zero-shot SVS model with style transfer and style control based on various prompts. TCSinger 2 mainly includes three key modules: 1) Blurred Boundary Content (BBC) Encoder, predicts duration, extends content embedding, and applies masking to the boundaries to enable smooth transitions. 2) Custom Audio Encoder, uses contrastive learning to extract aligned representations from singing, speech, and textual prompts. 3) Flow-based Custom Transformer, leverages Cus-MOE, with F0 supervision, enhancing both the synthesis quality and style modeling of the generated singing voice. Experimental results show that TCSinger 2 outperforms baseline models in both subjective and objective metrics across multiple related tasks. Singing voice samples are available at https://aaronz345.github.io/TCSinger2Demo/.

歌声合成零样本多语言风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。