无需训练即可合成新歌手声音,还能精细控制演唱风格。
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
- 用聚类向量量化压缩风格信息,构建紧凑表征空间。
- 联合预测风格与发音时长,提升生成质量与控制精度。
- 通过自适应归一化增强音色细节,支持跨语言风格迁移。
零样本歌唱语音合成(SVS)结合风格迁移与控制,旨在仅凭音频和文本提示,生成具有未见音色与风格(包括演唱方式、情绪、节奏、技巧和发音)的高质量歌声。然而,歌唱风格的多维度特性给建模、迁移与控制带来巨大挑战。现有模型常无法为未见过的歌手生成富有风格细节的歌声。为此,我们提出TCSinger,首个支持跨语言语音与歌唱风格迁移的零样本SVS模型,并具备多层级风格控制能力。TCSinger包含三个核心模块:1)聚类风格编码器采用聚类向量量化模型,稳定地将风格信息压缩至紧凑潜在空间;2)风格与持续时间语言模型(S&D-LM)同时预测风格信息与音素持续时间,相互促进;3)风格自适应解码器采用新型梅尔风格自适应归一化方法,生成更具细节的歌声。实验表明,TCSinger在合成质量、歌手相似度与风格可控性方面优于所有基线模型,在零样本风格迁移、多层级风格控制、跨语言风格迁移及语音到歌唱风格迁移等任务中均表现优异。歌声样例可访问 https://aaronz345.github.io/TCSingerDemo/。
原文摘要 · Abstract (English)
Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunciation) from audio and text prompts. However, the multifaceted nature of singing styles poses a significant challenge for effective modeling, transfer, and control. Furthermore, current SVS models often fail to generate singing voices rich in stylistic nuances for unseen singers. To address these challenges, we introduce TCSinger, the first zero-shot SVS model for style transfer across cross-lingual speech and singing styles, along with multi-level style control. Specifically, TCSinger proposes three primary modules: 1) the clustering style encoder employs a clustering vector quantization model to stably condense style information into a compact latent space; 2) the Style and Duration Language Model (S\&D-LM) concurrently predicts style information and phoneme duration, which benefits both; 3) the style adaptive decoder uses a novel mel-style adaptive normalization method to generate singing voices with enhanced details. Experimental results show that TCSinger outperforms all baseline models in synthesis quality, singer similarity, and style controllability across various tasks, including zero-shot style transfer, multi-level style control, cross-lingual style transfer, and speech-to-singing style transfer. Singing voice samples can be accessed at https://aaronz345.github.io/TCSingerDemo/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。