arXiv:2502.12572cs.SD2025-02AAAI被引 33

TechSinger实现多语言歌声合成中七种发声技巧的精准控制。

TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow Matching

  • 基于流匹配生成模型,实现发声技巧的可调控生成。
  • 支持五种语言、七种发声技巧,语音质量与控制精度均优于现有方法。
  • 用户可用自然语言指定声音特征,适合音乐创作与个性化音色设计者。

歌声合成技术在生成自然高质语音方面取得了显著进展,但现有方法很少能对音量强度、混声、假声、气泡音和气息音等发声技巧进行精确控制,限制了合成语音的表现力。本文提出TechSinger,一种支持五种语言和七种发声技巧的可控歌声合成系统。该系统采用基于流匹配的生成模型,提升对多种发声技巧的表达控制能力。为增强训练数据多样性,我们构建了音素级发声技巧检测模型,自动标注数据集中的技巧标签。此外,通过提示式技巧预测模型,用户可使用自然语言指定期望的声音属性,实现细粒度控制。实验结果表明,TechSinger在音频质量与技巧特定控制方面均显著优于现有方法。音频样例见 https://gwx314.github.io/tech-singer/。

原文摘要 · Abstract (English)

Singing voice synthesis has made remarkable progress in generating natural and high-quality voices. However, existing methods rarely provide precise control over vocal techniques such as intensity, mixed voice, falsetto, bubble, and breathy tones, thus limiting the expressive potential of synthetic voices. We introduce TechSinger, an advanced system for controllable singing voice synthesis that supports five languages and seven vocal techniques. TechSinger leverages a flow-matching-based generative model to produce singing voices with enhanced expressive control over various techniques. To enhance the diversity of training data, we develop a technique detection model that automatically annotates datasets with phoneme-level technique labels. Additionally, our prompt-based technique prediction model enables users to specify desired vocal attributes through natural language, offering fine-grained control over the synthesized singing. Experimental results demonstrate that TechSinger significantly enhances the expressiveness and realism of synthetic singing voices, outperforming existing methods in terms of audio quality and technique-specific control. Audio samples can be found at https://gwx314.github.io/tech-singer/.

歌声合成流匹配多语言技巧控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。