arXiv:2410.05620eess.AS2024-10

用歌声转换模型提升语音合成中的跨说话人风格迁移效果

Improving Data Augmentation-based Cross-Speaker Style Transfer for TTS with Singing Voice, Style Filtering, and F0 Matching

  • 用预训练歌声转换模型将源说话人表达性语音转为目标说话人声音
  • 通过音高匹配减少音色差异带来的声调偏差,提升语音自然度
  • 用风格分类器筛选最富有表现力的音频,仅需几分钟中性数据即可达成顶尖效果

语音合成中的跨说话人风格迁移目标是将源说话人的表达性语音风格迁移到仅有中性语音的目标说话人。本文提出使用预训练的歌声转换(SVC)模型,将源说话人的表达性语音转换为目标说话人声音。在转换过程中,采用基频(F0)匹配技术以缓解音色差异显著的说话人之间的声调偏差。同时,引入风格分类器过滤出最具表现力的输出音频用于语音合成训练。实验表明,该方法仅需目标说话人几分钟的中性语音数据即可达到当前最优水平,而其他方法通常需要数小时。感知评估显示,使用SVC和风格过滤器后,具有较强发声努力的风格在自然度和风格强度上均有提升;同时,提出的F0匹配算法显著增强了说话人相似性。

原文摘要 · Abstract (English)

The goal of cross-speaker style transfer in TTS is to transfer a speech style from a source speaker with expressive data to a target speaker with only neutral data. In this context, we propose using a pre-trained singing voice conversion (SVC) model to convert the expressive data into the target speaker's voice. In the conversion process, we apply a fundamental frequency (F0) matching technique to mitigate tonal variances between speakers with significant timbral differences. A style classifier filter is proposed to select the most expressive output audios for the TTS training. Our approach is comparable to state-of-the-art with only a few minutes of neutral data from the target speaker, while other methods require hours. A perceptual assessment showed improvements brought by the SVC and the style filter in naturalness and style intensity for the styles that display more vocal effort. Also, increased speaker similarity is obtained with the proposed F0 matching algorithm.

语音合成风格迁移歌声转换音高匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。