arXiv:2409.17364eess.AScs.SD2024-09中稿 · SynData4GenAI 2024被引 2

用语音转换生成的合成数据提升跨说话人语音合成的自然度和相似度。

Exploring synthetic data for cross-speaker style transfer in style representation based TTS

  • 用语音转换模型生成合成语音,辅助跨说话人风格迁移。
  • 合成数据使语音自然度提升,说话人相似度提高23.6%。
  • 该方法可扩展至跨语言场景,增强口音迁移能力。

在低资源表达性语音数据场景下,将跨说话人风格迁移引入文本到语音(TTS)模型面临挑战,因需分离音频中的说话人与风格信息。语音转换(VC)模型可为目标说话人生成表达性语音,用于训练TTS模型,但其生成质量与风格迁移能力直接影响整体性能。本文探索使用VC模型生成的合成数据辅助TTS模型完成跨说话人风格迁移任务。同时,采用音色扰动与原型角损失对风格编码器进行预训练,以缓解说话人信息泄露。实验表明,使用合成数据可显著提升跨说话人场景下的语音自然度与说话人相似度。进一步拓展至跨语言场景,有效增强口音迁移能力。

原文摘要 · Abstract (English)

Incorporating cross-speaker style transfer in text-to-speech (TTS) models is challenging due to the need to disentangle speaker and style information in audio. In low-resource expressive data scenarios, voice conversion (VC) can generate expressive speech for target speakers, which can then be used to train the TTS model. However, the quality and style transfer ability of the VC model are crucial for the overall TTS model quality. In this work, we explore the use of synthetic data generated by a VC model to assist the TTS model in cross-speaker style transfer tasks. Additionally, we employ pre-training of the style encoder using timbre perturbation and prototypical angular loss to mitigate speaker leakage. Our results show that using VC synthetic data can improve the naturalness and speaker similarity of TTS in cross-speaker scenarios. Furthermore, we extend this approach to a cross-language scenario, enhancing accent transfer.

语音合成风格迁移合成数据跨说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。