arXiv:2409.17452eess.AScs.CL2024-09被引 5

用跨语言描述控制语音,实现多语言自然可控合成

Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control

  • 用另一语言的描述数据训练控制模型,与目标语言TTS模型共享声纹与风格表征
  • 在英日双语上均实现高自然度与强可控性,日语无配对数据仍有效
  • 适合需要跨语言语音定制的场景,如多语种虚拟主播、个性化语音助手

我们提出一种新型基于描述的可控文本转语音方法,具备跨语言语音控制能力。为解决目标语言缺乏音视频描述配对数据的问题,将目标语言训练的TTS模型与另一语言训练的描述控制模型结合,后者将输入的文本描述映射为TTS模型的条件特征。两个模型通过自监督学习(SSL)共享解耦的音色和风格表征,实现如保留原音色的同时控制说话风格等解耦语音控制。由于基于SSL的音色与风格表征具有语言无关性,共享嵌入空间使跨语言语音特性控制得以实现。在英语和日语上的实验表明,即使未使用任何日语音视频描述配对数据,该方法仍能在两种语言上实现高自然度与强可控性。

原文摘要 · Abstract (English)

We propose a novel description-based controllable text-to-speech (TTS) method with cross-lingual control capability. To address the lack of audio-description paired data in the target language, we combine a TTS model trained on the target language with a description control model trained on another language, which maps input text descriptions to the conditional features of the TTS model. These two models share disentangled timbre and style representations based on self-supervised learning (SSL), allowing for disentangled voice control, such as controlling speaking styles while retaining the original timbre. Furthermore, because the SSL-based timbre and style representations are language-agnostic, combining the TTS and description control models while sharing the same embedding space effectively enables cross-lingual control of voice characteristics. Experiments on English and Japanese TTS demonstrate that our method achieves high naturalness and controllability for both languages, even though no Japanese audio-description pairs are used.

语音合成跨语言可控生成描述控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。