arXiv:2501.05976eess.AS2025-01中稿 · publication at the…被引 5

仅用5分钟语音+噪声增强,实现高质量低资源语音合成

Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron

  • 用噪声增强低资源语音数据,结合多采样训练策略
  • 仅需4个高资源说话人,5分钟目标语音即达高保真度
  • 适合新说话人、新语言的语音合成,降低数据门槛

近年来,多项文本到语音系统被提出,用于零样本、少样本和低资源场景下的自然语音合成。然而,这些方法通常需要来自多个说话人的训练数据,导致说话人之间语音质量差异大,限制了低资源说话人可达到的上限。本文提出一种新方法,仅需5分钟目标说话人语音,通过噪声增强低资源数据,并在训练中采用多种采样技术,实现高质量语音合成。该方法仅需4个高质量、高资源说话人数据,易于获取且实用。相比最先进的零样本方法HierSpeech++和近期的低资源方法AdapterMix,本方法在保持相近自然度的同时,显著提升了说话人相似度。该方法还可减少新说话人和新语言语音合成的数据需求。

原文摘要 · Abstract (English)

In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically require training with data from many different speakers. The speech quality across the speaker set typically is diverse and imposes an upper limit on the quality achievable for the low-resource speaker. In the current work, we achieve high-quality speech synthesis using as little as five minutes of speech from the desired speaker by augmenting the low-resource speaker data with noise and employing multiple sampling techniques during training. Our method requires only four high-quality, high-resource speakers, which are easy to obtain and use in practice. Our low-complexity method achieves improved speaker similarity compared to the state-of-the-art zero-shot method HierSpeech++ and the recent low-resource method AdapterMix while maintaining comparable naturalness. Our proposed approach can also reduce the data requirements for speech synthesis for new speakers and languages.

语音合成低资源噪声增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。