arXiv:2412.20155cs.SDcs.AI2024-12中稿 · ICASSP 2025被引 6

用少量高质量样本实现稳定音色适配的语音合成方法

Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting

  • 用预训练数据中的优质样本作为音调参考,保持语调一致性
  • 在少量甚至嘈杂目标样本下仍能准确还原说话人音色
  • 适合语音助手等个性化服务场景,对样本质量要求低

由于个性化语音助手等应用需求,说话人自适应语音合成受到广泛关注。现有方法常对目标语音样本的数量或质量敏感。为此,本文提出 Stable-TTS,一种新型说话人自适应语音合成框架,利用高质预训练数据集中的小部分样本(称为先验样本)作为参考。Stable-TTS 通过先验样本的高质量韵律信息实现韵律一致性,同时有效捕捉目标说话人的音色特征。此外,在微调阶段引入先验保持损失,防止对目标样本过拟合,从而保留对先验样本的合成能力。大量实验表明,即使在目标样本数量有限或存在噪声的情况下,Stable-TTS 仍表现出优异性能。

原文摘要 · Abstract (English)

Speaker-adaptive Text-to-Speech (TTS) synthesis has attracted considerable attention due to its broad range of applications, such as personalized voice assistant services. While several approaches have been proposed, they often exhibit high sensitivity to either the quantity or the quality of target speech samples. To address these limitations, we introduce Stable-TTS, a novel speaker-adaptive TTS framework that leverages a small subset of a high-quality pre-training dataset, referred to as prior samples. Specifically, Stable-TTS achieves prosody consistency by leveraging the high-quality prosody of prior samples, while effectively capturing the timbre of the target speaker. Additionally, it employs a prior-preservation loss during fine-tuning to maintain the synthesis ability for prior samples to prevent overfitting on target samples. Extensive experiments demonstrate the effectiveness of Stable-TTS even under limited amounts of and noisy target speech samples.

语音合成说话人适配小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。