用零样本语音合成数据增强,提升小样本个性化语音合成质量
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
- 通过轻量级域嵌入区分真实与合成语音,避免混淆
- 在极少量目标数据下仍能保持说话人相似度,提升音质与可懂性
- 无需修改基础模型,适合资源受限场景的语音合成应用
我们研究将零样本文本到语音(ZS-TTS)作为低资源个性化语音合成的数据增强来源。尽管合成数据能提供丰富的语言和语音多样性,但直接混合大量合成语音与有限的真实录音会导致微调时说话人特征退化。为此,我们提出ZeSTA,一种简单的域条件训练框架:通过轻量级域嵌入区分真实与合成语音,并结合真实数据过采样,在不修改基础架构的前提下,稳定极小规模目标数据下的适应过程。在LibriTTS和一个内部数据集上,使用两种ZS-TTS源的实验表明,该方法在保持可懂性和感知质量的同时,显著提升了说话人相似度。音频示例可在官网获取。
原文摘要 · Abstract (English)
We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide linguistically rich and phonetically diverse speech, naively mixing large amounts of synthetic speech with limited real recordings often leads to speaker similarity degradation during fine-tuning. To address this issue, we propose ZeSTA, a simple domain-conditioned training framework that distinguishes real and synthetic speech via a lightweight domain embedding, combined with real-data oversampling to stabilize adaptation under extremely limited target data, without modifying the base architecture. Experiments on LibriTTS and an in-house dataset with two ZS-TTS sources demonstrate that our approach improves speaker similarity over naive synthetic augmentation while preserving intelligibility and perceptual quality. Audio samples are available on our web page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。