通过特征平滑增强训练,提升通用声码器合成自然度
Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems
- 对输入声学特征随机加线性平滑滤波,增强泛化能力
- 在Tacotron 2和FastSpeech 2上分别提升11.99%和12.05%的评分
- 无需修改模型结构,适配任意声码器与声学模型
尽管通用声码器已在多种语音上实现高质量波形生成,但将其集成到文本转语音(TTS)任务中常导致合成质量下降。为解决此问题,本文提出一种新颖的通用声码器训练增强方法:在训练时随机对输入声学特征应用线性平滑滤波,促进声码器在广泛平滑程度下的泛化能力。该方法显著缓解了训练与推理间的不匹配问题,即使声学模型输出过度平滑的特征,也能提升合成语音的自然度。值得注意的是,该方法适用于任意声码器,无需架构修改或依赖特定声学模型。实验验证了所提声码器的优越性,在集成Tacotron 2和FastSpeech 2声学模型时,分别获得11.99%和12.05%的平均意见分(MOS)提升。
原文摘要 · Abstract (English)
While universal vocoders have achieved proficient waveform generation across diverse voices, their integration into text-to-speech (TTS) tasks often results in degraded synthetic quality. To address this challenge, we present a novel augmentation technique for training universal vocoders. Our training scheme randomly applies linear smoothing filters to input acoustic features, facilitating vocoder generalization across a wide range of smoothings. It significantly mitigates the training-inference mismatch, enhancing the naturalness of synthetic output even when the acoustic model produces overly smoothed features. Notably, our method is applicable to any vocoder without requiring architectural modifications or dependencies on specific acoustic models. The experimental results validate the superiority of our vocoder over conventional methods, achieving 11.99% and 12.05% improvements in mean opinion scores when integrated with Tacotron 2 and FastSpeech 2 TTS acoustic models, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。