用增强音频提示实现零样本语音合成,提升真实场景下的语音自然度。
Zero-Shot TTS With Enhanced Audio Prompts: Bsc Submission For The 2026 Wildspoof Challenge TTS Track
- 采用非自回归架构与灵活时长建模,提升语音韵律自然性。
- 多阶段增强后合成语音达4.21分UTMOS,3.47分DNSMOS。
- 适用于高噪声环境下语音生成,适合语音合成竞赛应用。
我们评估了两种非自回归架构StyleTTS2和F5-TTS,以应对真实场景中自发性语音的特点。模型采用灵活时长建模提升语调自然度。为处理声学噪声,我们使用Sidon模型构建多阶段增强流水线,显著优于标准Demucs,在信号质量上表现更优。实验表明,对增强后的音频进行微调可提升鲁棒性,最高达到4.21 UTMOS和3.47 DNSMOS。此外,我们分析了参考提示的质量与长度对零样本合成性能的影响,验证了该方法在真实语音生成中的有效性。
原文摘要 · Abstract (English)
We evaluate two non-autoregressive architectures, StyleTTS2 and F5-TTS, to address the spontaneous nature of in-the-wild speech. Our models utilize flexible duration modeling to improve prosodic naturalness. To handle acoustic noise, we implement a multi-stage enhancement pipeline using the Sidon model, which significantly outperforms standard Demucs in signal quality. Experimental results show that finetuning enhanced audios yields superior robustness, achieving up to 4.21 UTMOS and 3.47 DNSMOS. Furthermore, we analyze the impact of reference prompt quality and length on zero-shot synthesis performance, demonstrating the effectiveness of our approach for realistic speech generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。