用EMA和双评分筛选,让语音合成更自然且难被检测。
Natural Yet Challenging to Detect: Robust In-the-Wild TTS through EMA and Dual-Scoring Prompt Selection -- Submission for WildSpoof 2026 TTS Track
- 用EMA稳定训练,提升模型泛化能力
- 通过大模型双评分筛选高质量输入,提升语音保真度
- 在多个反语音欺骗系统中表现最佳,适合高仿真需求场景
本文介绍了我们针对WildSpoof挑战赛语音合成赛道的提交方案:基于F5-TTS架构构建F5-TTS-DPS模型。通过在监督微调中引入指数移动平均(EMA),稳定训练过程并提升泛化性能。为增强语音合成保真度,利用大语言模型(LLMs)与大音频语言模型(LALMs)进行双评分提示筛选,过滤参考音频与文本提示,确保输入质量并缓解噪声数据中的对齐问题。实验表明,F5-TTS-DPS在开发集上取得UTMOS 3.20、说话人相似性0.51的优异表现。更重要的是,在三种先进SASV系统下,其a-DCF得分分别达到0.1582、0.5233和0.2562,为所有提交中最佳,表明生成语音最难被检测,具有最高自然度与真实性。结合具有竞争力的词错误率(WER)表现,验证了该方法在生成高保真语音与强欺骗能力方面的有效性。
原文摘要 · Abstract (English)
In this technical report, we describe our submission for the WildSpoof Challenge TTS Track: Text-to-Speech with In-the-Wild Data. We introduce F5-TTS-DPS, a model built upon the F5-TTS architecture. Our approach integrates Exponential Moving Average (EMA) into supervised fine-tuning to stabilize training and improve generalization. To enhance synthesis fidelity, we leverage large language models (LLMs) and large audio language models (LALMs) for dual-scoring prompt selection, filtering reference audio and text prompts to ensure quality while addressing alignment issues in noisy datasets. Experimental evaluation demonstrates that F5-TTS-DPS achieves strong performance with UTMOS of 3.20 and speaker similarity of 0.51 on the development set. More importantly, our model achieves the best a-DCF scores of 0.1582, 0.5233, and 0.2562 across three advanced SASV systems among all submissions, indicating our synthesized speech is the most difficult to detect and exhibits the highest degree of naturalness and authenticity. Combined with competitive WER performance, these results validate the effectiveness of our approach in generating natural-sounding speech with strong spoofing capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。