arXiv:2512.17293cs.SDcs.AI2025-12

用自净化流匹配提升语音合成抗噪声能力,实测表现最佳。

Robust TTS Training via Self-Purifying Flow Matching for the WildSpoof 2026 TTS Track

  • 通过对比条件与无条件流匹配损失,自动识别并处理异常语音对。
  • 在真实场景下实现最低词错误率(WER),感知质量排名第二。
  • 适合追求高效鲁棒语音合成的开发者或工业应用团队。

本文针对野战语音伪造挑战赛的语音合成赛道,提出一种轻量级文本到语音(TTS)系统。该方法基于开源模型Supertonic,采用自净化流匹配(SPFM)进行微调,以增强对真实环境语音的鲁棒性。SPFM通过比较每个样本的条件与无条件流匹配损失,将可疑的文本-语音对路由至无条件训练,同时保留其声学信息。实验表明,该模型在所有参赛队伍中取得最低词错误率(WER),在感知指标如UTMOS和DNSMOS上位列第二。结果表明,结合显式噪声处理机制如SPFM,开放权重架构如Supertonic可有效适应多样化真实语音场景。

原文摘要 · Abstract (English)

This paper presents a lightweight text-to-speech (TTS) system developed for the WildSpoof Challenge TTS Track. Our approach fine-tunes the recently released open-weight TTS model, \textit{Supertonic}\footnote{\url{https://github.com/supertone-inc/supertonic}}, with Self-Purifying Flow Matching (SPFM) to enable robust adaptation to in-the-wild speech. SPFM mitigates label noise by comparing conditional and unconditional flow matching losses on each sample, routing suspicious text--speech pairs to unconditional training while still leveraging their acoustic information. The resulting model achieves the lowest Word Error Rate (WER) among all participating teams, while ranking second in perceptual metrics such as UTMOS and DNSMOS. These findings demonstrate that efficient, open-weight architectures like Supertonic can be effectively adapted to diverse real-world speech conditions when combined with explicit noise-handling mechanisms such as SPFM.

语音合成鲁棒性流匹配噪声处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。