arXiv:2605.22083cs.SDcs.LG2026-05中稿 · INTERSPEECH 2026

通过增强对抗性流匹配,提升语音合成对错位的鲁棒性。

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching

  • 引入保持长度的重复与跳过增强,直接惩罚真实错误模式。
  • 在零样本场景下,英文字符错误率从0.48%降至0.35%。
  • 无需外部对齐器,可无缝集成到现有语音合成流程中。

尽管基于流匹配的文本到语音(TTS)系统在零样本说话人相似性和自然度方面表现优异,但仍面临内容保真度问题,尤其在对齐不完美时易出现跳词和重复错误。本文提出RobustSpeechFlow,一种训练策略,通过在对比流匹配中引入保持长度的重复与跳词潜在增强,提升对齐鲁棒性。该方法无需外部对齐器或偏好数据,直接惩罚真实故障模式,并可轻松融入现有流水线。在Seed-TTS-eval上,仅使用0.06B参数,将词错误率(WER)从1.44降至1.38;在ZERO500基准测试中,跨不同说话人和语调条件均实现一致可懂度提升:在NFE=24时,英文字符错误率(CER)由0.48%降至0.35%,韩文CER由0.81%降至0.57%。

原文摘要 · Abstract (English)

While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repeat errors from imperfect alignment. We propose RobustSpeechFlow, a training strategy that improves alignment robustness by extending contrastive flow matching with length-preserving repeat and skip latent augmentations. Requiring no external aligners or preference data, our method directly penalizes realistic failure modes and readily integrates into existing pipelines. On Seed-TTS-eval, it reduces the word error rate (WER) from 1.44 to 1.38 using only 0.06B parameters. On our ZERO500 benchmark, it delivers consistent intelligibility improvements across diverse speaker and prosody conditions; at NFE=24, it reduces English character error rate (CER) from 0.48\% to 0.35\% and Korean CER from 0.81\% to 0.57\%. Audio samples: https://robustspeechflow.github.io/

语音合成流匹配鲁棒性TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。