arXiv:2602.11933cs.CL2026-02中稿 · INTERSPEECH2026

用文本攻击生成语音模型的鲁棒性,无需真实对抗语音数据。

Cross-Modal Robustness Transfer (CMRT): Training Robust Speech Translation Models Using Adversarial Text

  • 从文本攻击迁移鲁棒性到语音,避免生成对抗语音数据。
  • 跨语言对实验显示平均提升3 BLEU以上,抗攻击能力显著增强。
  • 适合关注语音翻译鲁棒性的研究者与工业落地团队。

端到端语音翻译(E2E-ST)虽有长足进步,但现有模型多在干净数据上评估,忽略真实场景中的形态鲁棒性挑战,如非母语或方言语音中的词形变化。本文将针对词形变化的文本对抗攻击方法拓展至语音域,证明当前先进E2E-ST模型对此类攻击高度脆弱。尽管对抗训练在文本任务中有效,但生成高质量对抗语音数据成本高、技术难。为此,我们提出跨模态鲁棒性迁移(CMRT)框架,实现从文本模态向语音模态的鲁棒性迁移,训练中无需对抗语音数据。在四个语言对上的大量实验表明,CMRT平均提升超过3 BLEU,建立无需生成对抗语音数据的新基准。

原文摘要 · Abstract (English)

End-to-End Speech Translation (E2E-ST) has seen significant advancements, yet current models are primarily benchmarked on curated, "clean" datasets. This overlooks critical real-world challenges, such as morphological robustness to inflectional variations common in non-native or dialectal speech. In this work, we adapt a text-based adversarial attack targeting inflectional morphology to the speech domain and demonstrate that state-of-the-art E2E-ST models are highly vulnerable it. While adversarial training effectively mitigates such risks in text-based tasks, generating high-quality adversarial speech data remains computationally expensive and technically challenging. To address this, we propose Cross-Modal Robustness Transfer (CMRT), a framework that transfers adversarial robustness from the text modality to the speech modality. Our method eliminates the requirement for adversarial speech data during training. Extensive experiments across four language pairs demonstrate that CMRT improves adversarial robustness by an average of more than 3 BLEU points, establishing a new baseline for robust E2E-ST without the overhead of generating adversarial speech.

语音翻译对抗攻击鲁棒性跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。