arXiv:2601.01852eess.AScs.AI2026-01被引 1

提出新型多目标攻击,同时破坏语音识别准确率与效率。

MORE: Multi-Objective Adversarial Attacks on Speech Recognition

  • 构建分阶段对抗优化框架,协同降低识别准确率与推理速度。
  • 通过重复加倍机制使输出文本长度显著增加,词错误率保持高位。
  • 适合研究模型鲁棒性或安全防御的开发者参考。

大型语音识别模型(如 Whisper)在真实场景中广泛应用,其对微小输入扰动的鲁棒性至关重要。现有研究主要关注对抗攻击导致的准确率下降,而对推理效率的影响仍缺乏探索。为此,本文提出 MORE:一种多目标重复加倍激励攻击,通过分阶段排斥-锚定机制,联合降低识别准确率与推理效率。具体地,将多目标对抗优化重构为分层框架,逐步实现双重目标;并引入新颖的重复鼓励加倍目标(REDO),通过维持准确率下降并周期性翻倍预测序列长度,诱导模型产生冗长错误文本。实验表明,相比基线方法,MORE 能持续生成显著更长的转录文本,同时保持高词错误率,验证了其在多目标对抗攻击中的有效性。

原文摘要 · Abstract (English)

The emergence of large-scale automatic speech recognition (ASR) models such as Whisper has greatly expanded their adoption across diverse real-world applications. Ensuring robustness against even minor input perturbations is therefore critical for maintaining reliable performance in real-time environments. While prior work has mainly examined accuracy degradation under adversarial attacks, robustness with respect to efficiency remains largely unexplored. This narrow focus provides only a partial understanding of ASR model vulnerabilities. To address this gap, we conduct a comprehensive study of ASR robustness under multiple attack scenarios. We introduce MORE, a multi-objective repetitive doubling encouragement attack, which jointly degrades recognition accuracy and inference efficiency through a hierarchical staged repulsion-anchoring mechanism. Specifically, we reformulate multi-objective adversarial optimization into a hierarchical framework that sequentially achieves the dual objectives. To further amplify effectiveness, we propose a novel repetitive encouragement doubling objective (REDO) that induces duplicative text generation by maintaining accuracy degradation and periodically doubling the predicted sequence length. Overall, MORE compels ASR models to produce incorrect transcriptions at a substantially higher computational cost, triggered by a single adversarial input. Experiments show that MORE consistently yields significantly longer transcriptions while maintaining high word error rates compared to existing baselines, underscoring its effectiveness in multi-objective adversarial attack.

语音识别对抗攻击鲁棒性多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。