arXiv:2503.19591cs.SDcs.CR2025-03中稿 · ICME 2025被引 4

通过优化声学特征提升语音对抗样本跨模型迁移能力

Boosting the Transferability of Audio Adversarial Examples with Acoustic Representation Optimization

  • 利用低层声学表示引导扰动生成,避免依赖具体模型
  • 在三种ASR模型上提升攻击迁移率,且音频质量不变
  • 可插拔集成现有攻击方法,适合真实场景隐蔽攻击

随着自动语音识别(ASR)系统的广泛应用,其对对抗攻击的脆弱性已受到广泛关注。然而,现有大多数对抗样本仅针对特定模型生成,缺乏跨模型迁移能力。在现实场景中,攻击者通常无法获取目标模型的详细信息,导致基于查询的攻击不可行。为此,我们提出声学表示优化技术,将对抗扰动对齐于语音表示模型提取的底层声学特征。与依赖模型特定高层抽象不同,该方法利用在多种ASR架构间保持一致的底层声学表示。通过引入声学表示损失,引导扰动向这些稳健的低层特征靠近,从而在不降低音频质量的前提下显著提升对抗样本的跨模型迁移能力。该方法具有即插即用特性,可与任意现有攻击方法结合。我们在三种现代ASR模型上进行了评估,实验结果表明,该方法显著提升了已有方法生成的对抗样本的迁移性能。

原文摘要 · Abstract (English)

With the widespread application of automatic speech recognition (ASR) systems, their vulnerability to adversarial attacks has been extensively studied. However, most existing adversarial examples are generated on specific individual models, resulting in a lack of transferability. In real-world scenarios, attackers often cannot access detailed information about the target model, making query-based attacks unfeasible. To address this challenge, we propose a technique called Acoustic Representation Optimization that aligns adversarial perturbations with low-level acoustic characteristics derived from speech representation models. Rather than relying on model-specific, higher-layer abstractions, our approach leverages fundamental acoustic representations that remain consistent across diverse ASR architectures. By enforcing an acoustic representation loss to guide perturbations toward these robust, lower-level representations, we enhance the cross-model transferability of adversarial examples without degrading audio quality. Our method is plug-and-play and can be integrated with any existing attack methods. We evaluate our approach on three modern ASR models, and the experimental results demonstrate that our method significantly improves the transferability of adversarial examples generated by previous methods while preserving the audio quality.

语音对抗迁移攻击声学特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。