研究音视频模型的抗攻击弱点,提出两类新攻击与防御方法。
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
- 从时间与模态角度设计两类对抗攻击,利用时序冗余和音视频不匹配
- 在Kinetics-Sounds数据集上显著降低模型性能,验证攻击有效性
- 提出高效防御框架,提升抗攻击能力与训练效率,适合多模态安全研究者
尽管音视频学习通过融合多种感官模态增强了对现实世界的理解,但这种整合也引入了新的对抗攻击漏洞。本文从时间和模态两个维度全面研究音视频模型的对抗鲁棒性。提出两种强大攻击:1)利用连续时间片段间固有时序冗余的时序不变性攻击;2)制造音频与视觉模态之间不一致的模态错位攻击。这些攻击旨在全面评估模型面对多样化威胁的脆弱性。为应对上述攻击,我们引入一种新型音视频对抗训练框架,解决了传统对抗训练中的关键挑战,包括针对多模态数据的高效扰动生成方法及对抗课程学习策略。在Kinetics-Sounds数据集上的大量实验表明,所提出的时空与模态相关攻击能有效削弱模型性能,达到当前最优水平;而所提防御方法显著提升了对抗鲁棒性与训练效率。
原文摘要 · Abstract (English)
While audio-visual learning equips models with a richer understanding of the real world by leveraging multiple sensory modalities, this integration also introduces new vulnerabilities to adversarial attacks. In this paper, we present a comprehensive study of the adversarial robustness of audio-visual models, considering both temporal and modality-specific vulnerabilities. We propose two powerful adversarial attacks: 1) a temporal invariance attack that exploits the inherent temporal redundancy across consecutive time segments and 2) a modality misalignment attack that introduces incongruence between the audio and visual modalities. These attacks are designed to thoroughly assess the robustness of audio-visual models against diverse threats. Furthermore, to defend against such attacks, we introduce a novel audio-visual adversarial training framework. This framework addresses key challenges in vanilla adversarial training by incorporating efficient adversarial perturbation crafting tailored to multi-modal data and an adversarial curriculum strategy. Extensive experiments in the Kinetics-Sounds dataset demonstrate that our proposed temporal and modality-based attacks in degrading model performance can achieve state-of-the-art performance, while our adversarial training defense largely improves the adversarial robustness as well as the adversarial training efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。