针对多模态模型脆弱性差异,提出针对性增强鲁棒性的训练方法。
Vulnerability-Aware Robust Multimodal Adversarial Training
- 基于攻击目标的一阶近似量化各模态脆弱性,识别薄弱环节。
- 对高脆弱模态施加定向正则化,提升整体鲁棒性,准确率不降。
- 在三个数据集上分别提升12.73%、22.21%、11.19%鲁棒性,适合安全敏感场景。
多模态学习通过融合多种模态在各类任务中表现优异,但模态间的相互依赖也使其更易受对抗攻击。现有方法或仅针对特定模态攻击,或无差别地攻击所有模态,忽视了不同模态对最终鲁棒性的贡献差异,导致鲁棒性不足。为此,本文提出脆弱性感知的多模态对抗训练方法(VARMAT),该方法通过探针-训练框架,先显式量化各模态的脆弱性(基于攻击目标的一阶近似),再引入针对高脆弱模态的定向正则化项,引导模型在保持任务精度的同时提升鲁棒性。实验表明,该方法在包含多种模态的多个多模态数据集上均显著增强鲁棒性,分别实现12.73%、22.21%和11.19%的提升,揭示了现有对抗训练中的重要盲区。
原文摘要 · Abstract (English)
Multimodal learning has shown significant superiority on various tasks by integrating multiple modalities. However, the interdependencies among modalities increase the susceptibility of multimodal models to adversarial attacks. Existing methods mainly focus on attacks on specific modalities or indiscriminately attack all modalities. In this paper, we find that these approaches ignore the differences between modalities in their contribution to final robustness, resulting in suboptimal robustness performance. To bridge this gap, we introduce Vulnerability-Aware Robust Multimodal Adversarial Training (VARMAT), a probe-in-training adversarial training method that improves multimodal robustness by identifying the vulnerability of each modality. To be specific, VARMAT first explicitly quantifies the vulnerability of each modality, grounded in a first-order approximation of the attack objective (Probe). Then, we propose a targeted regularization term that penalizes modalities with high vulnerability, guiding robust learning while maintaining task accuracy (Training). We demonstrate the enhanced robustness of our method across multiple multimodal datasets involving diverse modalities. Finally, we achieve {12.73%, 22.21%, 11.19%} robustness improvement on three multimodal datasets, revealing a significant blind spot in multimodal adversarial training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。