提出可通用的跨模态攻击框架,提升对抗样本迁移能力。
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
- 分层优化图像与文本的扰动路径,避免局部最优。
- 在多个模型和数据集上实现更强的攻击迁移性。
- 适合研究模型鲁棒性或防御机制的研究者。
现有视觉-语言模型的对抗攻击多为样本特定,扩展至大规模数据集或新场景时计算开销巨大。为此,本文提出分层精炼攻击(HRA),一种面向视觉-语言模型的通用多模态攻击框架。针对图像模态,通过利用历史梯度与预测未来梯度的时间层次结构,优化迭代路径,避免局部极小值并稳定通用扰动学习;针对文本模态,分层建模文本重要性,综合考虑句内与句间贡献,识别出全局影响显著的词,作为通用文本扰动。在多种下游任务、视觉-语言模型及数据集上的大量实验表明,所提通用多模态攻击具有优异的迁移性能。
原文摘要 · Abstract (English)
Existing adversarial attacks for VLP models are mostly sample-specific, resulting in substantial computational overhead when scaled to large datasets or new scenarios. To overcome this limitation, we propose Hierarchical Refinement Attack (HRA), a multimodal universal attack framework for VLP models. For the image modality, we refine the optimization path by leveraging a temporal hierarchy of historical and estimated future gradients to avoid local minima and stabilize universal perturbation learning. For the text modality, it hierarchically models textual importance by considering both intra- and inter-sentence contributions to identify globally influential words, which are then used as universal text perturbations. Extensive experiments across various downstream tasks, VLP models, and datasets, demonstrate the superior transferability of the proposed universal multimodal attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。