arXiv:2509.23917cs.CV2025-09被引 1

发现细粒度任务的对抗样本更易跨任务攻击CLIP,提出新框架提升攻击效果。

Bridging the Task Gap: Multi-Task Adversarial Transferability in CLIP and Its Derivatives

  • 设计任务感知特征聚合损失,增强对抗扰动的跨任务泛化能力
  • 在多任务上平均攻击成功率提升超39%,不增加扰动预算
  • 揭示了多任务CLIP中对抗样本的转移机制,适合安全与鲁棒性研究者

作为通用视觉-语言预训练模型,CLIP在图像-文本对齐任务中表现优异,广泛应用于图像分类和图文检索等下游任务。然而,在细粒度任务如目标检测和语义分割上表现较弱。尽管已有多种变体尝试改进其性能,但其对对抗扰动的鲁棒性仍缺乏深入研究。理解对抗样本在不同任务间的迁移特性,是评估CLIP泛化边界与安全风险的关键。本文系统分析了基于CLIP的模型在图像-文本检索、目标检测和语义分割任务下,面对对抗扰动时的跨任务迁移行为。发现从细粒度任务(如目标检测、语义分割)生成的对抗样本往往比粗粒度任务具有更强的迁移能力,能更有效地攻击原始CLIP模型。受此启发,我们提出多任务对抗性CLIP(MT-AdvCLIP)框架,引入任务感知特征聚合损失,并生成具备更强跨任务泛化能力的扰动。实验结果表明,该方法在多个公开数据集上显著提升了对抗迁移成功率(多任务平均攻击成功率达39%以上),且不增加扰动预算。本研究揭示了多任务CLIP模型中对抗样本的迁移机制,为多任务鲁棒性评估与对抗样本设计提供了新视角。

原文摘要 · Abstract (English)

As a general-purpose vision-language pretraining model, CLIP demonstrates strong generalization ability in image-text alignment tasks and has been widely adopted in downstream applications such as image classification and image-text retrieval. However, it struggles with fine-grained tasks such as object detection and semantic segmentation. While many variants aim to improve CLIP on these tasks, its robustness to adversarial perturbations remains underexplored. Understanding how adversarial examples transfer across tasks is key to assessing CLIP's generalization limits and security risks. In this work, we conduct a systematic empirical analysis of the cross-task transfer behavior of CLIP-based models on image-text retrieval, object detection, and semantic segmentation under adversarial perturbations. We find that adversarial examples generated from fine-grained tasks (e.g., object detection and semantic segmentation) often exhibit stronger transfer potential than those from coarse-grained tasks, enabling more effective attacks against the original CLIP model. Motivated by this observation, we propose a novel framework, Multi-Task Adversarial CLIP (MT-AdvCLIP), which introduces a task-aware feature aggregation loss and generates perturbations with enhanced cross-task generalization capability. This design strengthens the attack effectiveness of fine-grained task models on the shared CLIP backbone. Experimental results on multiple public datasets show that MT-AdvCLIP significantly improves the adversarial transfer success rate (The average attack success rate across multiple tasks is improved by over 39%.) against various CLIP-derived models, without increasing the perturbation budget. This study reveals the transfer mechanism of adversarial examples in multi-task CLIP models, offering new insights into multi-task robustness evaluation and adversarial example design.

对抗样本CLIP多任务鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。