arXiv:2507.07709cs.CV2025-07ICCV被引 1

构建跨任务对抗攻击基准,测试统一视觉语言模型的漏洞。

One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models

  • 提出基于区域的攻击框架CRAFT,精准操控目标物体分类
  • 在四个任务上实现92.3%的同步误分类率,显著优于现有方法
  • 适合安全研究者评估多任务模型鲁棒性

统一视觉语言模型(VLMs)近年来取得显著进展,通过不同指令即可在共享计算架构下灵活处理多样任务。这种基于指令的控制机制带来独特安全挑战:对抗输入需对不可预测的任务指令保持有效性。本文提出CrossVLAD,一个从MSCOCO数据集精心构建的基准数据集,利用GPT-4辅助标注,系统评估统一VLMs上的跨任务对抗攻击。CrossVLAD聚焦于对象变更目标,在四个下游任务中一致操纵目标物体分类,并提出新颖的成功率度量,衡量所有任务中的同步误分类表现,严格评估对抗迁移能力。为应对该挑战,我们提出CRAFT(跨任务区域攻击框架,带标记对齐),一种高效的区域中心攻击方法。在Florence-2及其他主流统一VLMs上的大量实验表明,本方法在整体跨任务攻击性能和目标物体变更成功率方面均优于现有方法,凸显其在多任务场景下有效干扰统一VLMs的能力。

原文摘要 · Abstract (English)

Unified vision-language models(VLMs) have recently shown remarkable progress, enabling a single model to flexibly address diverse tasks through different instructions within a shared computational architecture. This instruction-based control mechanism creates unique security challenges, as adversarial inputs must remain effective across multiple task instructions that may be unpredictably applied to process the same malicious content. In this paper, we introduce CrossVLAD, a new benchmark dataset carefully curated from MSCOCO with GPT-4-assisted annotations for systematically evaluating cross-task adversarial attacks on unified VLMs. CrossVLAD centers on the object-change objective-consistently manipulating a target object's classification across four downstream tasks-and proposes a novel success rate metric that measures simultaneous misclassification across all tasks, providing a rigorous evaluation of adversarial transferability. To tackle this challenge, we present CRAFT (Cross-task Region-based Attack Framework with Token-alignment), an efficient region-centric attack method. Extensive experiments on Florence-2 and other popular unified VLMs demonstrate that our method outperforms existing approaches in both overall cross-task attack performance and targeted object-change success rates, highlighting its effectiveness in adversarially influencing unified VLMs across diverse tasks.

对抗攻击视觉语言模型安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。