arXiv:2505.21967cs.CL2025-05被引 5

揭示视觉语言模型的对抗攻击漏洞,提出系统评估框架与安全对齐标准。

Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack

  • 分两阶段评估攻击:区分拒绝类型与有害意图实现程度。
  • 发现传统攻击可绕过模型安全机制,暴露新威胁面。
  • 提出理想行为规范,指导多模态系统安全对齐。

大型视觉语言模型(LVLMs)在多模态任务中表现出卓越能力,但其融合视觉输入的特性带来了新的攻击面,导致新型安全漏洞。本文通过系统性表征分析,揭示为何传统对抗攻击能绕过LVLM中的安全机制。我们提出一种两阶段评估框架:第一阶段区分指令不合规、直接拒绝和成功对抗利用;第二阶段量化模型输出满足有害提示意图的程度,并将拒绝行为细分为直接拒绝、软拒绝和部分仍具帮助性的拒绝。最后,我们引入一种规范性框架,定义面对有害提示时理想的模型行为,为多模态系统的安全对齐提供原则性目标。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have shown remarkable capabilities across a wide range of multimodal tasks. However, their integration of visual inputs introduces expanded attack surfaces, thereby exposing them to novel security vulnerabilities. In this work, we conduct a systematic representational analysis to uncover why conventional adversarial attacks can circumvent the safety mechanisms embedded in LVLMs. We further propose a novel two stage evaluation framework for adversarial attacks on LVLMs. The first stage differentiates among instruction non compliance, outright refusal, and successful adversarial exploitation. The second stage quantifies the degree to which the model's output fulfills the harmful intent of the adversarial prompt, while categorizing refusal behavior into direct refusals, soft refusals, and partial refusals that remain inadvertently helpful. Finally, we introduce a normative schema that defines idealized model behavior when confronted with harmful prompts, offering a principled target for safety alignment in multimodal systems.

对抗攻击视觉语言模型安全对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。