arXiv:2504.03735cs.CRcs.AI2025-04被引 2

通过改变角色与图像位置,暴露多模态模型对输入结构的脆弱性

Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots

  • 设计新型结构扰动攻击,利用角色混淆和图像位置变化诱导有害输出
  • 在8种设置下测试,攻击成功率显著提升,且在残差流中体现负面拒绝特征
  • 提出对抗训练方法,使模型专注内容而非输入结构,有效降低攻击成功率

多模态语言模型通常通过后训练对齐来防止有害内容生成,但现有对齐主要针对助手角色,忽略用户角色的对齐,且依赖固定提示结构中的特殊标记,导致模型在输入偏离预期时易受攻击。本文提出角色-模态攻击(RMA),通过混淆用户与助手角色,并改变图像标记的位置,诱导模型生成有害输出。与修改查询内容的攻击不同,RMA仅操纵输入结构而不改动查询本身。我们在多个视觉语言模型上系统评估了八种不同场景下的攻击效果,发现攻击可组合增强,且其在残差流中的负向拒绝投影显著增加,符合以往成功攻击的特征。为缓解此问题,我们提出一种对抗训练策略:在多种受扰动的有害与良性提示上训练模型,使其不再敏感于角色混淆和模态操控,仅关注查询内容,从而显著降低攻击成功率(ASR),同时保持模型通用能力。

原文摘要 · Abstract (English)

Multimodal Language Models (MMLMs) typically undergo post-training alignment to prevent harmful content generation. However, these alignment stages focus primarily on the assistant role, leaving the user role unaligned, and stick to a fixed input prompt structure of special tokens, leaving the model vulnerable when inputs deviate from these expectations. We introduce Role-Modality Attacks (RMA), a novel class of adversarial attacks that exploit role confusion between the user and assistant and alter the position of the image token to elicit harmful outputs. Unlike existing attacks that modify query content, RMAs manipulate the input structure without altering the query itself. We systematically evaluate these attacks across multiple Vision Language Models (VLMs) on eight distinct settings, showing that they can be composed to create stronger adversarial prompts, as also evidenced by their increased projection in the negative refusal direction in the residual stream, a property observed in prior successful attacks. Finally, for mitigation, we propose an adversarial training approach that makes the model robust against input prompt perturbations. By training the model on a range of harmful and benign prompts all perturbed with different RMA settings, it loses its sensitivity to Role Confusion and Modality Manipulation attacks and is trained to only pay attention to the content of the query in the input prompt structure, effectively reducing Attack Success Rate (ASR) while preserving the model's general utility.

多模态安全对抗攻击模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。