arXiv:2510.13698cs.CV2025-10被引 4

通过增强视觉注意力提升多模态安全检测,低成本防御图像隐含恶意攻击。

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment

  • 利用简洁视觉上下文强化关键区域注意力,精准识别风险
  • 仅需少量校准数据即可评估风险,显著降低部署开销
  • 适用于多种多模态越狱攻击,兼顾安全性与推理效率

当前AI模型在包含图像中隐含恶意意图的多模态查询下仍易出错。尽管广泛采用多模态安全数据集训练以对齐安全策略,但数据标注与训练成本高昂。为降低成本,近期研究探索推理时对齐,但普遍存在跨多样化多模态越狱攻击泛化性差、需额外前向传播或复杂校准流程等问题。本文发现,视觉注意力对安全关键图像区域关注不足是多模态安全失效的核心原因。基于此,提出多模态风险自适应引导(MoRAS),通过简明视觉上下文增强对关键区域的注意力,实现准确的风险评估。该风险信号支持直接拒绝响应,减少推理开销,且对多种越狱攻击保持泛化能力。值得注意的是,MoRAS仅需少量校准集即可估计多模态风险,大幅降低预部署负担。在多个基准和多模态大模型(MLLM)骨干网络上进行实证验证,结果表明,相较前沿推理时防御方法,MoRAS始终能有效缓解越狱攻击,保持模型效用,并降低计算开销。

原文摘要 · Abstract (English)

Even modern AI models often remain vulnerable to multimodal queries in which harmful intent is embedded in images. A widely used approach for safety alignment is training with extensive multimodal safety datasets, but the costs of data curation and training are often prohibitive. To mitigate these costs, inference-time alignment has recently been explored, but they often lack generalizability across diverse multimodal jailbreaks and still incur notable overhead due to extra forward passes for response refinement or heavy pre-deployment calibration procedures. Here, we identify insufficient visual attention to safety-critical image regions as one of the key causes of multimodal safety failures. Building on this insight, we propose Multimodal Risk-Adaptive Steering (MoRAS), which enhances safety-critical visual attention via concise visual contexts for accurate multimodal risk assessment. This risk signal enables risk-adaptive steering for direct refusals, reducing inference overhead while remaining generalizable across diverse multimodal jailbreaks. Notably, MoRAS requires only a small calibration set to estimate multimodal risk, substantially reducing pre-deployment overhead. We conduct various empirical validations across multiple benchmarks and MLLM backbones, and observe that the proposed MoRAS consistently mitigates jailbreaks, preserves utility, and reduces computational overhead compared to state-of-the-art inference-time defenses.

多模态安全视觉注意力风险检测推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。