arXiv:2507.16856cs.CVcs.AI2025-07中稿 · Safe and Trustwort…被引 3

让视觉语言模型读懂潜在恶意意图,生成更安全回复。

SIA: Enhancing Safety via Intent Awareness for Vision-Language Models

  • 通过图像描述+少样本思维链推断隐含意图
  • 在多个安全基准上显著降低有害输出
  • 无需重训练,适合部署于现有模型

随着视觉语言模型(VLMs)在真实场景中广泛应用,以往被忽视的安全风险日益凸显。特别是看似无害的多模态输入组合可能暗藏有害意图,导致模型生成不安全输出。现有方法难以应对这种由模态间交互引发的潜在风险。本文提出SIA(Safety via Intent Awareness),一种无需训练的意图感知安全框架,可主动检测多模态输入中的有害意图,并据此引导生成安全响应。SIA采用三阶段流程:(1) 通过图像描述进行视觉抽象;(2) 基于少样本思维链(CoT)提示进行意图推理;(3) 依据意图条件生成回应。通过动态适配从图文对中推断出的隐含意图,SIA在不需大规模重训练的前提下有效缓解有害输出。在SIUO、MM-SafetyBench和HoliSafe等安全基准上的大量实验表明,SIA持续提升安全性,优于现有训练自由方法。

原文摘要 · Abstract (English)

With the growing deployment of Vision-Language Models (VLMs) in real-world applications, previously overlooked safety risks are becoming increasingly evident. In particular, seemingly innocuous multimodal inputs can combine to reveal harmful intent, leading to unsafe model outputs. While multimodal safety has received increasing attention, existing approaches often fail to address such latent risks, especially when harmfulness arises only from the interaction between modalities. We propose SIA (Safety via Intent Awareness), a training-free, intent-aware safety framework that proactively detects harmful intent in multimodal inputs and uses it to guide the generation of safe responses. SIA follows a three-stage process: (1) visual abstraction via captioning; (2) intent inference through few-shot chain-of-thought (CoT) prompting; and (3) intent-conditioned response generation. By dynamically adapting to the implicit intent inferred from an image-text pair, SIA mitigates harmful outputs without extensive retraining. Extensive experiments on safety benchmarks, including SIUO, MM-SafetyBench, and HoliSafe, show that SIA consistently improves safety and outperforms prior training-free methods.

多模态安全意图识别VLM生成安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。