arXiv:2603.19092cs.CVcs.AI2026-03被引 3

用语义提示可操控视觉语言模型的安全判断,揭示其依赖联想而非真实理解。

SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues

  • 通过文本、视觉和认知干预,不改变场景内容即可引导模型安全决策。
  • 多模型实验表明安全判断对语义提示高度敏感,易受虚假关联影响。
  • 适合关注多模态安全漏洞的研究者或系统开发者参考。

视觉语言模型(VLMs)在现实世界与具身应用中越来越多地用于安全决策,但其判断所依赖的视觉证据尚不明确。本文研究是否可通过简单语义提示来引导VLM的多模态安全行为。提出一种语义引导框架,通过控制性文本、视觉和认知干预实现决策引导,而无需改变原始场景内容。为此构建SAVeS基准,用于评估语义提示下的情境安全,配套评估协议可区分行为拒绝、基于依据的安全推理与错误拒绝。跨多个VLM及一个先进基准的实验表明,安全决策对语义提示极为敏感,反映出模型更依赖学习到的视觉-语言关联,而非真实的视觉理解。进一步证明自动化引导流程可利用这些机制,揭示多模态安全系统存在潜在漏洞。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly deployed in real-world and embodied settings where safety decisions depend on visual context. However, it remains unclear which visual evidence drives these judgments. We study whether multimodal safety behavior in VLMs can be steered by simple semantic cues. We introduce a semantic steering framework that applies controlled textual, visual, and cognitive interventions without changing the underlying scene content. To evaluate these effects, we propose SAVeS, a benchmark for situational safety under semantic cues, together with an evaluation protocol that separates behavioral refusal, grounded safety reasoning, and false refusals. Experiments across multiple VLMs and an additional state-of-the-art benchmark show that safety decisions are highly sensitive to semantic cues, indicating reliance on learned visual-linguistic associations rather than grounded visual understanding. We further demonstrate that automated steering pipelines can exploit these mechanisms, highlighting a potential vulnerability in multimodal safety systems.

视觉语言模型安全判断语义提示多模态漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。