发现视觉语言模型隐性推理存在安全漏洞,可致无害输入生成危险输出。
Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language Models
- 提出隐性推理安全新概念,揭示多模态输入下模型内在风险
- 构建首个针对该问题的数据集SSUI,验证漏洞普遍存在
- 仅用简单上下文学习即显著缓解风险,适合安全研究者参考
大型视觉语言模型在多模态输入下面临日益严峻的安全挑战。本文提出隐性推理安全这一新概念,指出即使看似无害的组合输入,也可能因模型内部缺陷或隐藏推理路径导致不安全输出。为验证此问题,我们构建了首个相关数据集Safe Semantics, Unsafe Interpretations(SSUI)。实验表明,仅通过简单的上下文学习(In-Context Learning)结合SSUI,即可显著降低此类隐性多模态威胁,凸显提升跨模态隐性推理安全性的紧迫性。
原文摘要 · Abstract (English)
Large Vision-Language Models face growing safety challenges with multimodal inputs. This paper introduces the concept of Implicit Reasoning Safety, a vulnerability in LVLMs. Benign combined inputs trigger unsafe LVLM outputs due to flawed or hidden reasoning. To showcase this, we developed Safe Semantics, Unsafe Interpretations, the first dataset for this critical issue. Our demonstrations show that even simple In-Context Learning with SSUI significantly mitigates these implicit multimodal threats, underscoring the urgent need to improve cross-modal implicit reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。