提出视觉主导导致安全失效,用轻量监控提升多模态模型个性化安全
When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

- 发现视觉信息过早干扰文本风险信号,导致模型忽略用户特殊需求
- 86-99%响应直接输出,无一模型在个性化安全上超2.6/5
- 设计PRISM监控机制,可精准预测需延迟回答的场景
视觉语言模型在高风险场景中常因缺乏用户个人背景而产生不安全回应。我们构建了包含5,181个场景的MPS-Bench基准,覆盖12个高危领域,每张真实图像配有隐藏用户画像。评估8个前沿模型发现,它们几乎总是直接回应(86-99%),且无一在个性化安全上超过2.6/5。分析表明存在视觉主导现象:视觉信息早期进入文本表征并抑制文本风险信号。因果干预揭示双阶段机制——视觉情感先在早期层传入文本流,再通过修改后的文本影响最终决策,使后期修复不可靠。为此提出PRISM,一种基于双向跨模态调制的轻量输入监控器,可准确预测需延迟响应的查询。PRISM实现0.978 AUC,严格优于所有测试模型的安全-效用权衡前沿。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly deployed in high-stakes settings, where a response that is reasonable in general may still be unsafe for a particular user whose medical, emotional, or situational context is unknown to the model. We study this problem of personalized safety in multimodal systems and introduce MPS-Bench, a benchmark of 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile. Evaluating eight frontier VLMs, we find that they almost always respond directly (86-99%) rather than seek missing context, and none exceeds 2.6/5 on personalized safety. To understand why these failures arise, we analyze multimodal interactions and identify visual dominance: visual information enters text representations early and suppresses textual risk signals during multimodal fusion. Causal interventions reveal a two-stage mechanism in which visual affect is first transferred into the text stream in early layers and then shapes the final decision through this altered text representation, making late-stage internal remediation unreliable. Motivated by this mechanism, we propose PRISM, a lightweight input monitor that uses bidirectional cross-modal modulation to predict when a query is likely to require deferral. PRISM achieves 0.978 AUC and strictly dominates the safety-utility Pareto frontier across all tested models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。