arXiv:2509.19212cs.CLcs.AI2025-09被引 6

让多模态大模型更懂上下文安全,避免误拒或漏判。

Steering Multimodal Large Language Models Decoding for Context-Aware Safety

  • 通过对比真实图与噪声图,动态识别视觉敏感词元。
  • 结合场景理解调节拒绝行为,提升安全判断准确性。
  • 不依赖特定模型,适用于各类多模态大模型。

多模态大语言模型(MLLMs)在现实应用中日益普及,但其做出上下文感知安全决策的能力仍有限。现有方法常面临过度敏感(无端拒绝良性请求)与敏感不足(未检测到视觉关联风险)的平衡难题,导致安全对齐存在持续缺口。为此,我们提出安全感知对比解码(SafeCoDe),一种轻量级、模型无关的解码框架,可基于多模态上下文动态调整词元生成。SafeCoDe分两阶段运行:(1) 对比解码机制,通过对比真实图像与高斯噪声图像,突出受视觉上下文影响的词元;(2) 全局感知词元调制策略,融合场景级推理与词元级调整,根据预测的安全判定结果自适应调整拒绝行为。在多种MLLM架构和安全基准上的广泛实验表明,SafeCoDe在敏感性、过度敏感性及通用安全性评估中均显著改善上下文感知拒绝行为,同时保持模型有用性。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet their ability to make context-aware safety decisions remains limited. Existing methods often fail to balance oversensitivity (unjustified refusals of benign queries) and undersensitivity (missed detection of visually grounded risks), leaving a persistent gap in safety alignment. To address this issue, we introduce Safety-aware Contrastive Decoding (SafeCoDe), a lightweight and model-agnostic decoding framework that dynamically adjusts token generation based on multimodal context. SafeCoDe operates in two stages: (1) a contrastive decoding mechanism that highlights tokens sensitive to visual context by contrasting real and Gaussian-noised images, and (2) a global-aware token modulation strategy that integrates scene-level reasoning with token-level adjustment to adapt refusals according to the predicted safety verdict. Extensive experiments across diverse MLLM architectures and safety benchmarks, covering undersensitivity, oversensitivity, and general safety evaluations, show that SafeCoDe consistently improves context-sensitive refusal behaviors while preserving model helpfulness.

多模态安全对齐解码优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。