提出新安全评估框架,识别多模态模型隐性风险后果
OOD-MMSafe: Advancing MLLM Safety from Harmful Intent to Hidden Consequences
- 构建因果链风险评估基准,检测模型对潜在危害的感知能力
- 高容量模型失败率高达67.5%,静态对齐导致安全推理瓶颈
- 引入动态自蒸馏机制,使模型在复杂场景中风险识别准确率提升至94%以上
尽管多模态大模型(MLLM)的安全对齐已受到广泛关注,现有范式主要针对恶意意图或情境违规。本文提出将安全边界转向后果驱动型安全,这对自主与具身智能体的稳健部署至关重要。为此,我们构建了包含455个精心筛选的图文对的基准OOD-MMSafe,用于评估模型在上下文依赖的因果链中识别潜在危害的能力。分析发现,前沿模型普遍存在因果盲视现象,最高67.5%的失败率出现在高容量闭源模型中,并揭示静态对齐存在性能天花板,随着模型容量增长,其表现趋于格式化而非真正提升安全推理。为突破此瓶颈,我们提出后果感知安全策略优化(CASPO)框架,将模型内在推理作为动态参考,实现分词级自蒸馏奖励。实验表明,CASPO显著增强后果预测能力,使Qwen2.5-VL-7B和Qwen3-VL-4B的风险识别失败率分别降至7.3%和5.7%,同时保持整体有效性。
原文摘要 · Abstract (English)
While safety alignment for Multimodal Large Language Models (MLLMs) has gained significant attention, current paradigms primarily target malicious intent or situational violations. We propose shifting the safety frontier toward consequence-driven safety, a paradigm essential for the robust deployment of autonomous and embodied agents. To formalize this shift, we introduce OOD-MMSafe, a benchmark comprising 455 curated query-image pairs designed to evaluate a model's ability to identify latent hazards within context-dependent causal chains. Our analysis reveals a pervasive causal blindness among frontier models, with the highest 67.5% failure rate in high-capacity closed-source models, and identifies a preference ceiling where static alignment yields format-centric failures rather than improved safety reasoning as model capacity grows. To address these bottlenecks, we develop the Consequence-Aware Safety Policy Optimization (CASPO) framework, which integrates the model's intrinsic reasoning as a dynamic reference for token-level self-distillation rewards. Experimental results demonstrate that CASPO significantly enhances consequence projection, reducing the failure ratio of risk identification to 7.3% for Qwen2.5-VL-7B and 5.7% for Qwen3-VL-4B while maintaining overall effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。