arXiv:2601.14127cs.CVcs.CL2026-01ACL被引 3

多图推理越强的AI模型,反而越容易产生安全问题。

The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning

  • 构建首个针对多图推理安全的评测基准MIR-SafetyBench。
  • 19个模型中,推理能力越强,安全漏洞率越高。
  • 模型常以回避回答伪装安全,实际风险隐匿其中。

随着多模态大语言模型(MLLMs)在复杂多图指令下推理能力增强,其潜在安全风险也日益凸显。本文提出MIR-SafetyBench,首个专注于多图推理安全的评测基准,包含2,676个实例,涵盖9类多图关系。对19个MLLM的全面评估显示:推理能力越强的模型,在MIR-SafetyBench上表现越差,安全风险更高。除攻击成功率外,发现大量被标记为“安全”的回应实为表面化、回避性或误解所致。进一步分析表明,不安全生成的注意力熵普遍低于安全生成,提示模型可能因过度专注任务完成而忽略安全约束。代码与数据已公开于https://github.com/thu-coai/MIR-SafetyBench。

原文摘要 · Abstract (English)

As Multimodal Large Language Models (MLLMs) acquire stronger reasoning capabilities to handle complex, multi-image instructions, this advancement may pose new safety risks. We study this problem by introducing MIR-SafetyBench, the first benchmark focused on multi-image reasoning safety, which consists of 2,676 instances across a taxonomy of 9 multi-image relations. Our extensive evaluations on 19 MLLMs reveal a troubling trend: models with more advanced multi-image reasoning can be more vulnerable on MIR-SafetyBench. Beyond attack success rates, we find that many responses labeled as safe are superficial, often driven by misunderstanding or evasive, non-committal replies. We further observe that unsafe generations exhibit lower attention entropy than safe ones on average. This internal signature suggests a possible risk that models may over-focus on task solving while neglecting safety constraints. Our code and data are available at https://github.com/thu-coai/MIR-SafetyBench.

多模态安全模型风险推理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。