arXiv:2602.03263cs.AI2026-02被引 3

测试多模态大模型跨模态安全性的基准,发现模型常依赖语言主导而非真实理解图文联合意图。

CSR-Bench: A Benchmark for Evaluating the Cross-modal Safety and Reliability of MLLMs

  • 设计四类压力测试场景,强制模型融合图文信息判断安全行为
  • 16个主流模型普遍存在安全意识弱、文本干扰下表现下降等缺陷
  • 揭示安全提升可能源于拒绝策略而非真正理解,适合研究安全对齐的学者

多模态大语言模型(MLLMs)能处理文本与图像的交互,但其安全行为可能受单模态捷径驱动,而非真正的跨模态理解。我们提出CSR-Bench,一个通过四种压力测试交互模式评估跨模态可靠性的基准,涵盖安全、过度拒绝、偏见和幻觉,共61种细粒度类型。每项任务均需整合图像与文本理解,同时提供配对纯文本控制组,以诊断模态引发的行为变化。我们评估了16个前沿MLLMs,发现系统性跨模态对齐差距:模型安全意识薄弱,在文本干扰下语言主导明显,且从纯文本到多模态输入性能持续下降。还观察到减少过度拒绝与保持安全非歧视行为之间存在明显权衡,表明部分看似安全的提升可能来自拒绝导向的启发式策略,而非稳健的意图理解。警告:本文包含不安全内容。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) enable interaction over both text and images, but their safety behavior can be driven by unimodal shortcuts instead of true joint intent understanding. We introduce CSR-Bench, a benchmark for evaluating cross-modal reliability through four stress-testing interaction patterns spanning Safety, Over-rejection, Bias, and Hallucination, covering 61 fine-grained types. Each instance is constructed to require integrated image-text interpretation, and we additionally provide paired text-only controls to diagnose modality-induced behavior shifts. We evaluate 16 state-of-the-art MLLMs and observe systematic cross-modal alignment gaps. Models show weak safety awareness, strong language dominance under interference, and consistent performance degradation from text-only controls to multimodal inputs. We also observe a clear trade-off between reducing over-rejection and maintaining safe, non-discriminatory behavior, suggesting that some apparent safety gains may come from refusal-oriented heuristics rather than robust intent understanding. WARNING: This paper contains unsafe contents.

多模态安全模型评估可靠性测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。