构建多模态安全评估基准,揭示大模型跨模态推理短板
MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

- 设计1196个需融合视觉语音文本的多模态安全场景
- 模型对细微或非物理风险识别率低,依赖明显视觉/听觉线索
- 适合关注多模态安全与模型可信赖性的研究者
现有多模态安全评测仅针对视觉输入,无法评估同时处理视觉、音频和文本的Omni大语言模型(LLM)。我们提出MCBench,一个包含1196个场景的基准测试,覆盖四个安全类别,需融合多种模态才能准确评估安全风险。每个不安全场景均配有最小差异的安全对照场景,用于评估模型敏感性。对先进模型的评估显示显著挑战:Omni LLM在面对细微或非物理风险时表现不佳,但在存在显著视觉或听觉线索时表现更好。分析推理轨迹发现,尽管模型能提取各模态信息,却常无法有效整合以做出安全判断。结果表明当前Omni LLM在安全关键场景中缺乏稳健的跨模态推理能力,凸显了改进架构与训练策略的必要性。
原文摘要 · Abstract (English)
Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and text. We introduce MCBench, a benchmark with 1196 scenarios spanning four safety categories that require integrating multiple modalities for accurate safety assessment. Each unsafe scenario is paired with a minimally different safe counterpart to assess model sensitivity. Our evaluations of state-of-the-art models reveal significant challenges. Omni LLMs struggle with subtle or non-physical risks but perform better when salient visual or acoustic cues are present. Analysis of reasoning traces shows that, although models can extract modality-specific information, they often fail to integrate these cues effectively for safety judgments. Our findings reveal that current Omni LLMs lack robust cross-modal reasoning in safety-critical settings, underscoring the need for improved architectures and training strategies for multimodal safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。