提出多模态安全评估框架,揭示大模型在图文联合理解中的致命缺陷。
VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
- 构建17种安全模式的细粒度分类体系,通过真实图像与人工标注评估
- 模型在图文联合判断时准确率降至20%-55%,远低于单模态90%以上
- 34%错误源于单模态正确但组合推理失败,适合安全评测与模型改进研究
多模态大模型的安全评估常将视觉与语言输入分开处理,忽视了图文结合后良性内容可能变得有害的风险。现有方法难以区分明确危险与边界案例,导致过度拦截或漏判。我们提出视觉语言安全理解(VLSU)框架,通过细粒度严重性分类和17种安全模式的组合分析,系统评估多模态安全。基于真实图像与人工标注,构建包含8,187个样本、覆盖15类危害的大规模基准。对11个顶尖模型的评估显示:尽管单模态安全信号识别准确率达90%以上,但在需要图文联合推理时准确率暴跌至20%-55%。最严重的是,34%的错误发生在单模态分类正确的情况下,表明模型缺乏组合推理能力。此外,模型难以平衡拒绝危险内容与回应边界案例——例如,指令表述可使Gemini-1.5对边界内容的过度拦截率从62.4%降至10.4%,但同时使危险内容的拒绝率从90.8%降至53.9%。该框架揭示了当前模型在图文联合理解与对齐上的系统性缺陷,为推进鲁棒多模态安全研究提供关键测试基准。
原文摘要 · Abstract (English)
Safety evaluation of multimodal foundation models often treats vision and language inputs separately, missing risks from joint interpretation where benign content becomes harmful in combination. Existing approaches also fail to distinguish clearly unsafe content from borderline cases, leading to problematic over-blocking or under-refusal of genuinely harmful content. We present Vision Language Safety Understanding (VLSU), a comprehensive framework to systematically evaluate multimodal safety through fine-grained severity classification and combinatorial analysis across 17 distinct safety patterns. Using a multi-stage pipeline with real-world images and human annotation, we construct a large-scale benchmark of 8,187 samples spanning 15 harm categories. Our evaluation of eleven state-of-the-art models reveals systematic joint understanding failures: while models achieve 90%-plus accuracy on clear unimodal safety signals, performance degrades substantially to 20-55% when joint image-text reasoning is required to determine the safety label. Most critically, 34% of errors in joint image-text safety classification occur despite correct classification of the individual modalities, further demonstrating absent compositional reasoning capabilities. Additionally, we find that models struggle to balance refusing unsafe content while still responding to borderline cases that deserve engagement. For example, we find that instruction framing can reduce the over-blocking rate on borderline content from 62.4% to 10.4% in Gemini-1.5, but only at the cost of under-refusing on unsafe content with refusal rate dropping from 90.8% to 53.9%. Overall, our framework exposes weaknesses in joint image-text understanding and alignment gaps in current models, and provides a critical test bed to enable the next milestones in research on robust vision-language safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。