首个统一多模态模型安全评估基准,覆盖7种模态组合
UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
- 设计共享目标框架,跨任务对比多模态安全风险
- 评估15个主流模型,发现多图合成与多轮对话中违规率更高
- 适用于研究多模态系统安全的学者与工程师
统一多模态模型(UMMs)具备强大的跨模态能力,但引入了单任务模型中未见的新安全风险。尽管其迅速发展,现有安全评估基准仍分散于不同任务与模态,难以全面评估复杂系统的潜在漏洞。为填补这一空白,我们提出UniSAFE,首个面向多模态模型系统级安全评估的综合性基准,涵盖7种输入/输出模态组合,覆盖传统任务与新型多模态上下文图像生成场景。UniSAFE采用共享目标设计,将共性风险场景映射到特定任务的输入输出配置中,实现受控的跨任务安全失败比较。该基准包含6,802个精心筛选的实例,我们用其评估了15个最先进的UMM(包括专有和开源模型)。结果揭示当前模型存在严重漏洞,尤其在多图合成与多轮交互设置下安全违规显著增加,图像输出任务普遍比文本输出任务更易出错。这些发现凸显了加强多模态系统级安全对齐的迫切需求。代码与数据已公开于https://github.com/segyulee/UniSAFE。
原文摘要 · Abstract (English)
Unified Multimodal Models (UMMs) offer powerful cross-modality capabilities but introduce new safety risks not observed in single-task models. Despite their emergence, existing safety benchmarks remain fragmented across tasks and modalities, limiting the comprehensive evaluation of complex system-level vulnerabilities. To address this gap, we introduce UniSAFE, the first comprehensive benchmark for system-level safety evaluation of UMMs across 7 I/O modality combinations, spanning conventional tasks and novel multimodal-context image generation settings. UniSAFE is built with a shared-target design that projects common risk scenarios across task-specific I/O configurations, enabling controlled cross-task comparisons of safety failures. Comprising 6,802 curated instances, we use UniSAFE to evaluate 15 state-of-the-art UMMs, both proprietary and open-source. Our results reveal critical vulnerabilities across current UMMs, including elevated safety violations in multi-image composition and multi-turn settings, with image-output tasks consistently more vulnerable than text-output tasks. These findings highlight the need for stronger system-level safety alignment for UMMs. Our code and data are publicly available at https://github.com/segyulee/UniSAFE
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。