首个统一多模态大模型安全评测基准,揭示融合模型安全性能下降问题
Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models

- 构建六类安全维度的统一多模态安全评测框架
- 发现统一模型无法保持原始大模型的安全对齐,生成任务安全性显著降低
- 开源统一模型安全表现远低于专用模型,适合关注多模态安全的研究者
统一多模态大模型(UMLMs)将理解与生成能力集成于单一架构中。尽管统一架构拓展了多模态能力,其安全影响仍重要但研究不足。现有安全评测主要针对孤立的理解或生成任务,难以评估统一框架下多样任务的综合安全性。为此,我们提出Uni-SafeBench,一个涵盖七种任务类型的综合性评测基准,包含六大安全类别。为实现严谨评估,我们开发Uni-Judger框架,有效解耦上下文安全与内在安全。基于对Uni-SafeBench的全面评估,发现当前统一模型未能一致保留底层语言模型的原始安全对齐;此外,开源统一模型在生成任务上的安全表现显著低于专门用于生成或理解的多模态大模型。
原文摘要 · Abstract (English)
Unified Multimodal Large Models (UMLMs) integrate understanding and generation capabilities within a single architecture. While unified architectures expand multimodal capabilities, their safety implications remain important yet underexplored. Existing safety benchmarks predominantly focus on isolated understanding or generation tasks, failing to evaluate the holistic safety of UMLMs when handling diverse tasks under a unified framework. To address this, we introduce Uni-SafeBench, a comprehensive benchmark featuring a taxonomy of six major safety categories across seven task types. To ensure rigorous assessment, we develop Uni-Judger, a framework that effectively decouples contextual safety from intrinsic safety. Based on comprehensive evaluations across Uni-SafeBench, we find that the original safety alignment of the underlying LLM is not consistently preserved in current unified models. Moreover, open-source UMLMs exhibit much lower safety performance than multimodal large models specialized for either generation or understanding tasks, particularly on the generation side.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。