构建首个统一评估多模态大模型安全性的综合基准
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
- 设计高质合成数据管道,覆盖61类风险与4种模态组合
- 首次同时评估模型漏洞与过度敏感性,结果更真实可靠
- 涵盖中英文双语,适合安全研究者与模型开发者使用
尽管多模态大语言模型(MLLMs)取得了显著成就并广泛应用,其安全漏洞仍暴露严重隐患,亟需可靠的评估基准。现有安全评测基准在数据质量、覆盖范围和模态风险组合方面存在不足,导致评估结果虚高且矛盾,阻碍了安全问题的发现与治理。本文指出,有害查询的漏洞与无害查询的过度敏感应同步评估,而此前二者被分开处理。为此,我们提出统一安全评测基准(USB),是目前最全面的MLLM安全评估框架。该基准包含高质量查询、61个风险子类别、4种模态组合,并覆盖中英文及漏洞与过度敏感双重维度。通过对比发现,现有基准即使合并也远未全面。我们设计复杂的数据合成流程,生成大量填补空白的互补数据。结合开源数据与合成数据,实现每个风险子类别对应4种模态组合,推动更系统、更可信的安全评估。
原文摘要 · Abstract (English)
Despite their remarkable achievements and widespread adoption, Multimodal Large Language Models (MLLMs) have revealed significant security vulnerabilities, highlighting the urgent need for robust safety evaluation benchmarks. Existing MLLM safety benchmarks, however, fall short in terms of data quality and coverge, and modal risk combinations, resulting in inflated and contradictory evaluation results, which hinders the discovery and governance of security concerns. Besides, we argue that vulnerabilities to harmful queries and oversensitivity to harmless ones should be considered simultaneously in MLLMs safety evaluation, whereas these were previously considered separately. In this paper, to address these shortcomings, we introduce Unified Safety Benchmarks (USB), which is one of the most comprehensive evaluation benchmarks in MLLM safety. Our benchmark features high-quality queries, extensive risk categories, comprehensive modal combinations, and encompasses both vulnerability and oversensitivity evaluations. From the perspective of two key dimensions: risk categories and modality combinations, we demonstrate that the available benchmarks -- even the union of the vast majority of them -- are far from being truly comprehensive. To bridge this gap, we design a sophisticated data synthesis pipeline that generates extensive, high-quality complementary data addressing previously unexplored aspects. By combining open-source datasets with our synthetic data, our benchmark provides 4 distinct modality combinations for each of the 61 risk sub-categories, covering both English and Chinese across both vulnerability and oversensitivity dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。