构建统一评测平台,系统评估多模态大模型的越狱攻击与防御能力。
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
- 整合13种越狱攻击与15种防御策略,覆盖9大风险领域。
- 在10个开源和8个闭源模型上测试,揭示多模态模型普遍易受攻击。
- 提供可复现的评测工具,适合安全研究者和模型开发者使用。
多模态大语言模型虽具备统一感知与推理能力,但仍易受越狱攻击,导致有害行为。现有基准如JailBreakV-28K、MM-SafetyBench和HADES多聚焦于有限攻击场景,缺乏标准化防御评估,且无统一可复现工具。为此,我们提出OmniSafeBench-MM,一个涵盖13种代表性攻击方法、15种防御策略的综合性评测工具箱,数据集覆盖9大风险领域、50个细粒度类别,按咨询、指令、陈述三类提问类型设计,反映真实用户意图。评测采用三维标准:(1)危害性,采用从个体轻微伤害到社会级灾难的多层级量表;(2)回复与查询意图对齐程度;(3)回复详尽程度,支持安全与效用权衡分析。我们在10个开源与8个闭源多模态模型上开展实验,揭示其普遍脆弱性。通过整合数据、方法与评估,OmniSafeBench-MM为未来研究提供开放、可复现的标准化基础。代码已开源:https://github.com/jiaxiaojunQAQ/OmniSafeBench-MM。
原文摘要 · Abstract (English)
Recent advances in multi-modal large language models (MLLMs) have enabled unified perception-reasoning capabilities, yet these systems remain highly vulnerable to jailbreak attacks that bypass safety alignment and induce harmful behaviors. Existing benchmarks such as JailBreakV-28K, MM-SafetyBench, and HADES provide valuable insights into multi-modal vulnerabilities, but they typically focus on limited attack scenarios, lack standardized defense evaluation, and offer no unified, reproducible toolbox. To address these gaps, we introduce OmniSafeBench-MM, which is a comprehensive toolbox for multi-modal jailbreak attack-defense evaluation. OmniSafeBench-MM integrates 13 representative attack methods, 15 defense strategies, and a diverse dataset spanning 9 major risk domains and 50 fine-grained categories, structured across consultative, imperative, and declarative inquiry types to reflect realistic user intentions. Beyond data coverage, it establishes a three-dimensional evaluation protocol measuring (1) harmfulness, distinguished by a granular, multi-level scale ranging from low-impact individual harm to catastrophic societal threats, (2) intent alignment between responses and queries, and (3) response detail level, enabling nuanced safety-utility analysis. We conduct extensive experiments on 10 open-source and 8 closed-source MLLMs to reveal their vulnerability to multi-modal jailbreak. By unifying data, methodology, and evaluation into an open-source, reproducible platform, OmniSafeBench-MM provides a standardized foundation for future research. The code is released at https://github.com/jiaxiaojunQAQ/OmniSafeBench-MM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。