首个音视频多模态安全评测基准,揭示大模型在跨模态攻击下的脆弱性。
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
- 构建24种模态组合的2.3万条测试样本,支持跨模态一致性评估
- 11个主流模型中仅3个在双指标上超过0.6,音视频输入下安全性能显著下降
- 提出针对性评分体系,适合研究多模态安全与对齐方法的学者
集成视觉、听觉与文本处理能力的多模态大语言模型(OLLMs)面临严重安全风险,其对音视频联合有害输入防御薄弱,且在不同模态间表现出不一致的安全性能,导致简单的模态切换即可实现越狱。现有安全评测基准因缺乏音视频联合样本、模态覆盖有限及缺少平行测试用例,难以全面评估此类风险。为此,我们提出Omni-SafetyBench,首个面向OLLMs的综合性并行评测基准,包含从972个种子样本衍生出的23,328个测试实例,涵盖24种模态变体。针对复杂输入的理解挑战与跨模态一致性的重要性,我们设计了基于条件攻击成功率(C-ASR)和条件拒绝率(C-RR)的安全得分,以及跨模态安全一致性得分(CMSC-score)。对11个前沿OLLMs的评估发现:仅有3个模型在两项指标上均超过0.6,音视频输入下的安全性急剧下降。进一步评估现有安全对齐方法,揭示了多模态安全对齐的根本性挑战,凸显该领域亟需深入研究。
原文摘要 · Abstract (English)
Omni-modal Large Language Models (OLLMs) that integrate visual, auditory, and textual processing face severe safety risks. They exhibit fragile defenses against audio-visual joint harmful inputs and demonstrate inconsistent safety performance across different modalities, enabling simple modality-switching jailbreaks. However, existing safety benchmarks fail to comprehensively assess these risks due to the absence of audio-visual joint samples, limited modality coverage, and lack of parallel test cases for cross-modal consistency evaluation. To address these gaps, we introduce Omni-SafetyBench, the first comprehensive parallel benchmark for OLLM safety evaluation, featuring 23,328 test instances across 24 modality variations derived from 972 seed samples. Recognizing that complex inputs pose comprehension challenges and that cross-modal consistency is critical for OLLM safety, we propose tailored metrics: a Safety-score based on Conditional Attack Success Rate (C-ASR) and Conditional Refusal Rate (C-RR), and a Cross-Modal Safety Consistency score (CMSC-score). Evaluating 11 state-of-the-art OLLMs reveals severe vulnerabilities: only 3 models exceed 0.6 in both metrics, with safety degrading sharply for audio-visual inputs. Furthermore, evaluation of existing safety alignment methods on Omni-SafetyBench identifies fundamental challenges in OLLM safety alignment, highlighting urgent needs for enhanced research in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。