首个面向多模态大模型的安全检测基准,覆盖文本、图像、音频、视频四类风险内容。
OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models
- 构建涵盖中英文的多模态安全数据集,含1.8万条文本、4500张图等。
- 提出跨风险评分(MCRS)与可解释评估框架,提升检测可靠性。
- 发现9个主流多模态模型普遍存在安全漏洞,适合安全研究者参考。
随着多模态大语言模型(MLLMs)日益融入日常工具与智能代理,其生成不安全内容的风险引发广泛关注,包括有毒语言、偏见图像、隐私泄露及有害信息传播。现有安全评测基准在模态覆盖和评估维度上仍显不足,难以全面反映内容安全问题。本文提出OutSafe-Bench,首个专为多模态时代设计的综合性内容安全评测基准。该基准包含大规模数据集,覆盖四种模态:超过18,000条中英文文本提示、4,500张图像、450段音频及450个视频,均在九类关键内容风险类别中系统标注。除数据集外,我们引入多维交叉风险评分(MCRS),用于建模和评估不同风险类别间的重叠与相关性。为确保评估公平可靠,提出FairScore——一种可解释的自动化多评审加权聚合框架,通过自适应选择表现最佳模型作为评审团,减少单一模型判断偏差,增强评估可信度。对九个先进MLLMs的评估显示,这些模型普遍存有持续且显著的安全缺陷,凸显强化防护机制的紧迫性。
原文摘要 · Abstract (English)
Since Multimodal Large Language Models (MLLMs) are increasingly being integrated into everyday tools and intelligent agents, growing concerns have arisen regarding their possible output of unsafe contents, ranging from toxic language and biased imagery to privacy violations and harmful misinformation. Current safety benchmarks remain highly limited in both modality coverage and performance evaluations, often neglecting the extensive landscape of content safety. In this work, we introduce OutSafe-Bench, the first most comprehensive content safety evaluation test suite designed for the multimodal era. OutSafe-Bench includes a large-scale dataset that spans four modalities, featuring over 18,000 bilingual (Chinese and English) text prompts, 4,500 images, 450 audio clips and 450 videos, all systematically annotated across nine critical content risk categories. In addition to the dataset, we introduce a Multidimensional Cross Risk Score (MCRS), a novel metric designed to model and assess overlapping and correlated content risks across different categories. To ensure fair and robust evaluation, we propose FairScore, an explainable automated multi-reviewer weighted aggregation framework. FairScore selects top-performing models as adaptive juries, thereby mitigating biases from single-model judgments and enhancing overall evaluation reliability. Our evaluation of nine state-of-the-art MLLMs reveals persistent and substantial safety vulnerabilities, underscoring the pressing need for robust safeguards in MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。