构建多模态安全评估数据集Falcon,提升大模型生成内容的安全检测能力。
Falcon: A Cross-Modal Evaluation Dataset for Comprehensive Safety Perception
- 构建包含5.7万组图文对的跨模态安全数据集,覆盖13类危害
- FalconEye评测器在多个基准上准确率领先,可识别复杂场景中的有害内容
- 适合大模型安全审计、内容过滤系统研发者使用
现有大语言模型内容危害性评估方法已较为成熟,但针对多模态大语言模型(MLLMs)的评估仍不充分且缺乏深度。本文强调视觉信息在视觉问答(VQA)中对内容安全判断的关键作用,这一维度常被忽视。为此,我们提出Falcon,一个大规模跨模态安全数据集,包含57,515个VQA对,覆盖13种危害类别。数据集对图像、指令和回复中的有害属性提供明确标注,并附带判断依据说明。此外,我们基于Qwen2.5-VL-7B微调得到FalconEye专用评测器。实验表明,FalconEye在自建Falcon-test数据集及广泛使用的VLGuard与Beavertail-V两个基准上均实现最高整体准确率,证明其在复杂多模态对话场景中可靠识别有害内容的能力,具备作为实际安全审计工具的潜力。
原文摘要 · Abstract (English)
Existing methods for evaluating the harmfulness of content generated by large language models (LLMs) have been well studied. However, approaches tailored to multimodal large language models (MLLMs) remain underdeveloped and lack depth. This work highlights the crucial role of visual information in moderating content in visual question answering (VQA), a dimension often overlooked in current research. To bridge this gap, we introduce Falcon, a large-scale vision-language safety dataset containing 57,515 VQA pairs across 13 harm categories. The dataset provides explicit annotations for harmful attributes across images, instructions, and responses, thereby facilitating a comprehensive evaluation of the content generated by MLLMs. In addition, it includes the relevant harm categories along with explanations supporting the corresponding judgments. We further propose FalconEye, a specialized evaluator fine-tuned from Qwen2.5-VL-7B using the Falcon dataset. Experimental results demonstrate that FalconEye reliably identifies harmful content in complex and safety-critical multimodal dialogue scenarios. It outperforms all other baselines in overall accuracy across our proposed Falcon-test dataset and two widely-used benchmarks-VLGuard and Beavertail-V, underscoring its potential as a practical safety auditing tool for MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。