构建多模态对话安全数据集,提出可适配策略的防护框架
LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
- 设计自动化生成恶意多轮多模态对话的红队框架
- 在4484条对话上验证,显著优于现有模型和工具
- 适合关注AI对话安全与内容审核的研究者
随着视觉语言模型(VLMs)向交互式多轮对话发展,多模态多轮对话的安全问题日益突出,其特征包括恶意意图隐蔽、风险上下文累积及跨模态联合风险。现有针对单轮或单模态的审核方法难以应对。为此,我们构建了多模态多轮对话安全(MMDS)数据集,包含4,484条标注对话,并建立涵盖8个主维度和60个子维度的风险分类体系。在数据构建过程中,提出多模态多轮红队(MMRT)框架,用于自动化生成不安全对话。进一步提出LLaVAShield,可在指定政策维度下审计用户输入与助手回复的安全性。大量实验表明,该方法显著优于当前最优的VLMs和内容审核工具,具备强泛化能力与灵活策略适配性。此外,我们分析主流VLMs对有害输入的脆弱性,并评估关键组件贡献,深化了对多模态多轮对话安全机制的理解。
原文摘要 · Abstract (English)
As Vision-Language Models (VLMs) move into interactive, multi-turn use, safety concerns intensify for multimodal multi-turn dialogue, which is characterized by concealment of malicious intent, contextual risk accumulation, and cross-modal joint risk. These characteristics limit the effectiveness of content moderation approaches designed for single-turn or single-modality settings. To address these limitations, we first construct the Multimodal Multi-turn Dialogue Safety (MMDS) dataset, comprising 4,484 annotated dialogues and a comprehensive risk taxonomy with 8 primary and 60 subdimensions. As part of MMDS construction, we introduce Multimodal Multi-turn Red Teaming (MMRT), an automated framework for generating unsafe multimodal multi-turn dialogues. We further propose LLaVAShield, which audits the safety of both user inputs and assistant responses under specified policy dimensions in multimodal multi-turn dialogues. Extensive experiments show that LLaVAShield significantly outperforms state-of-the-art VLMs and existing content moderation tools while demonstrating strong generalization and flexible policy adaptation. Additionally, we analyze vulnerabilities of mainstream VLMs to harmful inputs and evaluate the contribution of key components, advancing understanding of safety mechanisms in multimodal multi-turn dialogues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。