ProGuard能主动识别未知安全风险,无需模型调整。
ProGuard: Towards Proactive Multimodal Safeguard
- 基于强化学习训练多模态模型,实现高效推理。
- 在未知风险检测上提升52.6%,描述能力提升64.8%。
- 适合需要主动防护的生成式AI安全场景。
生成模型的快速发展带来了持续涌现的多模态安全风险,现有防御方法面临局限。为此,我们提出ProGuard——一种视觉语言主动防护机制,可在不调整模型的前提下识别并描述分布外(OOD)安全风险。我们构建了一个包含87,000样本的模态平衡数据集,每个样本同时标注二元安全标签与层级化多模态安全类别,有效缓解模态偏差,确保文本、图像及图文输入的一致性审核。基于该数据集,我们纯通过强化学习训练视觉语言基础模型,实现高效简洁的推理。为在可控环境中模拟主动防护场景,我们引入分布外安全类别推理任务,并在强化学习目标中加入基于同义词库的相似性奖励,促使模型对未见的不安全类别生成简洁描述。实验表明,ProGuard在二元安全分类上性能接近闭源大模型,在不安全内容分类上显著优于现有开源防护模型。最突出的是,ProGuard展现出强大的主动防护能力,使分布外风险检测提升52.6%,描述能力提升64.8%。
原文摘要 · Abstract (English)
The rapid evolution of generative models has led to a continuous emergence of multimodal safety risks, exposing the limitations of existing defense methods. To address these challenges, we propose ProGuard, a vision-language proactive guard that identifies and describes out-of-distribution (OOD) safety risks without the need for model adjustments required by traditional reactive approaches. We first construct a modality-balanced dataset of 87K samples, each annotated with both binary safety labels and risk categories under a hierarchical multimodal safety taxonomy, effectively mitigating modality bias and ensuring consistent moderation across text, image, and text-image inputs. Based on this dataset, we train our vision-language base model purely through reinforcement learning (RL) to achieve efficient and concise reasoning. To approximate proactive safety scenarios in a controlled setting, we further introduce an OOD safety category inference task and augment the RL objective with a synonym-bank-based similarity reward that encourages the model to generate concise descriptions for unseen unsafe categories. Experimental results show that ProGuard achieves performance comparable to closed-source large models on binary safety classification, substantially outperforms existing open-source guard models on unsafe content categorization. Most notably, ProGuard delivers a strong proactive moderation ability, improving OOD risk detection by 52.6% and OOD risk description by 64.8%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。