提出新模型ShieldVLM,用多模态推理识别图文隐性毒性
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
- 通过跨模态反思推理检测图文组合中的隐性毒性
- 在2100条样本上实现比基线更高的隐性和显性毒性检测准确率
- 适合安全研究者、AI伦理团队及内容审核系统开发者
多模态文本-图像内容的毒性检测面临日益严峻挑战,尤其是多模态隐性毒性——各模态单独看似无害,但结合后却蕴含危险。此类毒性不仅出现在社交平台的公开表述中,也可能作为提示引发大型视觉语言模型(LVLMs)生成有害对话。尽管单模态文本或图像的过滤已取得进展,多模态隐性毒性检测仍严重不足。为此,我们系统构建了多模态隐性毒性(MMIT)分类体系,创建了一个包含2100个多模态语句与提示的MMIT数据集,涵盖7个风险类别(31个子类别)和5种典型跨模态关联模式。为推进该任务,我们提出ShieldVLM,一种通过反思式跨模态推理识别多模态陈述、提示及对话中隐性毒性的模型。实验表明,ShieldVLM在隐性和显性毒性检测上均优于现有强基线。模型与数据集将公开,以支持后续研究。警告:本文包含潜在敏感内容。
原文摘要 · Abstract (English)
Toxicity detection in multimodal text-image content faces growing challenges, especially with multimodal implicit toxicity, where each modality appears benign on its own but conveys hazard when combined. Multimodal implicit toxicity appears not only as formal statements in social platforms but also prompts that can lead to toxic dialogs from Large Vision-Language Models (LVLMs). Despite the success in unimodal text or image moderation, toxicity detection for multimodal content, particularly the multimodal implicit toxicity, remains underexplored. To fill this gap, we comprehensively build a taxonomy for multimodal implicit toxicity (MMIT) and introduce an MMIT-dataset, comprising 2,100 multimodal statements and prompts across 7 risk categories (31 sub-categories) and 5 typical cross-modal correlation modes. To advance the detection of multimodal implicit toxicity, we build ShieldVLM, a model which identifies implicit toxicity in multimodal statements, prompts and dialogs via deliberative cross-modal reasoning. Experiments show that ShieldVLM outperforms existing strong baselines in detecting both implicit and explicit toxicity. The model and dataset will be publicly available to support future researches. Warning: This paper contains potentially sensitive contents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。