发现图片组合隐含毒性,提出可解释的检测方法
Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine

- 构建多图联合分析框架,识别单图无害但合起来有害的情况
- 在7类风险上训练模型,80亿参数版本超越主流平台
- 输出带推理过程的安全判断,适合内容审核系统使用
多图内容在社交媒体中日益普遍,催生了一种新安全问题——多图隐含毒性(MIIT):单张图像看似无害,但组合后产生有害语义。现有商业审核系统因缺乏单图中的明确风险线索而难以应对。本文首次正式定义MIIT,并分析其三大检测挑战。为缓解数据稀缺问题,构建了MIIT-dataset,一个仅含图像的多图安全数据集,覆盖七类典型风险,通过自动化生成流程构建。最后,采用渐进式蒸馏推理监督训练MiShield模型,使其在做出安全判断时能提供涉及关联实体的明确分析。实验表明,MiShield-8B模型在性能上超越代表性审核服务和更大规模模型,展现出在该广泛视觉格式中的有效性与实用价值。警告:本文包含潜在敏感内容。
原文摘要 · Abstract (English)
Multi-image content has become an increasingly prevalent form of visual communication in social media, giving rise to a new safety issue, multi-image implicit toxicity (MIIT), where each image appears benign in isolation, but harmful semantics emerge when the images are interpreted jointly. MIIT is particularly challenging for existing commercial moderation APIs and models due to the lack of explicit risky cues in each image. This paper aims to study how to identify MIIT. We first provide a formal definition of MIIT and analyze three key challenges for its detection. To alleviate the scarcity of data in this area, we construct MIIT-dataset, an image-only multi-image safety dataset covering seven representative risk categories through an automatic generation pipeline. Finally, we train MiShield with progressively distilled reasoning supervision, enabling it to produce safety judgments accompanied by explicit analyses of the correlated entities that result in the hazards. Experiments show that MiShield-8B models outperform representative moderation services and even larger-scale models, revealing its effectiveness and practical value for this widely used visual format. Warning: This paper contains potentially sensitive content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。