用AI生成藏有仇恨信息的视觉错觉,现有审核模型几乎失效
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
- 用Stable Diffusion生成1860个视觉错觉,其中1571个成功隐藏仇恨信息
- 主流审核模型检测准确率低于24.5%,视觉语言模型低于10.2%
- 揭示模型只看表面细节,忽视隐藏信息,适合安全与内容审核研究者
文本到图像扩散模型的进展催生了新型数字艺术——视觉错觉,即通过视觉技巧制造对现实的不同感知。然而,攻击者可能滥用此类技术生成仇恨错觉,将特定仇恨信息嵌入看似无害的场景中,并在社交网络中传播。本文首次系统研究可扩展的仇恨错觉生成风险及其绕过现有内容审核模型的可能性。我们使用Stable Diffusion与ControlNet,基于62条仇恨信息生成1,860个光学幻象,其中1,571个成功嵌入仇恨内容,形成「仇恨错觉数据集」。利用该数据集,评估六种审核分类器和九种视觉语言模型(VLMs)对仇恨错觉的识别能力。实验结果表明现有模型存在显著漏洞:审核分类器检测准确率低于0.245,VLMs低于0.102。进一步分析发现,其视觉编码器主要关注图像表层细节,忽视了隐藏的信息层。为此,我们探索了图像变换与训练策略等初步缓解措施,识别出最有效的应对路径。
原文摘要 · Abstract (English)
Recent advances in text-to-image diffusion models have enabled the creation of a new form of digital art: optical illusions--visual tricks that create different perceptions of reality. However, adversaries may misuse such techniques to generate hateful illusions, which embed specific hate messages into harmless scenes and disseminate them across web communities. In this work, we take the first step toward investigating the risks of scalable hateful illusion generation and the potential for bypassing current content moderation models. Specifically, we generate 1,860 optical illusions using Stable Diffusion and ControlNet, conditioned on 62 hate messages. Of these, 1,571 are hateful illusions that successfully embed hate messages, either overtly or subtly, forming the Hateful Illusion dataset. Using this dataset, we evaluate the performance of six moderation classifiers and nine vision language models (VLMs) in identifying hateful illusions. Experimental results reveal significant vulnerabilities in existing moderation models: the detection accuracy falls below 0.245 for moderation classifiers and below 0.102 for VLMs. We further identify a critical limitation in their vision encoders, which mainly focus on surface-level image details while overlooking the secondary layer of information, i.e., hidden messages. To address this risk, we explore preliminary mitigation measures and identify the most effective approaches from the perspectives of image transformations and training-level strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。