arXiv:2504.11707cs.CVcs.AI2025-04被引 10

构建百万级多模态NSFW数据集并提出抗攻击防御模型,提升生成图像安全检测可靠性。

Towards Safe Synthetic Image Generation On the Web: A Multimodal Robust NSFW Defense and Million Scale Dataset

  • 基于开源扩散模型生成百万级图文对抗样本数据集
  • 新防御模型在对抗攻击下仍保持高准确率与召回率,显著降低攻击成功率
  • 适合需要安全可控图像生成的平台开发者和研究者使用

近年来,文本到图像(T2I)模型在互联网上广泛应用,其生成超逼真图像的能力也引发了对不适宜工作场合(NSFW)内容泛滥和社会污染的新担忧。尽管现有方法通过NSFW过滤器和事后安全检查来防范滥用,但近期研究揭示这些手段易被文本与图像模态的对抗攻击突破。当前缺乏包含提示词与图像对及对抗样例的鲁棒多模态NSFW数据集。本文提出一个基于开源扩散模型构建的百万规模图文数据集;同时设计一种对对抗攻击鲁棒的多模态防御机制,可有效区分安全与NSFW内容。大量实验表明,该模型在准确率和召回率上优于现有最先进方法,并在多模态对抗攻击场景中大幅降低攻击成功率达(ASR),代码已开源。

原文摘要 · Abstract (English)

In the past years, we have witnessed the remarkable success of Text-to-Image (T2I) models and their widespread use on the web. Extensive research in making T2I models produce hyper-realistic images has led to new concerns, such as generating Not-Safe-For-Work (NSFW) web content and polluting the web society. To help prevent misuse of T2I models and create a safer web environment for users features like NSFW filters and post-hoc security checks are used in these models. However, recent work unveiled how these methods can easily fail to prevent misuse. In particular, adversarial attacks on text and image modalities can easily outplay defensive measures. %Exploiting such leads to the growing concern of preventing adversarial attacks on text and image modalities. Moreover, there is currently no robust multimodal NSFW dataset that includes both prompt and image pairs and adversarial examples. This work proposes a million-scale prompt and image dataset generated using open-source diffusion models. Second, we develop a multimodal defense to distinguish safe and NSFW text and images, which is robust against adversarial attacks and directly alleviates current challenges. Our extensive experiments show that our model performs well against existing SOTA NSFW detection methods in terms of accuracy and recall, drastically reducing the Attack Success Rate (ASR) in multimodal adversarial attack scenarios. Code: https://github.com/shahidmuneer/multimodal-nsfw-defense.

图像安全对抗防御多模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。