用隐藏层特征实时检测文本生成图像中的不当内容,准确率超95%。
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
- 从文本编码器隐状态提取有害特征,利用可分离性实现高效检测
- 在多个数据集上准确率超95%,推理效率提升十倍以上
- 支持实时解释与优化,适用于多种模型架构和少样本场景
随着文本到图像(T2I)模型的快速发展与广泛应用,其安全问题日益突出。恶意用户通过有害或对抗性提示生成不适宜工作场合(NSFW)图像,亟需有效防护机制保障输出合规。现有检测方法普遍存在准确率低、效率差等问题。本文提出AEIOU,一种自适应、高效、可解释、可优化且统一的防御框架,用于应对T2I模型中的NSFW提示。该框架从模型文本编码器的隐藏状态中提取NSFW特征,利用这些特征的可分离性实现快速检测,推理开销极小。同时支持实时结果解释,并可通过数据增强进行优化。框架具有高度通用性,适配多种T2I架构。大量实验表明,AEIOU显著优于商业与开源内容审核工具,在所有测试数据集上准确率均超过95%,效率提升至少十倍,对自适应攻击具有强鲁棒性,且在少样本与多标签场景下表现优异。
原文摘要 · Abstract (English)
As text-to-image (T2I) models advance and gain widespread adoption, their associated safety concerns are becoming increasingly critical. Malicious users exploit these models to generate Not-Safe-for-Work (NSFW) images using harmful or adversarial prompts, underscoring the need for effective safeguards to ensure the integrity and compliance of model outputs. However, existing detection methods often exhibit low accuracy and inefficiency. In this paper, we propose AEIOU, a defense framework that is adaptable, efficient, interpretable, optimizable, and unified against NSFW prompts in T2I models. AEIOU extracts NSFW features from the hidden states of the model's text encoder, utilizing the separable nature of these features to detect NSFW prompts. The detection process is efficient, requiring minimal inference time. AEIOU also offers real-time interpretation of results and supports optimization through data augmentation techniques. The framework is versatile, accommodating various T2I architectures. Our extensive experiments show that AEIOU significantly outperforms both commercial and open-source moderation tools, achieving over 95\% accuracy across all datasets and improving efficiency by at least tenfold. It effectively counters adaptive attacks and excels in few-shot and multi-label scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。