文本到图像模型易被多模态隐性攻击生成违规内容
Multimodal Pragmatic Jailbreak on Text-to-image Models
- 通过图文结合的隐性提示触发模型生成违规图像
- 9个主流模型均遭攻击,违规生成率10%至70%
- 传统过滤器对这种多模态攻击无效,适合安全研究者
扩散模型在图像质量和文本一致性上取得显著进展,但其安全性日益受到关注。本文提出一种新型越狱攻击,通过在文本提示中嵌入视觉文字,使图像与文字各自看似安全,组合后却生成违规内容。为系统研究该现象,我们构建了评估数据集,并对九个代表性文本到图像(T2I)模型进行了基准测试,包括两个闭源商业模型。实验结果揭示令人担忧的趋势:所有测试模型均受此类越狱影响,违规生成率在10%至70%之间,其中DALL·E 3表现最差。在实际场景中,常用关键词黑名单、自定义提示过滤和NSFW图像过滤等手段被用于风险缓解,但本研究发现这些过滤机制对本攻击无效。我们从文本渲染能力和训练数据角度分析了越狱成因。本工作为构建更安全可靠的T2I模型提供了基础。
原文摘要 · Abstract (English)
Diffusion models have recently achieved remarkable advancements in terms of image quality and fidelity to textual prompts. Concurrently, the safety of such generative models has become an area of growing concern. This work introduces a novel type of jailbreak, which triggers T2I models to generate the image with visual text, where the image and the text, although considered to be safe in isolation, combine to form unsafe content. To systematically explore this phenomenon, we propose a dataset to evaluate the current diffusion-based text-to-image (T2I) models under such jailbreak. We benchmark nine representative T2I models, including two closed-source commercial models. Experimental results reveal a concerning tendency to produce unsafe content: all tested models suffer from such type of jailbreak, with rates of unsafe generation ranging from around 10\% to 70\% where DALLE 3 demonstrates almost the highest unsafety. In real-world scenarios, various filters such as keyword blocklists, customized prompt filters, and NSFW image filters, are commonly employed to mitigate these risks. We evaluate the effectiveness of such filters against our jailbreak and found that, while these filters may be effective for single modality detection, they fail to work against our jailbreak. We also investigate the underlying reason for such jailbreaks, from the perspective of text rendering capability and training data. Our work provides a foundation for further development towards more secure and reliable T2I models. Project page at https://multimodalpragmatic.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。