自动检测文本生成图像模型的隐性安全漏洞,提升内容安全性。
Automatic Red Teaming for Implicit Vulnerabilities of Text-to-Image Models

- 用双智能体框架动态生成并优化隐蔽有害提示。
- 在多个模型上发现隐性漏洞,效果优于现有方法。
- 无需模型参数即可评估,适合安全审计与部署验证。
对文本生成图像(T2I)模型进行红队测试对安全部署至关重要,但针对隐性对抗提示仍具挑战。与可直接识别的显式对抗提示不同,隐性提示在表面上看似无害,却仍生成不当视觉内容。为此,我们提出对抗探测隐性漏洞框架AdvPIE,一种无需访问目标模型参数的多模态代理框架。AdvPIE采用策略智能体基于裁判智能体的反馈生成并优化隐性对抗提示。裁判智能体在全局和相对层面提供跨迭代的模态特定安全评估以生成有效反馈。为高效利用反馈,我们提出一种新的累积对抗解码策略,动态重加权词元分布,优先选择生成更危害图像的词元,同时保持采样多样性。在标准及安全对齐的T2I模型上进行的大量实验表明,AdvPIE能有效揭露隐性漏洞,显著优于多种基线方法。
原文摘要 · Abstract (English)
Red-teaming Text-to-Image (T2I) models is essential for safe deployment, yet it remains particularly challenging against implicit adversarial prompts. Unlike explicit adversarial prompts that can be readily identified and blocked, implicit ones are much harder to detect: the prompts appear benign on the text surface yet still lead to inappropriate visual content. To address this, we propose Adversarial Probing for Implicit VulnErabilities (AdvPIE), a multimodal agentic framework to expose implicit vulnerabilities without requiring access to the parameters of target models. AdvPIE adopts a policy agent to generate and refine implicit adversarial prompts based on the feedback from a judge agent. To construct informative feedback, the judge agent provides modality-specific safety evaluation at both global and relative levels across iterations. To effectively leverage the feedback, we propose a novel Cumulative Adversarial Decoding strategy for the policy agent, which dynamically reweights token distributions to favor tokens that lead to more harmful images while preserving sampling diversity. Extensive experiments on standard and safety-aligned T2I models show that AdvPIE1 effectively uncovers implicit vulnerabilities, outperforming various baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。