arXiv:2505.21074cs.LGcs.AI2025-05NeurIPS被引 7

用规则模型引导大模型攻击文本生成图像系统,无需内部信息也能突破防御。

Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling

  • 通过规则化反馈偏好,让大模型动态适应未知防御机制。
  • 在19个图像生成系统和3个商用API上验证,攻击成功率显著提升。
  • 适合安全评估人员、模型开发者用于检测潜在风险漏洞。

文本到图像(T2I)模型可能生成不当或有害内容,引发伦理与安全问题。红队测试对评估其安全性至关重要,但白盒方法受限于需访问内部结构,难以应用于闭源模型;现有黑盒方法常假设了解模型具体防御机制,限制了在真实商业API场景中的应用。核心挑战在于如何绕过未知且多样的防御机制。为此,我们提出一种基于规则偏好建模的红队攻击方法(RPG-RT),通过大模型迭代修改提示词并利用T2I系统的反馈来微调自身。该方法将每轮反馈视为先验信息,使大模型能动态适应未知防御。由于反馈常为粗粒度标签,难以直接使用,我们进一步引入规则化偏好建模,借助一组规则评估期望或不期望的反馈,实现对大模型自适应过程的细粒度控制。在19个具有不同安全机制的T2I系统、3个在线商用API服务以及文本到视频(T2V)模型上进行的大量实验,验证了本方法的优越性与实用性。

原文摘要 · Abstract (English)

Text-to-image (T2I) models raise ethical and safety concerns due to their potential to generate inappropriate or harmful images. Evaluating these models' security through red-teaming is vital, yet white-box approaches are limited by their need for internal access, complicating their use with closed-source models. Moreover, existing black-box methods often assume knowledge about the model's specific defense mechanisms, limiting their utility in real-world commercial API scenarios. A significant challenge is how to evade unknown and diverse defense mechanisms. To overcome this difficulty, we propose a novel Rule-based Preference modeling Guided Red-Teaming (RPG-RT), which iteratively employs LLM to modify prompts to query and leverages feedback from T2I systems for fine-tuning the LLM. RPG-RT treats the feedback from each iteration as a prior, enabling the LLM to dynamically adapt to unknown defense mechanisms. Given that the feedback is often labeled and coarse-grained, making it difficult to utilize directly, we further propose rule-based preference modeling, which employs a set of rules to evaluate desired or undesired feedback, facilitating finer-grained control over the LLM's dynamic adaptation process. Extensive experiments on nineteen T2I systems with varied safety mechanisms, three online commercial API services, and T2V models verify the superiority and practicality of our approach.

红队测试生成安全大模型攻击规则建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。