轻量级模块可精准拦截不当图像生成,且不影响正常使用。
SteerDiff: Steering towards Safe Text-to-Image Diffusion Models
- 在文本嵌入空间中识别并操控不当概念,实现安全干预。
- 在多种概念删除任务中表现优异,对红队攻击具有强鲁棒性。
- 无需微调主模型,适合快速部署于各类文生图系统。
文生图扩散模型虽能生成高质量、精确对齐文本的图像,但易被滥用于生成不当内容。现有安全措施多依赖大规模标注数据的文本分类器或类似ControlNet的方法,存在易被改写绕过、难以随模型规模扩展、灵活性差等问题。近期红队攻击研究更凸显了新防护范式的必要性。本文提出SteerDiff,一种轻量级适配器模块,作为用户输入与扩散模型间的中介,确保生成图像符合伦理安全标准,同时对用户体验影响极小。SteerDiff通过识别并操纵文本嵌入空间中的不当概念,引导模型避开有害输出。我们通过大量概念删除任务评估其有效性,并在多个红队攻击策略下验证其鲁棒性。此外,还探索了SteerDiff在概念遗忘任务中的潜力,证明其在文生图生成中的多功能性。
原文摘要 · Abstract (English)
Text-to-image (T2I) diffusion models have drawn attention for their ability to generate high-quality images with precise text alignment. However, these models can also be misused to produce inappropriate content. Existing safety measures, which typically rely on text classifiers or ControlNet-like approaches, are often insufficient. Traditional text classifiers rely on large-scale labeled datasets and can be easily bypassed by rephrasing. As diffusion models continue to scale, fine-tuning these safeguards becomes increasingly challenging and lacks flexibility. Recent red-teaming attack researches further underscore the need for a new paradigm to prevent the generation of inappropriate content. In this paper, we introduce SteerDiff, a lightweight adaptor module designed to act as an intermediary between user input and the diffusion model, ensuring that generated images adhere to ethical and safety standards with little to no impact on usability. SteerDiff identifies and manipulates inappropriate concepts within the text embedding space to guide the model away from harmful outputs. We conduct extensive experiments across various concept unlearning tasks to evaluate the effectiveness of our approach. Furthermore, we benchmark SteerDiff against multiple red-teaming strategies to assess its robustness. Finally, we explore the potential of SteerDiff for concept forgetting tasks, demonstrating its versatility in text-conditioned image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。