arXiv:2511.04834cs.LGcs.AI2025-11中稿 · NeurIPS被引 1

用隐式负向嵌入替代负向提示,提升文本到图像模型的安全性。

Prompt-Based Safety Guidance Is Ineffective for Unlearned Text-to-Image Diffusion Models

  • 用概念反演生成隐式负向嵌入,替代传统负向提示。
  • 在裸露与暴力基准上,防御成功率显著提升。
  • 无需修改现有模型,可无缝集成到安全防护流程中。

近期文本到图像生成模型的发展引发了对其在恶意输入下生成有害内容的担忧。为此,主要出现了两种方法:(1) 微调模型以消除有害概念;(2) 无需训练的引导方法,利用负向提示。然而我们发现,将这两种正交方法结合常导致防御性能边际甚至下降,表明二者存在关键不兼容性。本文提出一种概念简单但实验稳健的方法:用通过概念反演获得的隐式负向嵌入,替代训练无感方法中的负向提示。该方法无需修改任一原有方案,可轻松融入现有流程。我们在裸露与暴力基准上验证其有效性,结果显示防御成功率持续提升,同时保留输入提示的核心语义。

原文摘要 · Abstract (English)

Recent advances in text-to-image generative models have raised concerns about their potential to produce harmful content when provided with malicious input text prompts. To address this issue, two main approaches have emerged: (1) fine-tuning the model to unlearn harmful concepts and (2) training-free guidance methods that leverage negative prompts. However, we observe that combining these two orthogonal approaches often leads to marginal or even degraded defense performance. This observation indicates a critical incompatibility between two paradigms, which hinders their combined effectiveness. In this work, we address this issue by proposing a conceptually simple yet experimentally robust method: replacing the negative prompts used in training-free methods with implicit negative embeddings obtained through concept inversion. Our method requires no modification to either approach and can be easily integrated into existing pipelines. We experimentally validate its effectiveness on nudity and violence benchmarks, demonstrating consistent improvements in defense success rate while preserving the core semantics of input prompts.

文本生成扩散模型安全性负向提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。