arXiv:2508.02973cs.CV2025-08ICCV

无需外部资源,自动优化负向提示以提升文生图准确性

Diffusion Models with Adaptive Negative Sampling Without External Resources

  • 基于模型内部对否定的理解,动态生成适配的负向采样
  • 在多个评测中超越基线,人类偏好度提升2倍
  • 无需额外负向提示,通用性强且不需训练

扩散模型(DMs)能根据文本提示生成多样且高质量的图像,但其在提示遵循性和图像质量上差异较大。负向提示被提出以提高对提示的遵循度,通过指定图像不应包含的内容。已有研究显示存在理想负向提示可最大化正向提示的成功率。本文探索了负向提示与无分类器引导(CFG)的关系,提出一种无需外部资源的自适应负向采样方法(ANSWER),仅通过单一提示同时考虑正负条件。该方法利用扩散模型内部对否定的理解,提升生成结果对提示的忠实度。ANSWER为免训练技术,适用于支持CFG的任意模型,无需显式负向提示(后者往往丢失信息且不完整)。实验表明,将ANSWER加入现有扩散模型后,在多个基准测试中表现优于基线,人类评估中偏好度提升2倍。

原文摘要 · Abstract (English)

Diffusion models (DMs) have demonstrated an unparalleled ability to create diverse and high-fidelity images from text prompts. However, they are also well-known to vary substantially regarding both prompt adherence and quality. Negative prompting was introduced to improve prompt compliance by specifying what an image must not contain. Previous works have shown the existence of an ideal negative prompt that can maximize the odds of the positive prompt. In this work, we explore relations between negative prompting and classifier-free guidance (CFG) to develop a sampling procedure, {\it Adaptive Negative Sampling Without External Resources} (ANSWER), that accounts for both positive and negative conditions from a single prompt. This leverages the internal understanding of negation by the diffusion model to increase the odds of generating images faithful to the prompt. ANSWER is a training-free technique, applicable to any model that supports CFG, and allows for negative grounding of image concepts without an explicit negative prompts, which are lossy and incomplete. Experiments show that adding ANSWER to existing DMs outperforms the baselines on multiple benchmarks and is preferred by humans 2x more over the other methods.

扩散模型文生图负向提示无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。