arXiv:2510.26052cs.CVcs.AI2025-10被引 1

用视觉语言模型动态生成负面提示,提升图像生成质量。

Dynamic VLM-Guided Negative Prompting for Diffusion Models

  • 在去噪过程中分步生成图像并调用VLM生成适配的负面提示。
  • 实验表明负面引导强度与图文一致性存在权衡关系。
  • 适合需要精准控制生成内容的文本到图像任务用户。

我们提出一种新的扩散模型动态负面提示方法,利用视觉语言模型(VLM)在去噪过程中自适应生成负面提示。与传统使用固定负面提示的方法不同,本方法在特定去噪步骤生成中间图像预测,并通过VLM生成上下文相关的负面提示。我们在多个基准数据集上评估该方法,揭示了负面引导强度与文本-图像对齐效果之间的权衡关系。

原文摘要 · Abstract (English)

We propose a novel approach for dynamic negative prompting in diffusion models that leverages Vision-Language Models (VLMs) to adaptively generate negative prompts during the denoising process. Unlike traditional Negative Prompting methods that use fixed negative prompts, our method generates intermediate image predictions at specific denoising steps and queries a VLM to produce contextually appropriate negative prompts. We evaluate our approach on various benchmark datasets and demonstrate the trade-offs between negative guidance strength and text-image alignment.

扩散模型负向提示视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。