arXiv:2510.10715cs.GRcs.CV2025-10被引 3

用视觉语言模型引导生成,让AI画出更新颖的创意图像。

VLM-Guided Adaptive Negative Prompting for Creative Generation

  • 通过VLM分析中间生成结果,动态调整负向提示词
  • 在CLIP空间中提升图像新颖性,且不损失内容真实性
  • 无需训练,可直接用于复杂场景和组合提示

创意生成旨在创造新颖、意外且有价值的新样本,反映用户意图但无法预先设想。该任务旨在拓展人类想象力,发现存在于熟悉领域之间的未探索视觉概念。尽管文本到图像扩散模型能精准渲染符合提示的逼真图像,但仍难以生成真正新颖的内容。现有方法或依赖图像特征插值(局限于预定义类别),或需耗时的嵌入优化或模型微调。我们提出VLM-Guided Adaptive Negative Prompting,一种无需训练、仅在推理阶段执行的方法,可在保持生成物体有效性的同时促进创意生成。该方法利用视觉语言模型(VLM)分析生成过程中的中间输出,自适应地引导其远离常规视觉概念,从而激发新颖且令人惊喜的输出。我们通过统计指标在CLIP嵌入空间中评估创意性,涵盖新颖性和有效性。大量实验表明,该方法在保持极低计算开销的前提下,持续提升创意新颖性。此外,不同于仅生成单个对象的现有方法,本方法可扩展至复杂场景,如生成一致的创意物体集合,并在复杂构图提示中保持创造力。该方法可无缝集成至现有扩散模型流程,为突破文本描述限制、生成更具创意的输出提供实用路径。

原文摘要 · Abstract (English)

Creative generation is the synthesis of new, surprising, and valuable samples that reflect user intent yet cannot be envisioned in advance. This task aims to extend human imagination, enabling the discovery of visual concepts that exist in the unexplored spaces between familiar domains. While text-to-image diffusion models excel at rendering photorealistic scenes that faithfully match user prompts, they still struggle to generate genuinely novel content. Existing approaches to enhance generative creativity either rely on interpolation of image features, which restricts exploration to predefined categories, or require time-intensive procedures such as embedding optimization or model fine-tuning. We propose VLM-Guided Adaptive Negative-Prompting, a training-free, inference-time method that promotes creative image generation while preserving the validity of the generated object. Our approach utilizes a vision-language model (VLM) that analyzes intermediate outputs of the generation process and adaptively steers it away from conventional visual concepts, encouraging the emergence of novel and surprising outputs. We evaluate creativity through both novelty and validity, using statistical metrics in the CLIP embedding space. Through extensive experiments, we show consistent gains in creative novelty with negligible computational overhead. Moreover, unlike existing methods that primarily generate single objects, our approach extends to complex scenarios, such as generating coherent sets of creative objects and preserving creativity within elaborate compositional prompts. Our method integrates seamlessly into existing diffusion pipelines, offering a practical route to producing creative outputs that venture beyond the constraints of textual descriptions.

创意生成视觉语言模型扩散模型负向提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。