arXiv:2511.00686cs.CVcs.AI2025-11中稿 · NeurIPS

用新颖性搜索生成多样化图像,让单一提示产生创意多样性。

Evolve to Inspire: Novelty Search for Diverse Image Generation

  • 基于大语言模型演化提示,用CLIP衡量图像新颖性
  • 引入发射器机制,使生成结果在提示空间中分布更广
  • 在多个指标上超越现有方法,适合创意设计场景

文生图扩散模型虽能生成高保真图像,但输出多样性有限,难以支持探索性与创意任务。现有提示优化方法多聚焦美学质量,不适用于创造性视觉领域。为此,我们提出WANDER——一种基于新颖性搜索的图像生成方法,可从单个输入提示生成多样化的图像集合。WANDER直接作用于自然语言提示,利用大语言模型(LLM)进行语义演化,并通过CLIP嵌入量化图像新颖性。同时引入发射器机制,引导搜索进入提示空间的不同区域,显著提升生成图像多样性。在FLUX-DEV生成、GPT-4o-mini进行变异的实验中,WANDER在多样性指标上显著优于现有进化式提示优化基线。消融实验证明了发射器的有效性。

原文摘要 · Abstract (English)

Text-to-image diffusion models, while proficient at generating high-fidelity images, often suffer from limited output diversity, hindering their application in exploratory and ideation tasks. Existing prompt optimization techniques typically target aesthetic fitness or are ill-suited to the creative visual domain. To address this shortcoming, we introduce WANDER, a novelty search-based approach to generating diverse sets of images from a single input prompt. WANDER operates directly on natural language prompts, employing a Large Language Model (LLM) for semantic evolution of diverse sets of images, and using CLIP embeddings to quantify novelty. We additionally apply emitters to guide the search into distinct regions of the prompt space, and demonstrate that they boost the diversity of the generated images. Empirical evaluations using FLUX-DEV for generation and GPT-4o-mini for mutation demonstrate that WANDER significantly outperforms existing evolutionary prompt optimization baselines in diversity metrics. Ablation studies confirm the efficacy of emitters.

图像生成扩散模型多样性提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。