用进化算法优化提示词,让图像还原更准确自然。
PromptEvolver: Prompt Inversion through Evolutionary Optimization in Natural-Language Space

- 用遗传算法在自然语言空间进化提示词
- 黑盒模型下实现高保真图像重建
- 生成的提示词更自然易懂,适合可控生成
文本到图像生成技术发展迅速,但复杂场景的精准生成仍需大量试错。提示词反演的目标是恢复能忠实重建给定目标图像的文本提示。现有方法常导致重建效果不佳,且生成的提示词不自然、难以理解,影响透明性与可控性。本文提出 PromptEvolver,一种在自然语言空间中通过遗传算法优化提示词的方法,利用强大的视觉-语言模型引导进化过程,仅需生成模型的图像输出即可工作,适用于黑箱模型。我们在多个提示词反演基准上评估,结果表明其持续优于现有方法。
原文摘要 · Abstract (English)
Text-to-image generation has progressed rapidly, but faithfully generating complex scenes requires extensive trial-and-error to find the exact prompt. In the prompt inversion task, the goal is to recover a textual prompt that can faithfully reconstruct a given target image. Currently, existing methods frequently yield suboptimal reconstructions and produce unnatural, hard-to-interpret prompts that hinder transparency and controllability. In this work, we present PromptEvolver, a prompt inversion approach that generates natural-language prompts while achieving high-fidelity reconstructions of the target image. Our method uses a genetic algorithm to optimize the prompt, leveraging a strong vision-language model to guide the evolution process. Importantly, it works on black-box generation models by requiring only image outputs. Finally, we evaluate PromptEvolver across multiple prompt inversion benchmarks and show that it consistently outperforms competing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。