用大模型生成可读提示,让文生图更可控。
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
- 用大模型+CLIP分数生成可读提示,无需反向传播
- 生成提示与视觉概念对齐度高,语义连贯
- 适合需要精准控制文生图的设计师和开发者
文生图模型如DALL-E和Stable Diffusion已广泛应用于广告、个性化媒体和设计原型等领域。然而,生成有效提示仍需大量试错。现有提示反演方法(如软/硬提示)因可解释性差、提示不连贯而效果有限。为此,我们提出视觉引导解码(VGD),一种无梯度的方法,利用大语言模型(LLMs)和基于CLIP的指导生成语义一致且可读的提示。VGD借助LLMs强大的文本生成能力,结合CLIP分数确保提示与用户指定视觉概念对齐,无需额外训练即可提升提示的可解释性、泛化性和灵活性。实验表明,VGD在生成可理解、上下文相关提示方面优于现有技术,使与文生图模型的交互更直观可控。
原文摘要 · Abstract (English)
Text-to-image generative models like DALL-E and Stable Diffusion have revolutionized visual content creation across various applications, including advertising, personalized media, and design prototyping. However, crafting effective textual prompts to guide these models remains challenging, often requiring extensive trial and error. Existing prompt inversion approaches, such as soft and hard prompt techniques, are not so effective due to the limited interpretability and incoherent prompt generation. To address these issues, we propose Visually Guided Decoding (VGD), a gradient-free approach that leverages large language models (LLMs) and CLIP-based guidance to generate coherent and semantically aligned prompts. In essence, VGD utilizes the robust text generation capabilities of LLMs to produce human-readable prompts. Further, by employing CLIP scores to ensure alignment with user-specified visual concepts, VGD enhances the interpretability, generalization, and flexibility of prompt generation without the need for additional training. Our experiments demonstrate that VGD outperforms existing prompt inversion techniques in generating understandable and contextually relevant prompts, facilitating more intuitive and controllable interactions with text-to-image models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。