arXiv:2506.03067cs.CV2025-06

提出可解释的文本到图像生成提示反演方法,提升生成图像与原始提示的匹配度。

EDITOR: Effective and Interpretable Prompt Inversion for Text-to-Image Diffusion Models

  • 用预训练图像描述模型初始化嵌入,再在隐空间逆向优化
  • 在多个数据集上实现更高图像相似度和提示可读性
  • 适合需要可解释性提示还原的应用场景

文本到图像生成模型(如Stable Diffusion)已取得显著进展,能根据文本描述生成高质量图像。提示反演旨在识别生成特定图像所用的原始文本提示,对数据归属、模型溯源和水印验证具有重要意义。现有方法虽采用延迟投影优化词汇空间代表性,但仍存在语义流畅性和效率问题。高级图像描述模型或视觉大语言模型虽能生成可解释提示,但图像相似度不足。本文提出一种名为\sys的提示反演技术,包括使用预训练图像描述模型初始化嵌入,通过隐空间逆向工程优化,并利用嵌入到文本模型转换为文本。在MS COCO、LAION、Flickr和DiffusionDB等广泛数据集上的实验表明,该方法在图像相似度、文本对齐、提示可解释性和泛化能力方面均优于现有方法。进一步展示了生成提示在跨概念图像合成、概念操控、多概念演化生成和无监督分割中的应用。

原文摘要 · Abstract (English)

Text-to-image generation models~(e.g., Stable Diffusion) have achieved significant advancements, enabling the creation of high-quality and realistic images based on textual descriptions. Prompt inversion, the task of identifying the textual prompt used to generate a specific artifact, holds significant potential for applications including data attribution, model provenance, and watermarking validation. Recent studies introduced a delayed projection scheme to optimize for prompts representative of the vocabulary space, though challenges in semantic fluency and efficiency remain. Advanced image captioning models or visual large language models can generate highly interpretable prompts, but they often lack in image similarity. In this paper, we propose a prompt inversion technique called \sys for text-to-image diffusion models, which includes initializing embeddings using a pre-trained image captioning model, refining them through reverse-engineering in the latent space, and converting them to texts using an embedding-to-text model. Our experiments on the widely-used datasets, such as MS COCO, LAION, Flickr and DiffusionDB, show that our method outperforms existing methods in terms of image similarity, textual alignment, prompt interpretability and generalizability. We further illustrate the application of our generated prompts in tasks such as cross-concept image synthesis, concept manipulation, evolutionary multi-concept generation and unsupervised segmentation.

提示反演扩散模型可解释性图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。