同时逆向文本与噪声,让AI图像生成更准更可控
Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives

- 分两阶段逆向文本提示和潜在噪声,提升生成精度
- 在多个数据集上实现顶尖图像保真度,生成无明显瑕疵
- 适合需要精确图像编辑的科研与设计场景
提示逆向作为典型逆向工程方法,可使文本到图像扩散模型在无需大量提示调优的情况下生成目标图像。然而现有方法存在显著局限:(1) 基于梯度的方法不稳定且不可解释,常导致图像出现严重伪影;(2) 无梯度方法虽生成可读提示,但因缺乏细粒度细节对齐,难以保持视觉保真度。我们认为问题根源在于将提示逆向视为逆向工程的充分条件,忽视了编码结构信息的潜在噪声的关键作用。因此,我们提出Dualin(双逆向),一种两阶段方法,联合恢复目标图像的语义提示与潜在噪声。第一阶段结合视觉-语言模型CLIP与大语言模型,逆向生成忠实且人类可读的硬提示;第二阶段通过无条件DDIM逆向重建目标图像的精确潜在噪声,确保结构信息一致性。理论上,我们证明逆向噪声支持无需重优化的灵活图像编辑。大量实验表明,Dualin同时生成高质量逆向提示,并达到当前最优图像保真度。此外,Dualin为精准可控的图像编辑建立了坚实基础。
原文摘要 · Abstract (English)
Prompt inversion, as a typical reverse engineering technique, enables text-to-image (T2I) diffusion models to generate the desired target images without extensive prompt engineering. However, existing prompt inversion methods suffer from significant limitations: (1) gradient-based methods are unstable and uninterpretable, often resulting in generated images with severe artifacts; (2) gradient-free methods yield human-readable prompts but still fail to preserve visual fidelity due to the lack of fine-grained detail alignment. We contend that the limitations stem from treating prompt inversion as a sufficient condition for reverse engineering, ignoring the critical role of the latent noise that encodes structural information. Consequently, we propose Dualin (Dual inversion), a two-stage method that jointly recovers both the semantic prompt and latent noise of the target image. In the first stage, we integrate vision-language model, CLIP and large language model to invert a faithful, human-interpretable hard prompt. In the second stage, unconditional DDIM inversion reconstructs the exact latent noise of the target image, guaranteeing the consistency at the structural information level. Theoretically, we prove that the inverted noise enables flexible image editing without re-optimization. Extensive experiments on diverse datasets demonstrate that Dualin simultaneously generates high-quality inverted prompts and achieves state-of-the-art image fidelity. Additionally, Dualin can establish a robust foundation for the precise and controllable image editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。