用提示词优化让图像生成模型快速准确地解逆问题。
LATINO-PRO: LAtent consisTency INverse sOlver with PRompt Optimization
- 基于隐空间一致性模型,设计零样本反问题求解框架。
- 仅需8次神经函数计算即达当前最优重建质量。
- 自动校准提示词,适合图像重建与高效计算场景。
文本到图像的隐空间扩散模型(LDMs)在成像逆问题求解中展现出巨大潜力。然而,以即插即用、零样本方式使用这类模型仍具挑战,因需为未知图像找到合适文本提示。现有文本到图像的即插即用方法计算成本极高。本文提出一种专为嵌入生成模型而设计的新颖即插即用推断范式,重点关注将LDMs压缩为快速生成器的隐空间一致性模型(LCMs)。我们基于该框架提出首个利用LCMs先验的零样本即插即用框架LATINO。其条件机制避免了自动微分,在仅8次神经函数评估下即达到最先进性能。结果表明,LATINO能实现高精度重建,并显著优于以往方法在内存和计算效率上。随后,我们将LATINO嵌入经验贝叶斯框架,通过边际最大似然估计从观测数据自动校准文本提示。大量实验显示,提示词自校准大幅提升估计效果,使带提示词优化的LATINO在图像重建质量和计算效率上均定义新基准。代码已公开于 https://latino-pro.github.io
原文摘要 · Abstract (English)
Text-to-image latent diffusion models (LDMs) have recently emerged as powerful generative models with great potential for solving inverse problems in imaging. However, leveraging such models in a Plug & Play (PnP), zero-shot manner remains challenging because it requires identifying a suitable text prompt for the unknown image of interest. Also, existing text-to-image PnP approaches are highly computationally expensive. We herein address these challenges by proposing a novel PnP inference paradigm specifically designed for embedding generative models within stochastic inverse solvers, with special attention to Latent Consistency Models (LCMs), which distill LDMs into fast generators. We leverage our framework to propose LAtent consisTency INverse sOlver (LATINO), the first zero-shot PnP framework to solve inverse problems with priors encoded by LCMs. Our conditioning mechanism avoids automatic differentiation and reaches SOTA quality in as little as 8 neural function evaluations. As a result, LATINO delivers remarkably accurate solutions and is significantly more memory and computationally efficient than previous approaches. We then embed LATINO within an empirical Bayesian framework that automatically calibrates the text prompt from the observed measurements by marginal maximum likelihood estimation. Extensive experiments show that prompt self-calibration greatly improves estimation, allowing LATINO with PRompt Optimization to define new SOTAs in image reconstruction quality and computational efficiency. The code is available at https://latino-pro.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。