arXiv:2509.18619cs.CV2025-09中稿 · DICTA 2025被引 23

用双路径控制让图像修复更准更清晰,不需训练也不用逐图优化。

Prompt-Guided Dual Latent Steering for Inversion Problems

  • 分结构与语义两条路径,用提示词引导修复方向
  • 在FFHQ-1K和ImageNet-1K上均优于单隐向量方法
  • 基于最优控制理论,动态调整生成轨迹防失真

将受损图像反演至扩散模型的隐空间极具挑战。现有方法将图像编码为单一隐向量,难以平衡结构保真度与语义准确性,导致重建出现语义漂移,如细节模糊或属性错误。为此,我们提出无需训练的提示引导双隐向量调控(PDLS)框架,基于修正流模型实现稳定反演路径。PDLS将反演过程分解为两个互补流:一条保留结构完整性,另一条由提示词引导语义信息。我们将双路引导建模为最优控制问题,通过线性二次调节器(LQR)推导出闭式解。该控制器在每一步动态调控生成轨迹,防止语义漂移,同时保证细粒度细节,无需昂贵的逐图像优化。在FFHQ-1K和ImageNet-1K上的多种反演任务(包括高斯去模糊、运动去模糊、超分辨率和自由形修补)实验表明,PDLS生成的重构结果既更贴近原图,又更符合语义信息,优于单隐向量基线方法。

原文摘要 · Abstract (English)

Inverting corrupted images into the latent space of diffusion models is challenging. Current methods, which encode an image into a single latent vector, struggle to balance structural fidelity with semantic accuracy, leading to reconstructions with semantic drift, such as blurred details or incorrect attributes. To overcome this, we introduce Prompt-Guided Dual Latent Steering (PDLS), a novel, training-free framework built upon Rectified Flow models for their stable inversion paths. PDLS decomposes the inversion process into two complementary streams: a structural path to preserve source integrity and a semantic path guided by a prompt. We formulate this dual guidance as an optimal control problem and derive a closed-form solution via a Linear Quadratic Regulator (LQR). This controller dynamically steers the generative trajectory at each step, preventing semantic drift while ensuring the preservation of fine detail without costly, per-image optimization. Extensive experiments on FFHQ-1K and ImageNet-1K under various inversion tasks, including Gaussian deblurring, motion deblurring, super-resolution and freeform inpainting, demonstrate that PDLS produces reconstructions that are both more faithful to the original image and better aligned with the semantic information than single-latent baselines.

图像修复扩散模型隐空间反演双路径控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。