arXiv:2607.04731cs.CV2026-07被引 1

高引导值扩散反演为何失败?研究发现关键在提示与潜在表示的匹配度。

When Does High-CFG Diffusion Inversion Fail? A Controlled Study of Prompt--Latent Interactions

论文配图:When Does High-CFG Diffusion Inversion Fail? A Controlled Study of Prompt--Latent Interactions
图 1 · 摘自论文原文
  • 通过控制实验分析提示与潜在空间的交互,揭示反演失败的三类模式。
  • 总引导压力与重建质量相关,但无法解释中间类型提示-潜在对的成功与否。
  • 提出局部稳定干预方法,显著提升重建与提示间编辑效果。

文本引导的扩散反演是图像编辑的核心,将图像映射到初始潜在表示后,通过修改提示重放去噪过程实现编辑。实践中,反演常使用低于生成/编辑阶段的分类器自由引导(CFG)尺度,这种不匹配虽经验有效,却留下根本问题:当目标图像由高CFG轨迹生成时,何时能真正反演该轨迹?本研究在可控的生成-反演-重建设置中,已知真实初始潜在表示和去噪轨迹,采用现有扩散编辑基准中的提示,在高CFG下生成图像,并用固定点反演法以相同提示和引导设置重构。结果揭示三类提示级重建行为:易处理提示对多数初始潜在表示可重构,难处理提示对多数失败,中间类提示的成功取决于提示-潜在配对。为分析生成端,定义了逐步的提示压力,衡量CFG使去噪更新偏离无条件轨迹的程度。总压力与重建质量相关,能区分易与难提示,但无法解释中间对的成功或失败。文本侧分析显示,主要视觉主体和表述方式影响反演难度。最后评估了一种紧凑的轨迹一致性干预:仅在局部不稳定的反演步骤放松引导。该诊断检查提升了重建效果与提示间编辑性能,支持高CFG反演失败需局部、轨迹感知分析的观点。

原文摘要 · Abstract (English)

Text-guided diffusion inversion is central to image editing, where an image is mapped to an initial latent and then edited by replaying the denoising process under a modified prompt. In practice, however, inversion is often performed with a lower classifier-free guidance(CFG) scale than the one used for generation or editing. This mismatch is empirically useful but leaves a basic question unresolved: when a target image is generated by a high-CFG trajectory, when can that trajectory actually be inverted? We study this question in a controlled generation--inversion--reconstruction setting, where the true initial latent and denoising trajectory are known. Using prompts taken from an existing diffusion-editing benchmark, we generate images under high CFG and reconstruct them with fixed-point inversion using the same prompt and guidance setting. The results reveal three types of prompt-level reconstruction behavior: easy prompts that reconstruct for most initial latents, hard prompts that fail for most initial latents, and intermediate prompts whose success depends on the prompt--latent pairing. To analyze the generation side, we define prompt pressure, a step-wise measure of how strongly CFG moves the denoising update away from the unconditional trajectory. Total pressure correlates with reconstruction quality and separates easy from hard prompts, but it does not explain the success or failure of intermediate prompt--latent pairs. Text-side analyses further show that the main visual subject and wording can change inversion difficulty. Finally, we evaluate a compact trajectory-consistency intervention that relaxes guidance only at locally unstable inverse steps. This diagnostic check improves reconstruction and Prompt-to-Prompt editing in our controlled setting, supporting the view that high-CFG inversion failure requires local, trajectory-aware analysis.

扩散模型图像编辑反演失败提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。