arXiv:2508.10498cs.CV2025-08AAAI被引 6

通过路径正则化实现高效一致的图像编辑,无需反演且速度极快。

TweezeEdit: Consistent and Efficient Image Editing with Path Regularization

论文配图:TweezeEdit: Consistent and Efficient Image Editing with Path Regularization
图 1 · 摘自论文原文
  • 用梯度驱动的路径正则化替代反演锚点,直接控制生成过程。
  • 仅需12步(1.6秒/次)即可完成编辑,显著提升效率。
  • 适合需要快速、保持原图语义的实时图像编辑场景。

大规模预训练扩散模型使用户可通过文本引导编辑图像。然而,现有方法常过度匹配目标提示,而未能充分保留源图像语义。这些方法通常从源图像的反演噪声中显式或隐式生成目标图像,称为反演锚点。我们发现该策略在语义保留上表现不佳且效率低下,因编辑路径过长。为此,提出TweezeEdit——一种无需微调和反演的框架,通过正则化整个去噪路径而非依赖反演锚点,确保源图像语义保留并缩短编辑路径。在梯度驱动的正则化下,利用一致性模型沿直接路径高效注入目标提示语义。大量实验表明,TweezeEdit在语义保留与目标对齐方面优于现有方法,尤为突出的是仅需12步(每编辑1.6秒),具备实时应用潜力。

原文摘要 · Abstract (English)

Large-scale pre-trained diffusion models empower users to edit images through text guidance. However, existing methods often over-align with target prompts while inadequately preserving source image semantics. Such approaches generate target images explicitly or implicitly from the inversion noise of the source images, termed the inversion anchors. We identify this strategy as suboptimal for semantic preservation and inefficient due to elongated editing paths. We propose TweezeEdit, a tuning- and inversion-free framework for consistent and efficient image editing. Our method addresses these limitations by regularizing the entire denoising path rather than relying solely on the inversion anchors, ensuring source semantic retention and shortening editing paths. Guided by gradient-driven regularization, we efficiently inject target prompt semantics along a direct path using a consistency model. Extensive experiments demonstrate TweezeEdit's superior performance in semantic preservation and target alignment, outperforming existing methods. Remarkably, it requires only 12 steps (1.6 seconds per edit), underscoring its potential for real-time applications.

图像编辑扩散模型路径正则化实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。