arXiv:2509.12203cs.CV2025-09被引 3

让拖拽编辑更稳定,无需优化就能精准改图。

LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence

  • 用显式对应图替代隐式注意力匹配,提升控制精度。
  • 实现全强度反演,无需测试时优化,支持高质量修复与生成。
  • 适合需要精确几何控制和文本引导的复杂图像编辑场景。

拖拽编辑依赖注意力机制进行隐式点匹配,导致反演强度减弱且需昂贵的测试时优化(TTO),严重限制了扩散模型的生成能力,抑制高保真修复与文本引导生成。本文提出LazyDrag,首个适用于多模态扩散变压器的拖拽编辑方法,直接消除对隐式匹配的依赖。通过用户拖拽输入生成显式对应图作为可靠参考,增强注意力控制,实现稳定的全强度反演,首次在拖拽编辑任务中达成此目标。该方法无需TTO,释放模型生成潜力。因此,LazyDrag自然融合精确几何控制与文本引导,可实现此前无法完成的复杂编辑:如打开狗嘴并修复内部、生成网球等新物体,或对模糊拖拽做出上下文感知调整(如将手移入口袋)。此外,支持多轮操作,可同时进行移动与缩放。在DragBench上评估,其拖拽准确率与感知质量均优于基线,经VIEScore与人工评价验证。LazyDrag不仅达到新基准性能,更开辟了新的编辑范式。

原文摘要 · Abstract (English)

The reliance on implicit point matching via attention has become a core bottleneck in drag-based editing, resulting in a fundamental compromise on weakened inversion strength and costly test-time optimization (TTO). This compromise severely limits the generative capabilities of diffusion models, suppressing high-fidelity inpainting and text-guided creation. In this paper, we introduce LazyDrag, the first drag-based image editing method for Multi-Modal Diffusion Transformers, which directly eliminates the reliance on implicit point matching. In concrete terms, our method generates an explicit correspondence map from user drag inputs as a reliable reference to boost the attention control. This reliable reference opens the potential for a stable full-strength inversion process, which is the first in the drag-based editing task. It obviates the necessity for TTO and unlocks the generative capability of models. Therefore, LazyDrag naturally unifies precise geometric control with text guidance, enabling complex edits that were previously out of reach: opening the mouth of a dog and inpainting its interior, generating new objects like a ``tennis ball'', or for ambiguous drags, making context-aware changes like moving a hand into a pocket. Additionally, LazyDrag supports multi-round workflows with simultaneous move and scale operations. Evaluated on the DragBench, our method outperforms baselines in drag accuracy and perceptual quality, as validated by VIEScore and human evaluation. LazyDrag not only establishes new state-of-the-art performance, but also paves a new way to editing paradigms.

图像编辑扩散模型拖拽交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。