arXiv:2510.08181cs.CV2025-10

让文字指令与拖拽结合,精准移动图像物体并修改属性。

InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing

  • 将拖拽视为图像重建,分两支协同实现精准定位。
  • 融合文本指令与拖拽梯度,实现位置与语义双重控制。
  • 支持细粒度编辑,适合需要高精度交互的图像设计者。

文本到图像扩散模型在图像编辑中展现出巨大潜力,文本驱动和物体拖拽成为主流方法。但前者难以精确控制物体位置,后者仅限于静态重定位。为此,我们提出 InstructUDrag,一种基于扩散框架的联合方法,同时支持文本指令与物体拖拽。该框架将拖拽视为图像重建过程,分为两个协同分支:移动-重建分支利用能量梯度引导精确移动物体,并通过优化交叉注意力图提升定位精度;文本驱动编辑分支与重建分支共享梯度信号,确保变换一致,并实现对物体属性的细粒度控制。此外,我们采用 DDPM 反演并将先验信息注入噪声图,以保留被移动物体的结构。大量实验表明,InstructUDrag 能实现灵活、高保真的图像编辑,兼具物体精确定位与内容语义控制能力。

原文摘要 · Abstract (English)

Text-to-image diffusion models have shown great potential for image editing, with techniques such as text-based and object-dragging methods emerging as key approaches. However, each of these methods has inherent limitations: text-based methods struggle with precise object positioning, while object dragging methods are confined to static relocation. To address these issues, we propose InstructUDrag, a diffusion-based framework that combines text instructions with object dragging, enabling simultaneous object dragging and text-based image editing. Our framework treats object dragging as an image reconstruction process, divided into two synergistic branches. The moving-reconstruction branch utilizes energy-based gradient guidance to move objects accurately, refining cross-attention maps to enhance relocation precision. The text-driven editing branch shares gradient signals with the reconstruction branch, ensuring consistent transformations and allowing fine-grained control over object attributes. We also employ DDPM inversion and inject prior information into noise maps to preserve the structure of moved objects. Extensive experiments demonstrate that InstructUDrag facilitates flexible, high-fidelity image editing, offering both precision in object relocation and semantic control over image content.

图像编辑扩散模型交互式文本指令

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。