arXiv:2507.08285cs.GRcs.CV2025-07ICML被引 14

通过3D网格引导变形流,实现更精准的拖拽图像编辑

FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow Fields

  • 用3D网格和能量函数指导拖拽点的几何变形
  • 在VFD和DragBench上优于现有方法,结构保持更稳定
  • 构建了带真实标签的VFD视频拖拽数据集,便于评估

拖拽编辑通过点控实现精确物体操作,但现有方法仅关注对齐用户指定点,忽视整体几何结构,导致伪影或编辑不稳。本文提出FlowDrag,从图像构建3D网格,利用能量函数引导网格基于拖拽点变形,将所得位移投影至2D并融入UNet去噪过程,实现精确的控制点到目标点对齐,同时保持结构完整性。此外,现有拖拽编辑评测缺乏真实标签,难以准确评估编辑效果。为此,我们构建了VFD(VidFrameDrag)基准数据集,利用视频数据集中连续帧提供真实参考帧。FlowDrag在VFD Bench和DragBench上均超越现有方法。

原文摘要 · Abstract (English)

Drag-based editing allows precise object manipulation through point-based control, offering user convenience. However, current methods often suffer from a geometric inconsistency problem by focusing exclusively on matching user-defined points, neglecting the broader geometry and leading to artifacts or unstable edits. We propose FlowDrag, which leverages geometric information for more accurate and coherent transformations. Our approach constructs a 3D mesh from the image, using an energy function to guide mesh deformation based on user-defined drag points. The resulting mesh displacements are projected into 2D and incorporated into a UNet denoising process, enabling precise handle-to-target point alignment while preserving structural integrity. Additionally, existing drag-editing benchmarks provide no ground truth, making it difficult to assess how accurately the edits match the intended transformations. To address this, we present VFD (VidFrameDrag) benchmark dataset, which provides ground-truth frames using consecutive shots in a video dataset. FlowDrag outperforms existing drag-based editing methods on both VFD Bench and DragBench.

图像编辑3D感知拖拽生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。