arXiv:2509.25740cs.CV2025-09被引 5

用3D几何信息指导图像拖拽编辑,让物体旋转透视更自然

Dragging with Geometry: From Pixels to Geometry-Guided Image Editing

  • 用统一位移场融合3D几何与2D空间先验
  • 单次前向传播实现高保真、结构一致的编辑
  • 多点拖拽不冲突,适合复杂场景操作

交互式点驱动图像编辑可实现内容的精确灵活控制。然而,现有拖拽方法主要在2D像素平面上操作,缺乏对3D线索的利用,导致在旋转、透视等几何密集场景中常产生不准确、不一致的编辑结果。为此,我们提出一种新型几何引导的拖拽编辑方法GeoDrag,解决三大挑战:1)将3D几何线索融入像素级编辑;2)缓解仅依赖几何引导带来的不连续性;3)解决多点拖拽引发的冲突。基于联合编码3D几何与2D空间先验的统一位移场,GeoDrag可在单次前向传播中实现连贯、高保真、结构一致的编辑。此外,引入无冲突分区策略隔离编辑区域,有效避免干扰并保证一致性。大量实验验证了方法的有效性,在多种编辑场景中均展现出更优的精度、结构一致性和可靠的多点编辑能力。

原文摘要 · Abstract (English)

Interactive point-based image editing serves as a controllable editor, enabling precise and flexible manipulation of image content. However, most drag-based methods operate primarily on the 2D pixel plane with limited use of 3D cues. As a result, they often produce imprecise and inconsistent edits, particularly in geometry-intensive scenarios such as rotations and perspective transformations. To address these limitations, we propose a novel geometry-guided drag-based image editing method-GeoDrag, which addresses three key challenges: 1) incorporating 3D geometric cues into pixel-level editing, 2) mitigating discontinuities caused by geometry-only guidance, and 3) resolving conflicts arising from multi-point dragging. Built upon a unified displacement field that jointly encodes 3D geometry and 2D spatial priors, GeoDrag enables coherent, high-fidelity, and structure-consistent editing in a single forward pass. In addition, a conflict-free partitioning strategy is introduced to isolate editing regions, effectively preventing interference and ensuring consistency. Extensive experiments across various editing scenarios validate the effectiveness of our method, showing superior precision, structural consistency, and reliable multi-point editability. Project page: https://xinyu-pu.github.io/projects/geodrag.

图像编辑几何引导拖拽操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。