让图像中物体在三维空间精准移动,解决遮挡与深度错乱问题。
SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling

- 通过隐式建模3D空间结构,理解物体深度关系。
- 实验显示复杂场景下物体移动精度显著提升。
- 适合需要精确3D空间编辑的视觉生成任务。
当前图像编辑技术虽能实现物体操控,但在复杂场景中处理空间移动仍存在困难,如物体跨越不同深度层或部分被遮挡。多数方法仅依赖2D数据集中的先验信息,强调平面特征而缺乏对空间结构的支持。即使引入显式位置信息的方法,也无法准确捕捉真实3D空间关系,导致复杂场景中物体移动不准确。本文提出SpatialDiff,通过两项核心创新:(1)隐式3D空间建模,引入3D先验知识,使模型内部构建对三维空间结构的全面理解;(2)全局空间监督,约束潜在空间特征,使模型能够感知编辑操作引起的物体空间位置变化。实验表明,该方法显著提升了复杂场景中物体移动的准确性与保真度。
原文摘要 · Abstract (English)
Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span different depth layers or are partially occluded. Most image editing methods focus solely on prior information from 2D datasets, emphasizing planar features while lacking support for spatial structures. Even approaches that incorporate explicit positional information fail to capture true 3D spatial relationships, thus limiting accurate object movement in complex scenes. In this paper, we present SpatialDiff, a method that effectively captures 3D spatial structures, enabling precise and consistent object movements in complex scenes. Our core innovations are twofold: (1) Implicit 3D Spatial Modeling, which introduces 3D prior knowledge and enables the model to internally build a comprehensive understanding of the three-dimensional spatial structure; and (2) Global Spatial Supervision, which constrains the latent spatial features to enable the model to perceive changes in object spatial positions caused by editing operations. Experimental results demonstrate that our method significantly improves the accuracy and fidelity of spatial movement in complex scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。