用旋转位置编码实现深度感知的物体移动,保持场景一致性。
RoPEMover: Depth-Aware Object Relocation via Positional Embeddings

- 基于旋转位置编码构建深度感知空间场,控制物体位移
- 在真实图像上微调后仍能保持物体身份与光影一致性
- 适合需要精确场景重建的图像编辑任务
单图中移动物体需保持几何一致的时空重排,包括处理遮挡、展现新区域、维持阴影与反射的一致性。现有方法难以满足此需求,常导致场景不连贯。本文提出一种基于扩散变换器位置表示的几何感知物体运动方法。核心思想是旋转位置编码(RoPE)定义了可显式操控的结构化空间场,用于诱导可控运动。将2D RoPE扩展为包含3D空间结构的深度感知形式,实现一致的物体位移与场景感知更新。模型通过合成数据结合少量真实图像进行参数高效微调。即使真实监督极少,仍能在大范围位移下保持物体身份,生成新暴露区域的合理内容,并一致更新阴影与光照等场景依赖效应。在标准物体移动基准测试中,各项指标均达到当前最优性能。
原文摘要 · Abstract (English)
Moving an object in a single image requires geometry-consistent spatial rearrangement, including handling occlusions, revealing previously unseen regions, and maintaining coherent shadows and reflections. Existing approaches are not well suited to this setting and often fail to preserve such scene-level consistency. We address this problem by introducing a geometry-aware object motion method that operates directly on the positional representations of diffusion transformers. Our key insight is that rotary positional embeddings (RoPE) define a structured spatial field that can be explicitly manipulated to induce controlled motion. We extend 2D RoPE into a depth-aware formulation that encodes 3D spatial structure, enabling consistent object displacement and scene-aware updates. Our model is trained using synthetic data combined with a small set of real images via parameter-efficient fine-tuning. Despite minimal real supervision, it preserves object identity under large spatial displacements, generates plausible content in newly revealed regions, and consistently updates scene-dependent effects such as shadows and illumination. Experimental results on standard object motion benchmarks demonstrate state-of-the-art performance across all evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。