arXiv:2608.23238cs.CV2026-08

让360°全景图物体操作像普通图片一样简单,支持一键移动、插入和删除。

Mover360: Controllable Object Manipulation in 360° Panoramic Images

论文配图:Mover360: Controllable Object Manipulation in 360° Panoramic Images
图 1 · 摘自论文原文
  • 用统一指令地图实现点、框、掩码三种控制方式,点击即移物
  • 在合成与真实全景图上均超越现有方法,在细节还原和语义一致上表现更优
  • 专为全景图设计,适合虚拟现实、全景摄影等场景的编辑需求

我们提出Mover360,一个针对360°全景图像的可控物体操作框架。与透视图像不同,360°全景图在等距柱状投影(ERP)下存在水平环绕、纬度相关畸变和全局场景连续性,使现有透视编辑工具难以有效进行物体级编辑。Mover360以物体平移为核心任务,同时支持参考引导的插入和删除作为辅助任务。其界面通过将每项任务编码为固定提示和紧凑的ERP对齐指令图,统一点、边界框和掩码引导控制。默认点模式下,单击即可完成物体重定位,模型利用全景上下文和辅助深度条件推断合理大小、支撑与光照。结构上,Mover360是预训练扩散变换器的轻量级适配。为生成成对监督数据,我们构建了基于UE5的数据生成管道,实现表面感知物体放置与随机光照,生成大规模成对数据,并建立涵盖合成与真实全景图的双域基准。在两个测试域和两种评估协议下,Mover360在重建保真度、语义一致性与分布质量上均优于强基线方法。代码与基准数据集已公开。

原文摘要 · Abstract (English)

We present Mover360, a controllable object manipulation framework for 360° images. Unlike perspective images, 360° images in equirectangular projection (ERP) exhibit horizontal wrap-around, latitude-dependent distortion, and global scene continuity, which makes object-level edits difficult for existing perspective editors to produce and for users to specify. To address this, Mover360 centers on object Translation (relocating a specified object within an existing panorama) while supporting reference-guided Insert and Remove as auxiliary tasks. Its interface unifies point-, bbox-, and mask-guided control by encoding each task into a fixed prompt and a compact, ERP-aligned instruction map. In the default point mode, a single click relocates an object, allowing the model to infer a plausible size, support, and illumination using panoramic context and an auxiliary depth condition. Structurally, Mover360 is a lightweight adaptation of a pretrained diffusion transformer. To generate paired supervision, we construct a UE5 data-generation pipeline with surface-aware object placement and randomized illumination, yielding large-scale paired data and a dual-domain benchmark of synthetic and real panoramas with ground truth for all three tasks. Across both test domains and two evaluation protocols, Mover360 outperforms strong baselines for perspective editing, insertion, and inpainting in reconstruction fidelity, semantic consistency, and distributional quality. Code and our benchmark dataset are available at https://zhonghaoyi.github.io/Mover360/.

全景图像物体编辑扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。