arXiv:2602.11440cs.CV2026-02中稿 · ICLR被引 3

无需3D建模即可实现精准几何可控的图像物体操作

Ctrl&Shift: High-Quality Geometry-Aware Object Manipulation in Visual Generation

  • 分两阶段处理:先移除物体,再在相机位姿控制下参考修复
  • 在真实场景视频上实现高保真度与视角一致性,效果领先
  • 适合影视后期、AR创作等需要精细空间控制的场景

物体级操作——在保持场景真实性的前提下移动或旋转图像/视频中的物体——是影视后期、增强现实和创意编辑的核心需求。现有方法难以同时满足背景保留、视角变化下的几何一致性以及用户可控性三大目标。基于几何的方法虽控制精确但需显式3D重建且泛化差;扩散模型泛化好但缺乏细粒度几何控制。本文提出Ctrl&Shift,一种无需显式3D表示的端到端扩散框架,通过将操作分解为物体移除与受控参考修复两个阶段,并在统一扩散过程中编码相机位姿信息。设计多任务多阶段训练策略,分离背景、身份与姿态信号。构建可扩展的真实世界数据集生成流程,提供带有估计相对相机位姿的图像视频配对数据。大量实验证明,该方法在保真度、视角一致性与可控性方面达到当前最优。据我们所知,这是首个在不依赖显式3D建模的前提下,统一实现细粒度几何控制与真实世界泛化的物体操作框架。

原文摘要 · Abstract (English)

Object-level manipulation, relocating or reorienting objects in images or videos while preserving scene realism, is central to film post-production, AR, and creative editing. Yet existing methods struggle to jointly achieve three core goals: background preservation, geometric consistency under viewpoint shifts, and user-controllable transformations. Geometry-based approaches offer precise control but require explicit 3D reconstruction and generalize poorly; diffusion-based methods generalize better but lack fine-grained geometric control. We present Ctrl&Shift, an end-to-end diffusion framework to achieve geometry-consistent object manipulation without explicit 3D representations. Our key insight is to decompose manipulation into two stages, object removal and reference-guided inpainting under explicit camera pose control, and encode both within a unified diffusion process. To enable precise, disentangled control, we design a multi-task, multi-stage training strategy that separates background, identity, and pose signals across tasks. To improve generalization, we introduce a scalable real-world dataset construction pipeline that generates paired image and video samples with estimated relative camera poses. Extensive experiments demonstrate that Ctrl&Shift achieves state-of-the-art results in fidelity, viewpoint consistency, and controllability. To our knowledge, this is the first framework to unify fine-grained geometric control and real-world generalization for object manipulation, without relying on any explicit 3D modeling.

图像编辑扩散模型几何控制视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。