用3D框精确控制真实图像的物体变换,效果更自然。
Thinking in Boxes: 3D Editing in Real Images Made Easy

- 用户指定物体输入/输出3D框,将编辑转为几何问题。
- 在大幅运动和视角变化下仍保持物体与场景一致。
- 适合需要精准3D编辑的视觉设计与影视制作人员。
文本和2D条件接口对图像编辑中的空间变换控制能力弱且模糊,尤其在大范围物体运动和相机变化下。现有方法虽使用3D体素如盒子,但仅作为物体位置的粗略提示。本文提出将3D盒子作为结构化指令:用户输入并输出物体的3D框,将编辑视为明确的几何问题。该“以盒思”界面通过颜色编码盒子各面来表达3D朝向,实现对平移、旋转、缩放及视角变化的精确控制,同时保留场景与物体身份,并恢复此前未见的物体区域。为使变换与场景外观对齐,引入深度对齐的平面地板作为全局参考框架,通过深度感知线索着色。在该结构条件下,图像生成器可在大幅变换下生成一致结果。模型分两阶段训练——先在合成多物体场景上,再在少量来自Objectron的真实视频上——从而泛化至复杂真实图像。本方法直接作用于真实照片,在大规模3D编辑任务中显著优于近期最先进方法。
原文摘要 · Abstract (English)
Text and 2D-conditioning interfaces provide weak, ambiguous control over spatial transformations in image editing -- particularly under large object motions and camera changes. Prior work has used 3D primitives such as boxes, but only as loose conditioning signals indicating approximate object location rather than specifying the transformation. We instead use 3D boxes as structured specifications: the user provides the input and output boxes of the edit, casting editing as a well-posed geometry problem. This ``thinking in boxes'' interface, where each box face is color-coded to convey 3D orientation, gives precise control over translation, rotation, scaling, and viewpoint changes in real images while preserving scene and object identity, and recovering previously unseen object regions. To ground transformations in scene appearance, we introduce a depth-aligned planar floor as a global reference frame, shaded with depth-aware cues. Conditioned on this structure, an image generator produces consistent results under large transformations. Trained in two stages -- on synthetic multi-object scenes and a small set of real-world videos from Objectron -- the system generalizes to complex, in-the-wild real images. Our method operates directly on real photographs and substantially outperforms recent state-of-the-art methods on large 3D edits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。