用视觉地图精准控制物体在场景中的位置与朝向。
Controllable 3D Placement of Objects with Scene-Aware Diffusion Models
- 设计视觉地图+粗略掩码,实现高精度物体定位。
- 支持形状和朝向变化,保持背景完整不变。
- 适合汽车场景中对位置与外观的精细控制。
随着文本条件生成模型的发展,图像编辑能力显著提升。然而,精确控制物体在环境中的位置和朝向仍具挑战性,通常需精心设计的修复掩码或提示词。本文提出一种基于视觉地图与粗略物体掩码的条件信号,可有效解决歧义问题,同时支持物体形状和朝向的变化。该方法基于修复模型,天然保留背景完整性,不同于联合建模对象与背景的方法。我们在汽车场景中验证了该方法的有效性,通过新设计的物体放置任务评估编辑质量,不仅关注外观一致性,还衡量姿态与位置准确性,包括需复杂形变的情形。最后,展示了位置控制与外观控制的协同应用,可将现有物体精确放置于场景指定位置。
原文摘要 · Abstract (English)
Image editing approaches have become more powerful and flexible with the advent of powerful text-conditioned generative models. However, placing objects in an environment with a precise location and orientation still remains a challenge, as this typically requires carefully crafted inpainting masks or prompts. In this work, we show that a carefully designed visual map, combined with coarse object masks, is sufficient for high quality object placement. We design a conditioning signal that resolves ambiguities, while being flexible enough to allow for changing of shapes or object orientations. By building on an inpainting model, we leave the background intact by design, in contrast to methods that model objects and background jointly. We demonstrate the effectiveness of our method in the automotive setting, where we compare different conditioning signals in novel object placement tasks. These tasks are designed to measure edit quality not only in terms of appearance, but also in terms of pose and location accuracy, including cases that require non-trivial shape changes. Lastly, we show that fine location control can be combined with appearance control to place existing objects in precise locations in a scene.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。