arXiv:2608.06836cs.CV2026-08

让家具在单视角图像中真实摆放,靠生成式3D姿态推理解决尺度不确定问题。

GOPI: Generation-Oriented 3D Pose Inference for Furniture Insertion from Single-View RGB-D Indoor Scenes

论文配图:GOPI: Generation-Oriented 3D Pose Inference for Furniture Insertion from Single-View RGB-D Indoor Scenes
图 1 · 摘自论文原文
  • 用数据驱动迭代法推断家具在3D空间的合理位置。
  • 相比直接回归,生成姿势更符合真实布局,且与图像对齐稳定。
  • 适合做室内场景家具插入的生成模型研究者或设计师使用。

我们研究如何将新家具插入室内场景图像。在仅用单视角2D图像遮挡条件的情况下,插入家具相对于场景的物理尺度无法唯一确定,导致基于图像证据的物理合理摆放难以求解。为此,我们将任务重构为3D姿态推断与几何引导图像生成的结合,其中估计几何上合理的3D放置是可靠合成的关键。为此,我们提出两阶段框架:对于3D放置,引入GOPI——一种面向生成的3D姿态推断框架,通过数据驱动的迭代推理解决单视角家具插入的欠定性,生成几何上合理的物体位置;对于图像生成,设计一种几何引导的条件策略,将推断出的3D姿态投影到图像平面作为像素对齐约束,确保合成图像与底层3D几何的一致性。实验结果从3D姿态估计和图像合成两个角度验证了该框架的有效性。在3D放置方面,GOPI生成的姿态具有更强的几何可行性,并与参考布局有更好的一致性,优于直接回归和基线方法;在图像生成方面,我们的方法在不同家具尺度下均保持与投影3D几何的一致性,展现出稳定的投影-生成对齐能力。

原文摘要 · Abstract (English)

We study the problem of inserting new furniture into indoor scene images. Under masked single-view 2D image-plane conditioning, however, the physical scale of the inserted furniture relative to the scene cannot be uniquely determined, making physically grounded furniture placement underdetermined from image evidence alone. We therefore reformulate the task as a combination of 3D pose inference and geometry-guided image generation, where estimating a geometrically plausible 3D placement is essential for reliable synthesis. To this end, we propose a two-stage framework. For 3D placement, we introduce GOPI, a generation-oriented 3D pose inference framework that addresses the underdetermined nature of single-view furniture insertion through data-driven iterative inference, producing geometrically plausible object placements. For image generation, we develop a geometry-guided conditioning strategy that projects the inferred 3D pose into the image plane as a pixel-aligned constraint, enforcing consistency between the synthesized image and the underlying 3D geometry. Experimental results validate the proposed framework from both 3D pose estimation and image synthesis perspectives. For 3D placement, GOPI produces poses with stronger geometric feasibility and better consistency with reference layouts than direct regression and vanilla baselines. For image synthesis, our method preserves alignment with the projected 3D geometry across different furniture scales, showing stable projection-generation alignment across the tested furniture scales.

3D姿态家具插入图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。