用一张全景图生成带结构布局的3D室内场景。
InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360° Image

- 基于全景图像,分三步重建空间结构与物体细节
- 在3D和2D指标上均优于现有方法,生成完整场景
- 适合虚拟现实、数字孪生等需要真实布局的场景
单图生成3D技术已实现高质量资产合成,但室内场景生成仍具挑战。现有方法侧重物体级生成,忽视了对下游应用至关重要的空间布局。单张图像视野有限,难以恢复连贯全局布局。为此,我们采用等距圆柱投影(ERP)表示的360°图像,提出InSpace框架,实现结构感知的3D室内场景生成。该框架包含三阶段:(1) 估计局部几何作为空间先验;(2) 通过视角选择性交叉注意力生成粗略结构;(3) 采用全局-局部混合注意力结合流匹配,生成精细布局与带纹理的物体几何。我们还构建了基于3D-FRONT的成对ERP-图像到3D场景数据集ERP-FRONT。实验表明,InSpace可从单张ERP图像生成完整3D室内场景及独立带纹理资产,在3D与2D评价指标上表现优异。
原文摘要 · Abstract (English)
Recent advances in single image-to-3D generation have enabled high-quality asset synthesis, yet extending these capabilities to indoor scene generation remains challenging. Existing methods focus on asset-level generation while neglecting the structural layout, which is essential for downstream applications and serves as the spatial anchor for grounding assets. However, a single image with a limited field of view lacks the spatial coverage to recover a coherent global layout. To this end, we use a 360° image represented in equirectangular projection (ERP) and propose InSpace, a structure-aware framework for 3D indoor scene generation. InSpace comprises three stages: (1) estimating partial scene geometry as spatial priors, (2) generating coarse scene structure with view-selective cross-attention, and (3) producing detailed layout and asset geometry with textures through a global-local hybrid attention, using flow matching. We also propose ERP-FRONT, a paired ERP-Image-to-3D indoor scene dataset based on 3D-FRONT. Experiments show that InSpace generates complete 3D indoor scenes with structural layout, along with separate textured assets from a single ERP image, achieving strong performance across 3D and 2D metrics. Project Page: https://kookie12.github.io/InSpace-Project-Page/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。