用少量图片重建可编辑的3D房间,自动推断结构与物体位置。
Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs

- 通过可执行场景程序表示,从少视角图推断房间结构和物体布局。
- 在11万张多视角图像上训练,实现视觉证据一致的精确空间重构。
- 适合需要快速生成可编辑3D场景的工业设计与虚拟现实应用。
从少量未标定的RGB视角合成结构化且可编辑的3D室内场景,不仅需生成高质量单个物体,还需推断房间结构、跨不完整观测关联物体,并恢复全局一致的空间配置。以往方法多依赖文本输入或连续视觉输入及额外先验(如人工标注掩码或精确3D布局),导致劳动密集且难以通用。本文提出 extsc{Scenix},一种基于可执行场景程序的稀疏视图3D场景重建框架,该结构化表示可直接实例化为可编辑3D场景。给定稀疏视图, extsc{Scenix}通过感知引导的资产实例化与闭环空间优化预测可执行场景程序。为此,我们构建了约11万张合成与真实室内场景的 extsc{Dataset},包含多视角图像、房间结构、以物体为中心的描述及度量空间标注。进一步引入观察一致监督,使目标场景与输入视图中的视觉证据对齐。在保留的 extsc{XScene}场景、真实室内图像及分布外的SpatialGen案例上,评估了结构化场景预测、物体定位与空间优化性能。
原文摘要 · Abstract (English)
Synthesizing a structured and editable 3D indoor scene from a few uncalibrated RGB views requires more than generating high-quality individual assets: a system must infer the room structure, associate objects across incomplete observations, and recover a globally consistent spatial configuration. Previous methods mainly focus on 3D scene generation with text input or require continuous visual inputs with additional priors, \ e.g., human-annotated masks or accurate 3D layouts, which makes these methods labor demanding and hard to apply in general cases. We present \textsc{Scenix}, a sparse-view 3D scene reconstruction framework via executable scene programs, a structured representation that can be directly instantiated into editable 3D scenes. Given sparse views, \textsc{Scenix} predicts executable scene programs through perception-grounded asset instantiation and closed-loop spatial refinement. % We present \method, a framework that predicts an executable scene representation from sparse views and realizes it through perception-grounded asset instantiation and closed-loop spatial refinement. To support this task, we construct \dataset, a dataset of approximately 110,000 synthetic and real indoor scenes with multiview imagery, room structures, object-centric descriptions, and metric spatial annotations. We further introduce observation-consistent supervision that aligns each target scene with the visual evidence available in its input views. Experiments on held-out \textsc{XScene} scenes, real indoor images, and out-of-distribution SpatialGen cases evaluate structured scene prediction, object grounding, and spatial refinement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。