arXiv:2603.13910cs.CV2026-03

用文本生成精确可导航的3D室内场景,解决尺度混乱问题

Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation

  • 基于文本预测全局3D布局,保持绝对坐标系一致
  • 融合全景扩散与视频扩散,实现10倍速高效场景覆盖
  • 支持渐进扩展和物体姿态精准传递,适合虚拟空间设计

我们提出GuidedSceneGen,一种文本到3D的生成框架,可生成具有度量精度、全局一致且语义可解释的室内场景。与以往存在几何漂移或尺度模糊的问题不同,本方法在整个生成过程中维持绝对世界坐标系。从文本描述出发,首先预测包含语义与几何结构的全局3D布局,作为后续阶段的引导代理。接着,利用语义与深度条件的全景扩散模型生成与全局布局对齐的360°图像,显著提升空间一致性。为探索未观测区域,采用由优化相机轨迹引导的视频扩散模型,在覆盖范围与避障之间取得平衡,采样速度较全路径探索快10倍。生成视图通过3D高斯点云融合,构建出一致且可全向导航的3D场景。该方法实现了从布局到重建的物体姿态与语义标签精准迁移,并支持无需重新对齐的渐进式场景扩展。定量结果与用户研究均表明,其3D一致性与布局合理性优于近期全景文本到3D基线方法。

原文摘要 · Abstract (English)

We present GuidedSceneGen, a text-to-3D generation framework that produces metrically accurate, globally consistent, and semantically interpretable indoor scenes. Unlike prior text-driven methods that often suffer from geometric drift or scale ambiguity, our approach maintains an absolute world coordinate frame throughout the entire generation process. Starting from a textual scene description, we predict a global 3D layout encoding both semantic and geometric structure, which serves as a guiding proxy for downstream stages. A semantics- and depth-conditioned panoramic diffusion model then synthesizes 360° imagery aligned with the global layout, substantially improving spatial coherence. To explore unobserved regions, we employ a video diffusion model guided by optimized camera trajectories that balances coverage and collision avoidance, achieving up to 10x faster sampling compared to exhaustive path exploration. The generated views are fused using 3D Gaussian Splatting, yielding a consistent and fully navigable 3D scene in absolute scale. GuidedSceneGen enables accurate transfer of object poses and semantic labels from layout to reconstruction, and supports progressive scene expansion without re-alignment. Quantitative results and a user study demonstrate greater 3D consistency and layout plausibility compared to recent panoramic text-to-3D baselines.

3D生成文本生成室内场景扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。