单张全景图秒级重建高保真3D室内场景
FastPano3D: Feed-Forward Indoor Panoramic 3D Reconstruction from a Single Image

- 直接从全景图生成可渲染的3D高斯表示,无需多视图监督或优化
- 推理速度比SOTA快156倍,参数量减半,仍保持高质量渲染
- 适合需要快速3D重建的应用,如虚拟现实、数字孪生
近期3D场景重建研究揭示了渲染质量、推理效率与数据依赖之间的复杂权衡。为实现从极少输入中快速重建细节丰富的3D室内场景,我们提出FastPano3D,一个端到端框架,可直接从单张全景图生成可渲染的3D高斯表示。与基于透视图的方法不同,全景图存在等距圆柱投影畸变和空间非均匀特征分布,使直接前馈生成高斯表示极具挑战。不同于依赖多视图监督或每场景优化的现有高斯喷溅方法,FastPano3D采用轻量特征编码器、自适应高斯采样及点云引导精炼策略,在无需测试时优化的情况下实现高效准确的场景生成。该方法可在数秒内重建高保真3D场景,推理速度较Pano2Room等先进方法提升最高达156倍,同时仅使用其一半参数。大量实验表明,FastPano3D的渲染质量可媲美基于NeRF与3DGS的重建,树立了快速单视角3D场景推断的新基准。
原文摘要 · Abstract (English)
Recent advances in 3D scene reconstruction have highlighted the intricate trade-offs among rendering quality, inference efficiency, and data dependency. To address the challenge of rapidly reconstructing detailed 3D indoor scenes from minimal input, we introduce FastPano3D, an end-to-end framework that directly generates renderable 3D Gaussian representations from a single panoramic image. Unlike perspective-based methods, panoramic images inherently suffer from equirectangular projection distortions and spatially non-uniform feature distributions, making direct feed-forward Gaussian generation particularly challenging. In contrast to existing Gaussian Splatting based methods that rely on multi-view supervision or per-scene optimization, FastPano3D employs a lightweight feature encoder, adaptive Gaussian sampling, and a point-cloud-guided refinement strategy to achieve efficient and accurate scene generation without any test-time optimization. Our approach reconstructs high-fidelity 3D scenes within seconds, achieving up to 156 times faster inference than prior state-of-the-art methods such as Pano2Room, while using only half the parameters. Extensive experiments demonstrate that FastPano3D delivers rendering quality comparable to NeRF- and 3DGS-based reconstructions, establishing a new benchmark for rapid, single-view 3D scene inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。