用全景图一次性重建3D场景,解决畸变难题。
PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery
- 基于球面感知的Transformer架构,支持单次前向计算
- 在PanoCity数据集上实现高精度姿态与深度估计
- 适合需要全景3D重建的自动驾驶与VR应用
全景图像提供360°视野,日益普及于消费设备中。然而其非针孔畸变给联合姿态估计与3D重建带来挑战。现有前馈模型针对透视相机设计,难以泛化至该场景。我们提出PanoVGGT,一种排列等变的Transformer框架,可单次前向传播同时预测相机姿态、深度图与3D点云。模型引入球面感知位置编码及全景特有三轴SO(3)旋转增强,实现球面域有效几何推理。为解决固有的全局坐标系模糊性,训练中引入随机锚定策略。此外,我们构建了大规模户外全景数据集PanoCity,含密集深度与6-DoF姿态标注。在PanoCity及标准基准上的大量实验表明,PanoVGGT达到具有竞争力的精度、强鲁棒性与提升的跨域泛化能力。代码与数据集将公开。
原文摘要 · Abstract (English)
Panoramic imagery offers a full 360° field of view and is increasingly common in consumer devices. However, it introduces non-pinhole distortions that challenge joint pose estimation and 3D reconstruction. Existing feed-forward models, built for perspective cameras, generalize poorly to this setting. We propose PanoVGGT, a permutation-equivariant Transformer framework that jointly predicts camera poses, depth maps, and 3D point clouds from one or multiple panoramas in a single forward pass. The model incorporates spherical-aware positional embeddings and a panorama-specific three-axis SO(3) rotation augmentation, enabling effective geometric reasoning in the spherical domain. To resolve inherent global-frame ambiguity, we further introduce a stochastic anchoring strategy during training. In addition, we contribute PanoCity, a large-scale outdoor panoramic dataset with dense depth and 6-DoF pose annotations. Extensive experiments on PanoCity and standard benchmarks demonstrate that PanoVGGT achieves competitive accuracy, strong robustness, and improved cross-domain generalization. Code and dataset will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。