arXiv:2507.21960cs.CV2025-07ICCV被引 7

无需精确位姿即可重建全景图,效果超越现有方法。

PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction

  • 将视角预训练迁移至全景域,实现无位姿重建。
  • 在无位姿条件下,新视角生成和深度估计均优于当前最佳方法。
  • 适合真实场景中缺乏精确位姿数据的3D环境重建任务。

宽基线全景重建已成为实现周围3D环境几何重建及生成高度逼真沉浸式新视图的关键方法。尽管现有方法在多个基准测试中表现优异,但普遍依赖精确的位姿信息。在真实场景中,获取精准位姿需额外计算资源且易受噪声干扰,限制了方法的普适性与实用性。本文提出PanoSplatt3R,一种无位姿的宽基线全景重建方法。通过扩展并适配视角域的基础重建预训练至全景域,实现强大的泛化能力。为保障领域迁移的无缝高效,引入跨注意力头的RoPE滚动机制,以在旋转位置编码中建模全景图像的水平周期性,仅对原机制做最小修改。大量实验表明,即便无位姿信息,PanoSplatt3R仍显著优于当前最优方法,在高质量新视图生成与深度估计精度上均表现突出,展现出巨大实际应用潜力。

原文摘要 · Abstract (English)

Wide-baseline panorama reconstruction has emerged as a highly effective and pivotal approach for not only achieving geometric reconstruction of the surrounding 3D environment, but also generating highly realistic and immersive novel views. Although existing methods have shown remarkable performance across various benchmarks, they are predominantly reliant on accurate pose information. In real-world scenarios, the acquisition of precise pose often requires additional computational resources and is highly susceptible to noise. These limitations hinder the broad applicability and practicality of such methods. In this paper, we present PanoSplatt3R, an unposed wide-baseline panorama reconstruction method. We extend and adapt the foundational reconstruction pretrainings from the perspective domain to the panoramic domain, thus enabling powerful generalization capabilities. To ensure a seamless and efficient domain-transfer process, we introduce RoPE rolling that spans rolled coordinates in rotary positional embeddings across different attention heads, maintaining a minimal modification to RoPE's mechanism, while modeling the horizontal periodicity of panorama images. Comprehensive experiments demonstrate that PanoSplatt3R, even in the absence of pose information, significantly outperforms current state-of-the-art methods. This superiority is evident in both the generation of high-quality novel views and the accuracy of depth estimation, thereby showcasing its great potential for practical applications. Project page: https://npucvr.github.io/PanoSplatt3R

全景重建无位姿预训练迁移3D生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。