arXiv:2605.13169cs.CVcs.AI2026-05被引 4

让大模型真正看懂360度全景图,突破传统视角局限。

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World

论文配图:PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World
图 1 · 摘自论文原文
  • 直接处理全景图的球面结构,不拆成多个视角。
  • 在多个评测中显著超越现有模型,最高提升27%。
  • 适合做机器人导航、虚拟现实等需要全局空间理解的任务。

多模态大模型在主流视角图像范式下仍难以实现空间理解,该范式继承了人类感知的窄视野局限。对于导航、机器人搜索和三维场景理解任务,360度全景感知通过一次性捕捉周围环境,提供一种超感能力。然而,现有多模态大模型流水线通常将全景图分解为多个视角视图,使等距圆柱投影(ERP)的球面结构几乎被忽略。本文研究全景原生理解,要求模型以连续、以观察者为中心的空间方式推理ERP全景图。我们首先定义关键能力:语义锚定、球面定位、参考系变换与深度感知的3D空间推理。随后构建大规模元数据生成管道,将混合来源的ERP全景图转化为几何感知、语言对齐、深度感知的监督信号,并构造与能力对齐的指令微调数据。在模型层面,提出PanoWorld,引入球面空间交叉注意力,将球面几何注入视觉流。进一步构建PanoSpace-Bench,用于评估ERP原生空间推理能力。实验表明,PanoWorld在PanoSpace-Bench、H* Bench和R2R-CE Val-Unseen基准上均显著优于开源及专有基线模型。结果表明,鲁棒全景推理需专用原生监督与几何感知模型适配。所有代码与数据将公开发布。

原文摘要 · Abstract (English)

Multimodal large laboratory models (MLLMs) still struggle with spatial understanding under the dominant perspective-image paradigm, which inherits the narrow field of view of human-like perception. For navigation, robotic search, and 3D scene understanding, 360-degree panoramic sensing offers a form of supersensing by capturing the entire surrounding environment at once. However, existing MLLM pipelines typically decompose panoramas into multiple perspective views, leaving the spherical structure of equirectangular projection (ERP) largely implicit. In this paper, we study pano-native understanding, which requires an MLLM to reason over an ERP panorama as a continuous, observer-centered space. To this end, we first define the key abilities for pano-native understanding, including semantic anchoring, spherical localization, reference-frame transformation, and depth-aware 3D spatial reasoning. We then build a large-scale metadata construction pipeline that converts mixed-source ERP panoramas into geometry-aware, language-grounded, and depth-aware supervision, and instantiate these signals as capability-aligned instruction tuning data. On the model side, we introduce PanoWorld with Spherical Spatial Cross-Attention, which injects spherical geometry into the visual stream. We further construct PanoSpace-Bench, a diagnostic benchmark for evaluating ERP-native spatial reasoning. Experiments show that PanoWorld substantially outperforms both proprietary and open-source baselines on PanoSpace-Bench, H* Bench, and R2R-CE Val-Unseen benchmarks. These results demonstrate that robust panoramic reasoning requires dedicated pano-native supervision and geometry-aware model adaptation. All source code and proposed data will be publicly released.

全景理解空间推理大模型360度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。