全景视觉从投影适配走向球面原生建模,突破几何畸变瓶颈。
Panoramic Scene Understanding: A Survey from Distortion-Aware Engineering to Sphere-Native Modeling

- 提出球面原生建模与几何感知令牌化,融合球面几何与大模型复用
- 现有方法多在旋转不变性与预训练复用间取平衡,动态感知仍平面化
- 深度、布局等任务进展不均,缺乏球面数据预训练基础模型
全景图像在单帧内捕捉完整视觉球面,提供传统相机无法获取的上下文。然而这种完整性伴随不可避免的几何代价:2-球面无法无失真映射到平面,所有投影均引入畸变,挑战标准视觉架构。本综述追溯全景场景理解从基于投影的适配与畸变感知工程,到球面原生建模的发展历程,反映对球面几何日益增强的重视。基础模型构成第四类,采用几何感知令牌化,在保留透视预训练权重的同时适配输入接口。我们跨五类任务进行回顾:密集预测、统一多任务理解、开放世界感知、视觉-语言推理与动态视频分析。各任务中,向球面几何的转向反复出现。然而实践中,领域并未选择最强球面原生算子(严格旋转等变但无法复用透视预训练主干,因而未规模化),而是趋向兼容性保持的中间路径:适度几何感知结合大规模预训练模型。此承诺程度不一:在密集预测中最为深入,而在动态感知中仍较浅,方法空间上球面感知,时间上仍平面化。基础模型适应推动全景深度估计最快,而布局、表面法向量与视频级理解仍基本未被探索。目前尚未有全景基础模型在球面数据上预训练。我们识别出五大评估缺口:球面面积加权指标、接缝一致性测试、极点鲁棒性分层、跨投影泛化能力与标准化开放世界协议。最后提出六点路线图,迈向通用全景智能。
原文摘要 · Abstract (English)
Panoramic images capture the full visual sphere in a frame, offering context unavailable to conventional cameras. Yet this completeness has an unavoidable geometric cost: the 2-sphere cannot be faithfully mapped to the plane, and every projection introduces distortions that challenge standard vision architectures. This survey traces panoramic scene understanding from projection-based adaptation and distortion-aware engineering to sphere-native modeling, reflecting increasing commitment to spherical geometry. Foundation models form a fourth family, geometry-aware tokenization, which adapts the input interface while reusing perspective-pretrained weights. We review these approaches across five task families: dense prediction, unified multi-task understanding, open-world perception, vision-language reasoning, and dynamic video analysis. Across tasks, the same shift toward spherical geometry recurs. In practice, however, the field has converged not on the strongest sphere-native operators, which are exactly rotation-equivariant but cannot reuse perspective-pretrained backbones and thus have not scaled, but on a compatibility-preserving middle ground combining moderate geometric awareness with large pretrained models. This commitment is uneven: deepest in dense prediction and shallowest in dynamic perception, where methods are spatially sphere-aware yet temporally planar. Foundation-model adaptation has advanced panoramic depth fastest, while layout, surface-normal, and video-level understanding remain largely unexplored. No panoramic foundation model has yet been pretrained on spherical data. We identify five evaluation gaps: spherical-area-weighted metrics, seam-consistency tests, polar-robustness stratification, cross-projection generalization, and standardized open-world protocols. We conclude with a six-point roadmap toward general-purpose panoramic intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。