arXiv:2608.29081cs.CV2026-08

提出生物启发的自适应全景分割模型,提升复杂视角下的鲁棒性。

AdapToPASS: Ambiguity-aware Adaptive Spherical Transformer for Panoramic Semantic Segmentation

论文配图:AdapToPASS: Ambiguity-aware Adaptive Spherical Transformer for Panoramic Semantic Segmentation
图 1 · 摘自论文原文
  • 基于生物感知机制,动态调整注意力以应对上下文与几何模糊。
  • 在未见过的球面变换下,斯坦福2D3D上提升13.38% mIoU,WildPASS上提升18.77%。
  • 轻量版参数少于200万,兼顾效率与对球面变换的鲁棒性。

球面变换器因其直接在球面几何上操作、缓解投影畸变,成为全景语义分割(PASS)的有前景框架。然而现有架构常假设标准球面结构和稳定视角,这在真实图像中因自由相机运动而难以满足,引入上下文与几何模糊。因此它们缺乏处理此类模糊的自适应机制,限制了对未见球面变换的鲁棒性。相比之下,生物感知天生具备模糊感知能力,能根据线索可靠性变化自适应调整,以维持复杂变换下的稳定解释。受此启发,我们系统分析现有PASS架构在多种未见球面变换下的表现,并提出AdapToPASS——一种新型生物启发球面变换器,可自适应建模上下文与几何模糊以实现鲁棒的PASS。其核心为自适应球面注意力(AdaSpA)模块,根据局部上下文模糊度动态调节注意力,模拟生物感知的上下文驱动适应性。为应对几何模糊,AdapToPASS采用双焦点球面表示,在视场与空间分辨率间取得平衡,并引入边界监督,借鉴生物视觉对边界的敏感性。在室内外语义分割任务中,AdapToPASS持续优于现有最先进方法。在未见球面变换下,其相对最佳基线在Stanford2D3D上提升+13.38% mIoU,WildPASS上提升+18.77%。我们还提出AdapToPASS-Swift,一种参数少于200万的轻量版本,在超越紧凑基线的同时保持对球面变换的鲁棒性。

原文摘要 · Abstract (English)

Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures often assume canonical spherical structure and stable viewpoints, which are frequently violated in real-world imagery due to unconstrained camera motion, introducing contextual and geometric ambiguity. Consequently, they lack adaptive mechanisms to handle such ambiguity, limiting robustness to unseen spherical transformations. In contrast, biological perception is inherently ambiguity-aware, adapting to fluctuations in cue reliability caused by geometric and contextual variations to maintain stable interpretation under complex transformations. Motivated by this, we first systematically analyze existing PASS architectures under various unseen spherical transformations. We then introduce AdapToPASS, a novel bio-inspired Spherical Transformer that adaptively models contextual and geometric ambiguities for robust PASS. At its core, Adaptive Spherical Attention (AdaSpA) blocks dynamically modulate attention according to local contextual ambiguity, mimicking adaptive, context-driven biological perception. To address geometric ambiguity, AdapToPASS employs Bifocal Spherical Representation to balance field of view and spatial resolution, together with boundary supervision inspired by the boundary-sensitive nature of biological vision. Across indoor and outdoor semantic segmentation, AdapToPASS consistently outperforms prior state-of-the-art methods. Under unseen spherical transformations, it surpasses the next-best method by +13.38% relative mIoU on Stanford2D3D and +18.77% on WildPASS. We further introduce AdapToPASS-Swift, a lightweight variant with fewer than 2M parameters, which surpasses compact baselines while retaining robustness to spherical transformations.

全景分割球面几何生物启发自适应注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。