arXiv:2601.07447cs.CV2026-01中稿 · ICPR 2026被引 4

将SAM模型适配全景图像分割,通过双视角融合提升精度。

PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion

  • 用多阶段特征和跨模态融合增强全景图理解能力
  • 在Stanford2D3DS和Matterport3D上达当前最佳性能
  • 适合全景图像语义分割与多模态感知研究者

现有图像基础模型主要基于透视图像训练,不适用于球面图像。PanoSAMic将预训练的分割一切(SAM)编码器引入全景图像语义分割任务,结合多模态输入。我们改进SAM编码器以输出多级特征,并设计新颖的时空-模态融合模块,使模型能为输入不同区域选择最相关模态与特征。此外,语义解码器采用球面注意力机制与双视角融合策略,缓解全景图像中的畸变与边缘不连续问题。PanoSAMic在Stanford2D3DS数据集上对RGB、RGB-D及RGB-D-N模态,在Matterport3D数据集上对RGB和RGB-D模态均取得当前最优结果。

原文摘要 · Abstract (English)

Existing image foundation models are not optimized for spherical images having been trained primarily on perspective images. PanoSAMic integrates the pre-trained Segment Anything (SAM) encoder to make use of its extensive training and integrate it into a semantic segmentation model for panoramic images using multiple modalities. We modify the SAM encoder to output multi-stage features and introduce a novel spatio-modal fusion module that allows the model to select the relevant modalities and best features from each modality for different areas of the input. Furthermore, our semantic decoder uses spherical attention and dual view fusion to overcome the distortions and edge discontinuity often associated with panoramic images. PanoSAMic achieves state-of-the-art (SotA) results on Stanford2D3DS for RGB, RGB-D, and RGB-D-N modalities and on Matterport3D for RGB and RGB-D modalities. https://github.com/dfki-av/PanoSAMic

全景分割SAM多模态球面图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。