arXiv:2509.20107cs.CVcs.AI2025-09被引 6

用视觉大模型提升高光谱图像语义分割效果

Hyperspectral Adapter for Semantic Segmentation with Vision Foundation Models

  • 设计光谱变换器与光谱感知空间先验模块提取多维特征
  • 在三个自动驾驶数据集上超越现有方法,实现最佳性能
  • 适合做高光谱遥感、机器人感知等需要精细材质识别的场景

高光谱成像(HSI)在多个窄波段中捕获密集光谱信息,具备增强机器人感知能力的潜力,尤其适用于材料复杂、光照变化或视觉挑战大的环境。然而,现有高光谱语义分割方法因依赖为RGB输入优化的架构和学习框架,表现不佳。本文提出一种新型高光谱适配器,利用预训练视觉基础模型有效学习高光谱数据。该架构引入光谱变换器和光谱感知空间先验模块,以提取丰富的空间-光谱特征;同时设计模态感知交互模块,通过专用提取与注入机制,实现高光谱表示与冻结视觉Transformer特征的有效融合。在三个基准自动驾驶数据集上的大量实验表明,所提方法直接使用高光谱输入,达到当前最优的语义分割性能,优于基于视觉和传统高光谱的分割方法。代码已公开于 https://hsi-adapter.cs.uni-freiburg.de。

原文摘要 · Abstract (English)

Hyperspectral imaging (HSI) captures spatial information along with dense spectral measurements across numerous narrow wavelength bands. This rich spectral content has the potential to facilitate robust robotic perception, particularly in environments with complex material compositions, varying illumination, or other visually challenging conditions. However, current HSI semantic segmentation methods underperform due to their reliance on architectures and learning frameworks optimized for RGB inputs. In this work, we propose a novel hyperspectral adapter that leverages pretrained vision foundation models to effectively learn from hyperspectral data. Our architecture incorporates a spectral transformer and a spectrum-aware spatial prior module to extract rich spatial-spectral features. Additionally, we introduce a modality-aware interaction block that facilitates effective integration of hyperspectral representations and frozen vision Transformer features through dedicated extraction and injection mechanisms. Extensive evaluations on three benchmark autonomous driving datasets demonstrate that our architecture achieves state-of-the-art semantic segmentation performance while directly using HSI inputs, outperforming both vision-based and hyperspectral segmentation methods. We make the code available at https://hsi-adapter.cs.uni-freiburg.de.

高光谱语义分割视觉大模型遥感感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。