提出多尺度光谱注意力模块,提升自动驾驶中高光谱图像分割精度
Multi-Scale Spectral Attention Module-based Hyperspectral Segmentation in Autonomous Driving Scenarios
- 用不同卷积核并行提取光谱特征,自适应融合增强表达能力
- 在多个数据集上平均提升mIoU 2.32%、mF1 2.88%
- 适合需要高光谱感知的自动驾驶视觉系统研发者
自动驾驶领域近年关注高光谱成像(HSI)在复杂天气与光照下的环境感知潜力。然而,高效处理高维光谱数据仍是挑战。本文提出一种多尺度注意力机制(MSAM),通过三个不同核大小(1-11)的并行一维卷积和自适应特征聚合,增强光谱特征提取。将MSAM嵌入UNet的跳跃连接中,在多个城市驾驶场景的HSI数据集上评估语义分割性能。消融实验表明,MSAM持续优于基线UNet-SC,平均提升mIoU 2.32%、mF1 2.88%,且在GPU性能上保持竞争力。研究发现最优卷积核组合具有数据集特异性,如(1;5;11)与(3;7;11)表现突出。该实证研究深化了对自动驾驶中HSI处理能力的理解,为车载部署中的自适应多尺度光谱特征提取奠定基础。
原文摘要 · Abstract (English)
Recent advances in autonomous driving (AD) have highlighted the potential of hyperspectral imaging (HSI) for enhanced environmental perception, particularly in challenging weather and lighting conditions. However, efficiently processing high-dimensional spectral data remains a significant challenge. This paper presents an empirical investigation of a Multi-Scale Attention Mechanism (MSAM) for enhanced spectral feature extraction through three parallel 1D convolutions with varying kernel sizes (1-11) and adaptive feature aggregation. By integrating MSAM into UNet's skip connections, we evaluate performance improvements in semantic segmentation across multiple HSI datasets for urban driving scenarios. Comprehensive ablation studies demonstrate that MSAM consistently outperforms baseline UNet-SC, achieving average improvements of 2.32% in mIoU and 2.88% in mF1, while maintaining competitive GPU performance against established attention mechanisms. Our findings reveal that optimal kernel combinations are dataset-specific, with configurations such as (1;5;11) and (3;7;11) demonstrating particularly strong performance. This empirical investigation advances understanding of HSI processing capabilities for AD applications and establishes a foundation for adaptive multi-scale spectral feature extraction in automotive deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。