用双向光谱-空间注意力提升无人机高光谱视觉感知效率
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
- 设计双向跨注意力模块,融合光谱与空间特征
- 在WHU-Hi-HongHu数据集上实现98.2%总体精度
- 适合需要实时感知的无人机导航与环境监测场景
本研究针对复杂环境下无人机导航可靠性下降的问题,提出一种基于高光谱成像(HSI)的深度学习架构。通过改进Mobile 3D Vision Transformer(MDvT),引入SpectralCA模块,利用双向跨注意力机制融合光谱与空间特征,在减少参数量和推理时间的同时提升性能。在WHU-Hi-HongHu数据集上的实验表明,该方法在总体精度、平均精度和Kappa系数上均优于现有方法,实现了无人机在导航、目标检测与地形分类任务中的高效实时感知。
原文摘要 · Abstract (English)
The relevance of this research lies in the growing demand for unmanned aerial vehicles (UAVs) capable of operating reliably in complex environments where conventional navigation becomes unreliable due to interference, poor visibility, or camouflage. Hyperspectral imaging (HSI) provides unique opportunities for UAV-based computer vision by enabling fine-grained material recognition and object differentiation, which are critical for navigation, surveillance, agriculture, and environmental monitoring. The aim of this work is to develop a deep learning architecture integrating HSI into UAV perception for navigation, object detection, and terrain classification. Objectives include: reviewing existing HSI methods, designing a hybrid 2D/3D convolutional architecture with spectral-spatial cross-attention, training, and benchmarking. The methodology is based on the modification of the Mobile 3D Vision Transformer (MDvT) by introducing the proposed SpectralCA block. This block employs bi-directional cross-attention to fuse spectral and spatial features, enhancing accuracy while reducing parameters and inference time. Experimental evaluation was conducted on the WHU-Hi-HongHu dataset, with results assessed using Overall Accuracy, Average Accuracy, and the Kappa coefficient. The findings confirm that the proposed architecture improves UAV perception efficiency, enabling real-time operation for navigation, object recognition, and environmental monitoring tasks. Keywords: SpectralCA, deep learning, computer vision, hyperspectral imaging, unmanned aerial vehicle, object detection, semi-supervised learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。