用轻量结构提升遥感小目标检测精度,兼顾效率与细节
Selective Structured State Space for Multispectral-fused Small Target Detection

- 引入增强型小目标模块与门控注意力机制,强化局部细节捕捉
- 多光谱融合模块有效提升小目标特征表达,准确率显著改善
- 适合遥感图像分析、安防监控等需高效识别微小目标场景
高分辨率遥感图像中的小目标检测面临识别精度低和计算成本高的双重挑战。传统Transformer的计算复杂度随图像分辨率呈平方增长,而CNN为扩大感受野需堆叠深层卷积,导致计算开销急剧上升。为此,本文采用Mamba的线性复杂度以提升效率,但其对小目标表现不佳,因小目标区域有限且语义信息少。为此,提出ESTD模块增强局部注意力,捕获细粒度特征;设计基于Mamba的CARG模块,聚焦空间与通道维度信息,共同提升模型对小目标的表征能力。此外,为强化小目标语义表达,构建了掩码增强的像素级多光谱融合(MEPF)模块,通过有效融合可见光与红外模态信息,显著增强目标特征。
原文摘要 · Abstract (English)
Target detection in high-resolution remote sensing imagery faces challenges due to the low recognition accuracy of small targets and high computational costs. The computational complexity of the Transformer architecture increases quadratically with image resolution, while Convolutional Neural Networks (CNN) architectures are forced to stack deeper convolutional layers to expand their receptive fields, leading to an explosive growth in computational demands. To address these computational constraints, we leverage Mamba's linear complexity for efficiency. However, Mamba's performance declines for small targets, primarily because small targets occupy a limited area in the image and have limited semantic information. Accurate identification of these small targets necessitates not only Mamba's global attention capabilities but also the precise capture of fine local details. To this end, we enhance Mamba by developing the Enhanced Small Target Detection (ESTD) module and the Convolutional Attention Residual Gate (CARG) module. The ESTD module bolsters local attention to capture fine-grained details, while the CARG module, built upon Mamba, emphasizes spatial and channel-wise information, collectively improving the model's ability to capture distinctive representations of small targets. Additionally, to highlight the semantic representation of small targets, we design a Mask Enhanced Pixel-level Fusion (MEPF) module for multispectral fusion, which enhances target features by effectively fusing visible and infrared multimodal information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。