用混合Mamba-CNN架构提升遥感小目标检测精度
RemoteDet-Mamba: A Hybrid Mamba-CNN Network for Multi-modal Object Detection in Remote Sensing Images
- 采用四方向扫描融合策略,同步提取局部特征与跨模态全局语义
- 在DroneVehicle数据集上优于主流方法,参数量和计算开销更低
- 适合无人机遥感中密集小目标检测,部署效率高
无人机遥感因信息获取快、成本低,广泛应用于应急响应等场景。但受成像距离远、成像机制复杂影响,遥感图像中的目标常存在尺寸小、分布密集、类间区分度低等问题。为此,本文提出一种名为RemoteDet-Mamba的多模态遥感目标检测网络,基于逐块四方向选择性扫描融合策略,同时学习单模态局部特征并融合跨模态块级全局语义信息,增强小目标可区分性,提升类间判别能力。所设计的轻量化融合机制有效解耦密集目标,降低计算复杂度。在DroneVehicle数据集上的实验表明,RemoteDet-Mamba相比当前主流方法表现更优,且参数量少、计算开销低,展现出良好的实际应用潜力。
原文摘要 · Abstract (English)
Unmanned Aerial Vehicle (UAV) remote sensing, with its advantages of rapid information acquisition and low cost, has been widely applied in scenarios such as emergency response. However, due to the long imaging distance and complex imaging mechanisms, targets in remote sensing images often face challenges such as small object size, dense distribution, and low inter-class discriminability. To address these issues, this paper proposes a multi-modal remote sensing object detection network called RemoteDet-Mamba, which is based on a patch-level four-direction selective scanning fusion strategy. This method simultaneously learns unimodal local features and fuses cross-modal patch-level global semantic information, thereby enhancing the distinguishability of small objects and improving inter-class discrimination. Furthermore, the designed lightweight fusion mechanism effectively decouples densely packed targets while reducing computational complexity. Experimental results on the DroneVehicle dataset demonstrate that RemoteDet-Mamba achieves superior detection performance compared to current mainstream methods, while maintaining low parameter count and computational overhead, showing promising potential for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。