首个航空多光谱目标检测基准,提升小目标识别精度
MODA: The First Challenging Benchmark for Multispectral Object Detection in Aerial Images
- 构建多光谱图像融合框架,整合光谱与空间信息
- 在1.4万张图像上实现33万标注,显著提升检测性能
- 适合遥感、无人机目标检测研究者使用
航空目标检测在真实场景中面临小目标和大背景干扰等挑战,基于RGB的检测器因信息不足而性能受限。多光谱图像(MSIs)在多个波段捕捉额外光谱线索,具有潜力。然而,训练数据缺乏是制约其发展的主要瓶颈。为此,我们提出了首个大规模航空多光谱目标检测数据集MODA,包含14,041幅多光谱图像和330,191个标注,覆盖多样且具挑战性的场景,为该领域提供全面的数据基础。同时,针对航空多光谱检测的固有难题,我们提出OSSDet框架,通过级联式光谱-空间调制结构优化目标感知,利用光谱相似性聚合相关特征以增强对象内关联,并通过对象感知掩码抑制无关背景。此外,跨光谱注意力在显式对象引导下进一步精炼目标表示。大量实验表明,OSSDet在参数量和效率相当的情况下超越现有方法。
原文摘要 · Abstract (English)
Aerial object detection faces significant challenges in real-world scenarios, such as small objects and extensive background interference, which limit the performance of RGB-based detectors with insufficient discriminative information. Multispectral images (MSIs) capture additional spectral cues across multiple bands, offering a promising alternative. However, the lack of training data has been the primary bottleneck to exploiting the potential of MSIs. To address this gap, we introduce the first large-scale dataset for Multispectral Object Detection in Aerial images (MODA), which comprises 14,041 MSIs and 330,191 annotations across diverse, challenging scenarios, providing a comprehensive data foundation for this field. Furthermore, to overcome challenges inherent to aerial object detection using MSIs, we propose OSSDet, a framework that integrates spectral and spatial information with object-aware cues. OSSDet employs a cascaded spectral-spatial modulation structure to optimize target perception, aggregates spectrally related features by exploiting spectral similarities to reinforce intra-object correlations, and suppresses irrelevant background via object-aware masking. Moreover, cross-spectral attention further refines object-related representations under explicit object-aware guidance. Extensive experiments demonstrate that OSSDet outperforms existing methods with comparable parameters and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。