针对密集分布的SAR图像目标检测,提出新型视觉Transformer模型
DenSe-AdViT: A novel Vision Transformer for Dense SAR Object Detection
- 设计密度感知模块,捕捉目标空间分布与密度信息
- 在RSDD和SIVED数据集上分别达79.8%和92.5% mAP
- 适合处理高密度车辆目标的遥感图像检测任务
视觉变换器(ViT)在合成孔径雷达(SAR)图像目标检测中表现优异,因其强大的全局特征提取能力。然而,其在多尺度局部特征提取方面存在不足,导致对小目标尤其是密集排列目标的检测性能受限。为此,我们提出用于密集SAR目标检测的密度敏感自适应标记视觉变换器(DenSe-AdViT)。设计密度感知模块(DAM),基于目标分布生成密度张量,并通过精心设计的目标函数实现对空间分布与密度的精准捕获。为融合卷积神经网络(CNN)增强的多尺度信息与变换器提取的全局特征,提出密度增强融合模块(DEFM),利用密度掩码和多源特征有效优化注意力聚焦于目标存活区域。实验表明,所提方法在包含大量密集分布车辆目标的RSDD与SIVED数据集上分别达到79.8%和92.5%的mAP。
原文摘要 · Abstract (English)
Vision Transformer (ViT) has achieved remarkable results in object detection for synthetic aperture radar (SAR) images, owing to its exceptional ability to extract global features. However, it struggles with the extraction of multi-scale local features, leading to limited performance in detecting small targets, especially when they are densely arranged. Therefore, we propose Density-Sensitive Vision Transformer with Adaptive Tokens (DenSe-AdViT) for dense SAR target detection. We design a Density-Aware Module (DAM) as a preliminary component that generates a density tensor based on target distribution. It is guided by a meticulously crafted objective metric, enabling precise and effective capture of the spatial distribution and density of objects. To integrate the multi-scale information enhanced by convolutional neural networks (CNNs) with the global features derived from the Transformer, Density-Enhanced Fusion Module (DEFM) is proposed. It effectively refines attention toward target-survival regions with the assist of density mask and the multiple sources features. Notably, our DenSe-AdViT achieves 79.8% mAP on the RSDD dataset and 92.5% on the SIVED dataset, both of which feature a large number of densely distributed vehicle targets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。