arXiv:2505.24489cs.CVcs.AI2025-05中稿 · the 29th Internati…被引 3

用可变形注意力提升遥感图像目标检测精度

Deformable Attention Mechanisms Applied to Object Detection, case of Remote Sensing

  • 引入可变形注意力机制改进DETR模型,适应遥感图像特征
  • 光学与SAR数据集上分别达95.12%和94.54%的F1分数
  • 适合遥感目标检测、多源影像分析的研究者参考

目标检测在遥感领域具有重要意义,因其图像具备稳定的地理覆盖和对象一致性。深度学习模型,特别是基于Transformer的架构,在视觉计算任务中表现突出。本文提出将采用可变形注意力机制的Deformable-DETR模型应用于遥感图像,涵盖光学与合成孔径雷达(SAR)两种模式。实验使用两个数据集:光学的Pleiades Aircraft数据集和SAR的船舶检测数据集(SSDD)。通过10折分层验证,模型在光学数据集上取得95.12%的F1分数,在SSDD上达94.54%,显著优于多种基于CNN和Transformer的基准模型,尤其在复杂遥感场景下表现优异。

原文摘要 · Abstract (English)

Object detection has recently seen an interesting trend in terms of the most innovative research work, this task being of particular importance in the field of remote sensing, given the consistency of these images in terms of geographical coverage and the objects present. Furthermore, Deep Learning (DL) models, in particular those based on Transformers, are especially relevant for visual computing tasks in general, and target detection in particular. Thus, the present work proposes an application of Deformable-DETR model, a specific architecture using deformable attention mechanisms, on remote sensing images in two different modes, especially optical and Synthetic Aperture Radar (SAR). To achieve this objective, two datasets are used, one optical, which is Pleiades Aircraft dataset, and the other SAR, in particular SAR Ship Detection Dataset (SSDD). The results of a 10-fold stratified validation showed that the proposed model performed particularly well, obtaining an F1 score of 95.12% for the optical dataset and 94.54% for SSDD, while comparing these results with several models detections, especially those based on CNNs and transformers, as well as those specifically designed to detect different object classes in remote sensing images.

目标检测遥感图像可变形注意力Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。