对比11种深度模型在卫星图像目标检测中的表现,发现基于Transformer的模型更优。
Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
- 采用11种检测算法,包含7个2020年后发表的Transformer模型
- 在3个开源高分辨率遥感数据集上验证,共训练33个模型
- 首次系统比较现代卷积与Transformer在遥感图像上的性能差异
2012年,AlexNet确立了深度卷积神经网络(DCNNs)在计算机视觉领域的领先地位,随后其在遥感等多个领域取得显著进展。随着视觉变换器(Visual Transformers)的出现,计算视觉迎来第二次现代飞跃。因此,理解各类基于变换器的神经网络在卫星影像上的表现至关重要。尽管变换器在自然语言处理和计算机视觉中表现出色,但尚未在大规模现代遥感数据上进行系统比较。本文研究基于变换器的神经网络在高分辨率光电卫星影像中的目标检测应用,在多个公开基准数据集上实现了前沿性能。研究对比了11种不同的边界框检测与定位算法,其中7个为2020年后发表,全部11个均发布于2015年之后。将5种基于变换器的架构与6种卷积网络在三个最新开源高分辨率遥感影像数据集上进行比较,这些数据集在规模和复杂度上各不相同。在完成33个深度神经网络的训练与评估后,进一步讨论并分析了不同特征提取方法与检测算法的性能表现。
原文摘要 · Abstract (English)
In 2012, AlexNet established deep convolutional neural networks (DCNNs) as the state-of-the-art in CV, as these networks soon led in visual tasks for many domains, including remote sensing. With the publication of Visual Transformers, we are witnessing the second modern leap in computational vision, and as such, it is imperative to understand how various transformer-based neural networks perform on satellite imagery. While transformers have shown high levels of performance in natural language processing and CV applications, they have yet to be compared on a large scale to modern remote sensing data. In this paper, we explore the use of transformer-based neural networks for object detection in high-resolution electro-optical satellite imagery, demonstrating state-of-the-art performance on a variety of publicly available benchmark data sets. We compare eleven distinct bounding-box detection and localization algorithms in this study, of which seven were published since 2020, and all eleven since 2015. The performance of five transformer-based architectures is compared with six convolutional networks on three state-of-the-art opensource high-resolution remote sensing imagery datasets ranging in size and complexity. Following the training and evaluation of thirty-three deep neural models, we then discuss and analyze model performance across various feature extraction methodologies and detection algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。