arXiv:2609.01800cs.CV2026-09

对比深度神经网络架构,提升合成孔径声呐目标识别准确率

Improved Automatic Target Recognition in Synthetic Aperture Sonar Imagery Using Large Deep Neural Networks

  • 比较现代CNN与Transformer模型在声呐图像中的表现
  • 大模型配合预训练和数据增强显著提升识别精度
  • 为声呐目标识别提供可复现的先进训练路线

合成孔径声呐(SAS)中的自动目标识别(ATR)主要依赖深度神经网络(DNN)。尽管卷积神经网络(CNN)是主流架构,但基于Transformer的模型在通用计算机视觉中已达到前沿水平,却在该领域应用较少。此外,研究人员在应对标注数据稀缺问题时,尝试使用数据增强和跨模态预训练权重,但效果参差不齐。本文比较了现代CNN与Transformer类DNN在SAS-ATR中的性能,探究网络规模、架构、预训练方法、数据增强及其他正则化手段对识别性能的影响,旨在构建最高性能模型,并为训练顶尖SAS-ATR系统提供可操作的指导路径。

原文摘要 · Abstract (English)

Automatic Target Recognition (ATR) in Synthetic Aperture Sonar (SAS) is a task largely dominated by deep neural networks (DNNs). Most SAS-ATR models use convolutional neural network (CNN) architectures whereas transformer-based architectures have had much less representation in the literature despite being state of the art in general computer vision (CV) research. Additionally, researchers have had mixed results in attempting to overcome challenges presented by a scarcity of labeled training data by using methods such as data augmentation and the use of pretrained weights from a variety of imaging modalities. In this work, we compare the performance of modern CNN and transformer-based DNNs to determine which architecture and training configurations elicit the highest performance in SAS-ATR. We investigate how network size, architecture, pretraining method, data augmentation and other forms of regularization affect SAS-ATR performance with a focus on producing the highest-performing model and providing a roadmap for training state-of-the-art SAS-ATR models.

目标识别声呐成像深度学习Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。