arXiv:2509.21722cs.CVeess.IV2025-09被引 4

用自监督学习提升雷达图像识别模型,显著超越现有最佳方案。

On the Status of Foundation Models for SAR Imagery

  • 在公开SAR数据上微调预训练模型,实现高效特征提取。
  • 新模型在低标注数据下表现优异,超越SARATR-X约12%准确率。
  • 为雷达图像领域构建可迁移的通用模型提供可行路径。

本文研究基础人工智能/机器学习模型在合成孔径雷达(SAR)目标识别任务中的可行性。受自然图像领域巨大进展启发,特别是基于网络规模数据和超大规模算力训练的自监督学习(SSL)模型,这些模型在下游任务中仅需少量标注数据即可适配,对分布偏移更具鲁棒性,且特征具有高度可迁移性。我们测试了当前最先进的视觉基础模型(如DINOv2、DINOv3、PE-Core),发现其直接应用于SAR时难以提取语义上有区分性的目标特征。随后,我们通过在SAR数据上自监督微调,训练出多个AFRL-DINOv2模型,达到新的SAR基础模型性能标杆,显著优于当前最佳模型SARATR-X。实验还分析了不同骨干网络与下游任务适配策略的性能权衡,并评估模型在扩展工作条件和低标注数据环境下的表现。尽管取得积极成果,但仍需长期探索。

原文摘要 · Abstract (English)

In this work we investigate the viability of foundational AI/ML models for Synthetic Aperture Radar (SAR) object recognition tasks. We are inspired by the tremendous progress being made in the wider community, particularly in the natural image domain where frontier labs are training huge models on web-scale datasets with unprecedented computing budgets. It has become clear that these models, often trained with Self-Supervised Learning (SSL), will transform how we develop AI/ML solutions for object recognition tasks - they can be adapted downstream with very limited labeled data, they are more robust to many forms of distribution shift, and their features are highly transferable out-of-the-box. For these reasons and more, we are motivated to apply this technology to the SAR domain. In our experiments we first run tests with today's most powerful visual foundational models, including DINOv2, DINOv3 and PE-Core and observe their shortcomings at extracting semantically-interesting discriminative SAR target features when used off-the-shelf. We then show that Self-Supervised finetuning of publicly available SSL models with SAR data is a viable path forward by training several AFRL-DINOv2s and setting a new state-of-the-art for SAR foundation models, significantly outperforming today's best SAR-domain model SARATR-X. Our experiments further analyze the performance trade-off of using different backbones with different downstream task-adaptation recipes, and we monitor each model's ability to overcome challenges within the downstream environments (e.g., extended operating conditions and low amounts of labeled data). We hope this work will inform and inspire future SAR foundation model builders, because despite our positive results, we still have a long way to go.

SAR图像自监督学习基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。