arXiv:2609.07753cs.CVcs.LG2026-09

用光学模型的类别特征指导雷达图像分类,解决标注数据少难题。

Cross-modal learning for SAR target recognition using optical vision foundation models

论文配图:Cross-modal learning for SAR target recognition using optical vision foundation models
图 1 · 摘自论文原文
  • 用冻结的光学视觉模型构建类别原型,跨模态对齐雷达图像特征。
  • 在UNICORNv2数据集上提升分类准确率,改善类别间区分度。
  • 无需配对图像,适合标注稀缺的遥感目标识别场景。

合成孔径雷达(SAR)因具备远距离、全天候成像能力,在众多成像应用中至关重要。然而,受限于标注数据少、强斑点噪声以及与光学图像显著的域差异,自动目标识别仍具挑战。相比之下,光电(EO)图像拥有海量数据、清晰视觉结构和强大的预训练基础模型。本文研究如何利用在光学数据上训练的视觉基础模型为SAR分类提供类别级监督。提出一种无配对的跨模态EO到SAR原型对齐框架:使用冻结的DINOv3视觉基础模型构建类别级光学原型,无需严格对应的EO/SAR图像对;随后训练SAR模型,使其嵌入向量与对应光学原型对齐。推理时SAR模型独立运行,不依赖光学图像。在包含严重斑点和类别不平衡的民用车辆数据集UNICORNv2上评估,该方法显著优于冻结DINOv3、仅微调SAR模型及无配对分布对齐基线。t-SNE可视化显示训练后SAR嵌入空间中各类别分离更清晰。结果表明,尽管光学基础模型训练于可见光谱,仍可为SAR图像分类提供可迁移信息,为跨模态大模型应用提供实用路径。

原文摘要 · Abstract (English)

Synthetic Aperture Radar (SAR) is an important modality in a wide range of imaging applications due to its versatile, long range and near all weather operating capabilities. However, Automatic Target Recognition (ATR) remains a challenging problem due to limited labelled data, the strong speckle in SAR images and the significant domain gap between SAR and more abundant optical imagery. In contrast, electro-optical (EO) imagery benefits from massive datasets, clearer visual structure and powerful foundation models. In this work, we investigate how vision foundation models trained on optical data can provide class level supervision for SAR classification. We propose a cross-modal EO to SAR prototype alignment framework in which a frozen EO encoder, based on a DINOv3 vision foundation model, is used to construct class level optical prototypes without requiring strict EO/SAR pairs. A SAR model is then trained to classify SAR images while aligning its embeddings to the corresponding EO class prototype. At inference time, the SAR model operates independently, without access to optical imagery. We evaluate our approach on the UNICORNv2 dataset, an EO and SAR dataset of civilian vehicles with heavily speckled images and severe class imbalance. EO prototype alignment improves SAR classification accuracy over frozen DINOv3, SAR only finetuning and unpaired distribution alignment baselines, and t-SNE visualizations provide qualitative evidence of clearer separation among classes in the trained SAR embedding space. These results suggest that optical vision foundation models, despite being trained on visible spectrum imagery, provide transferable information for SAR image classification, offering a practical method for using large scale pretrained vision foundation models across challenging sensing modalities.

跨模态SAR识别视觉基础模型原型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。