arXiv:2412.12565cs.CV2024-12被引 2

针对极端长尾分布的SAR图像分类,提出自监督与跨模态翻译联合策略。

PBVS 2024 Solution: Self-Supervised Learning and Sampling Strategies for SAR Classification in Extreme Long-Tail Distribution

  • 分两阶段训练:先自监督预训练,再通过SAR转EO提升多模态学习效果。
  • 在1000倍长尾分布下取得21.45%准确率,优于多数传统长尾方法。
  • 适合处理遥感图像中数据极度不均衡且模态差异大的场景。

多模态学习工作坊(PBVS 2024)旨在通过结合难以解读但不受天气和光照影响的合成孔径雷达(SAR)数据与光电(EO)数据,提升自动目标识别(ATR)系统性能。本任务聚焦于基于一组SAR-EO图像对及其标签,预测低分辨率航拍图像的类别。数据集呈现极端长尾分布,最大类与最小类样本量相差超1000倍,使常规长尾处理方法失效;同时SAR与EO数据域差异显著,影响标准多模态方法效果。为此,我们提出两阶段学习框架:利用自监督技术,并结合多模态学习与SAR到EO的图像转换,以更有效利用EO信息。在PBVS 2024多模态航拍图像分类任务(SAR分类)最终测试中,模型获得21.45%准确率、0.56 AUC和0.30总分,排名第九。

原文摘要 · Abstract (English)

The Multimodal Learning Workshop (PBVS 2024) aims to improve the performance of automatic target recognition (ATR) systems by leveraging both Synthetic Aperture Radar (SAR) data, which is difficult to interpret but remains unaffected by weather conditions and visible light, and Electro-Optical (EO) data for simultaneous learning. The subtask, known as the Multi-modal Aerial View Imagery Challenge - Classification, focuses on predicting the class label of a low-resolution aerial image based on a set of SAR-EO image pairs and their respective class labels. The provided dataset consists of SAR-EO pairs, characterized by a severe long-tail distribution with over a 1000-fold difference between the largest and smallest classes, making typical long-tail methods difficult to apply. Additionally, the domain disparity between the SAR and EO datasets complicates the effectiveness of standard multimodal methods. To address these significant challenges, we propose a two-stage learning approach that utilizes self-supervised techniques, combined with multimodal learning and inference through SAR-to-EO translation for effective EO utilization. In the final testing phase of the PBVS 2024 Multi-modal Aerial View Image Challenge - Classification (SAR Classification) task, our model achieved an accuracy of 21.45%, an AUC of 0.56, and a total score of 0.30, placing us 9th in the competition.

SAR分类长尾分布自监督学习多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。