SARFormer通过参数编码提升雷达图像识别效果,尤其适合多图处理。
SARFormer -- An Acquisition Parameter Aware Vision Transformer for Synthetic Aperture Radar Data
- 引入成像参数编码模块,引导ViT学习复杂雷达几何特征
- 在有限标注数据下,高度重建误差降低17%(RMSE)
- 适合遥感、地质测绘等需要多幅雷达图分析的场景
本文提出SARFormer,一种针对单幅或多幅合成孔径雷达(SAR)图像设计的改进视觉变换器(ViT)架构。由于SAR数据具有复杂的图像几何特性,我们设计了成像参数编码模块,显著指导学习过程,尤其在多图情况下表现更优。研究进一步探索自监督预训练,在标注数据有限条件下开展实验,并通过消融实验与基线模型对比,评估其在高程重建和分割等下游任务上的性能。结果表明,该方法在RMSE指标上相较基线模型最高提升17%。
原文摘要 · Abstract (English)
This manuscript introduces SARFormer, a modified Vision Transformer (ViT) architecture designed for processing one or multiple synthetic aperture radar (SAR) images. Given the complex image geometry of SAR data, we propose an acquisition parameter encoding module that significantly guides the learning process, especially in the case of multiple images, leading to improved performance on downstream tasks. We further explore self-supervised pre-training, conduct experiments with limited labeled data, and benchmark our contribution and adaptations thoroughly in ablation experiments against a baseline, where the model is tested on tasks such as height reconstruction and segmentation. Our approach achieves up to 17% improvement in terms of RMSE over baseline models
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。