arXiv:2504.11999cs.CV2025-04被引 12

用物理可解释的复数模型提升雷达图像理解能力

A Complex-valued SAR Foundation Model Based on Physically Inspired Representation Learning

  • 通过模拟极化分解过程,将像素散射建模为基函数加权组合
  • 在6个下游任务中达到顶尖性能,数据稀缺时仍具强泛化性
  • 适合遥感、地球观测等领域研究者,尤其关注可解释性的应用

遥感领域的视觉基础模型因在多种下游任务中表现出优异的泛化能力而受到广泛关注。合成孔径雷达(SAR)具备全天候、全天时成像能力,在地球观测中具有显著优势。然而,构建用于SAR图像解析的基础模型面临信息利用不足和可解释性差的挑战。本文提出一种基于复数SAR数据的遥感基础模型,通过模拟极化分解过程进行预训练:将像素散射强度表征为散射基函数与散射系数的加权组合,赋予模型物理可解释性。具体而言,构建一系列代表独立且有意义散射基的查询,与SAR特征在散射查询解码器中交互,输出对应散射系数。为指导预训练,设计极化分解损失与功率自监督损失:前者使预测系数与Yamaguchi系数对齐,后者从预测系数重构功率并与其输入图像功率对比。该基础模型在六个典型下游任务上验证性能,达到当前最优结果。值得注意的是,模型能提取稳定特征表示,在数据稀缺条件下仍展现强泛化能力。

原文摘要 · Abstract (English)

Vision foundation models in remote sensing have been extensively studied due to their superior generalization on various downstream tasks. Synthetic Aperture Radar (SAR) offers all-day, all-weather imaging capabilities, providing significant advantages for Earth observation. However, establishing a foundation model for SAR image interpretation inevitably encounters the challenges of insufficient information utilization and poor interpretability. In this paper, we propose a remote sensing foundation model based on complex-valued SAR data, which simulates the polarimetric decomposition process for pre-training, i.e., characterizing pixel scattering intensity as a weighted combination of scattering bases and scattering coefficients, thereby endowing the foundation model with physical interpretability. Specifically, we construct a series of scattering queries, each representing an independent and meaningful scattering basis, which interact with SAR features in the scattering query decoder and output the corresponding scattering coefficient. To guide the pre-training process, polarimetric decomposition loss and power self-supervision loss are constructed. The former aligns the predicted coefficients with Yamaguchi coefficients, while the latter reconstructs power from the predicted coefficients and compares it to the input image's power. The performance of our foundation model is validated on six typical downstream tasks, achieving state-of-the-art results. Notably, the foundation model can extract stable feature representations and exhibits strong generalization, even in data-scarce conditions.

遥感SAR基础模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。