arXiv:2607.23238cs.CV2026-07

提出新型SAR预训练目标,兼顾物理稳定与尺度兼容性

SARATR-X-v2: Scale-Aware Structural Pre-Training for SAR Foundation Models

论文配图:SARATR-X-v2: Scale-Aware Structural Pre-Training for SAR Foundation Models
图 1 · 摘自论文原文
  • 用六种固定结构提取器融合生成统一监督信号
  • 在12个SAR任务上达最优迁移性能,抗斑点扰动提升近两个数量级
  • 适合需要高鲁棒性的遥感图像分析研究者

掩码图像建模已成为合成孔径雷达(SAR)预训练的主流范式,但重建目标的设计仍不明确。本文认为,一个有效的SAR预训练目标需满足两个条件:(i) 物理根基的稳定性,即目标算子对相干成像固有的乘性斑点具有近似不变性;(ii) 语义尺度兼容性,即覆盖下游任务所需的异构空间尺度。这两个条件分别可实现,但联合难以兼顾:物理稳定性偏好固定算子,而语义尺度兼容性偏好数据驱动组合。为此,SARATR-X-v2在单一设计中协调二者。目标通过六个感受野跨度的固定结构提取器构建,涵盖盲区局部聚合到方向对数比区域对比,并通过可学习权重融合为统一的重建监督信号。在十二个SAR基准任务(分类、检测、分割)上,SARATR-X-v2均取得当前最优迁移性能。在合成斑点变化下,所提目标使学习表征的扰动漂移相比像素空间监督降低近两个数量级。结果共同确立了物理根基稳定性和语义尺度兼容性作为相干成像预训练目标设计的原理框架,表明有效SAR预训练不在于重建更多信号,而在于重建正确的结构目标。

原文摘要 · Abstract (English)

Masked image modeling has become a dominant paradigm for SAR pre-training, yet the design of the reconstruction target remains fundamentally unsettled. This article argues that a SAR pre-training target should satisfy two conditions to produce transferable representations: (i) physics-grounded stability, i.e., approximate invariance of the target operator to multiplicative speckle inherent in coherent imaging; and (ii) semantic scale compatibility, i.e., coverage of the heterogeneous spatial scales that downstream tasks demand. These two conditions are individually achievable but jointly difficult: physics-grounded stability favors fixed operators, while semantic scale compatibility favors data-driven composition. To this end, SARATR-X-v2 reconciles both within a single design. The target is constructed through fixed structural extractors spanning six receptive fields, from blind-spot local aggregation to directional log-ratio region contrast, and fused via learnable weights into one unified supervision signal for masked reconstruction. On twelve SAR benchmarks across classification, detection, and segmentation, SARATR-X-v2 achieves state-of-the-art transfer performance. Under synthetic speckle variation, the proposed target reduces perturbation drift in the learned representation by nearly two orders of magnitude relative to pixel-space supervision. Taken together, these results establish physics-grounded stability and semantic scale compatibility as a principled framework for pre-training target design under coherent imaging, and suggest that effective SAR pre-training is not about reconstructing more signal, but about reconstructing the right structural target.

SAR预训练结构建模斑点抑制遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。