首个百亿参数的跨域雷达影像分割模型,解决不同传感器间图像差异难题。
CrossEarth-SAR: A SAR-Centric and Billion-Scale Geospatial Foundation Model for Domain Generalizable Semantic Segmentation
- 基于物理引导的稀疏专家混合架构,融合成像原理提升泛化能力。
- 在22个子任务上超越前代方法超10%平均交并比,多域迁移表现优异。
- 适合遥感、地理信息、灾害监测等领域研究者使用。
合成孔径雷达(SAR)实现全球全天候地球观测,但因成像机制多样,传感器与区域间的域偏移严重制约其语义泛化能力。为此,我们提出CrossEarth-SAR,首个基于新型物理引导稀疏专家混合(MoE)架构构建的百亿级SAR视觉基础模型,专为跨域语义分割设计。为支持大规模预训练,我们构建了CrossEarth-SAR-200K数据集,整合公开与私有SAR影像,采用弱监督与全监督方式标注。同时,推出包含8类域差距的22个子基准组成的评估套件,建立首个统一的SAR语义分割域泛化标准。大量实验表明,CrossEarth-SAR在20个基准上达到领先水平,在部分多域迁移任务中相比以往方法提升超过10% mIoU。所有代码、基准与数据集将开源。
原文摘要 · Abstract (English)
Synthetic Aperture Radar (SAR) enables global, all-weather earth observation. However, owing to diverse imaging mechanisms, domain shifts across sensors and regions severely hinder its semantic generalization. To address this, we present CrossEarth-SAR, the first billion-scale SAR vision foundation model built upon a novel physics-guided sparse mixture-of-experts (MoE) architecture incorporating physical descriptors, explicitly designed for cross-domain semantic segmentation. To facilitate large-scale pre-training, we develop CrossEarth-SAR-200K, a weakly and fully supervised dataset that unifies public and private SAR imagery. We also introduce a benchmark suite comprising 22 sub-benchmarks across 8 distinct domain gaps, establishing the first unified standard for domain generalization semantic segmentation on SAR imagery. Extensive experiments demonstrate that CrossEarth-SAR achieves state-of-the-art results on 20 benchmarks, surpassing previous methods by over 10\% mIoU on some benchmarks under multi-gap transfer. All code, benchmark and datasets will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。