arXiv:2410.22629cs.CV2024-10TPAMI被引 46

首个面向遥感语义分割的跨域泛化基础模型,提升多场景适应能力。

CrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic Segmentation

  • 通过地球风格注入与多任务训练,增强模型跨域泛化能力。
  • 在32种跨域设置下显著优于现有方法,最高提升12.7%精度。
  • 适合遥感、地理信息、环境监测等需跨区域部署的场景使用。

遥感领域泛化(RSDG)已成为关键研究方向,旨在开发能在多种场景下有效泛化的模型。尽管遥感图像存在位置、波段和传感器类型等多样性的显著域间差异,该领域研究仍较薄弱:(1) 现有跨域方法多聚焦于领域自适应(DA),仅针对已知域进行适配,难以应对未见域;(2) 针对语义分割任务的RSDG研究极少,现有模型针对特定未知域设计,对其他未知场景易出现欠拟合;(3) 现有遥感基础模型更关注域内性能,忽视跨域泛化。为此,我们提出首个面向RSDG语义分割的视觉基础模型CrossEarth。CrossEarth通过定制的数据级地球风格注入流程与模型级多任务训练策略,实现强跨域泛化。此外,我们构建了包含32个跨域设置的RSDG基准,覆盖不同地区、光谱波段、平台与气候条件,为未来RSDG模型提供全面评估框架。大量实验表明,CrossEarth在该基准上显著优于现有最先进方法。

原文摘要 · Abstract (English)

The field of Remote Sensing Domain Generalization (RSDG) has emerged as a critical and valuable research frontier, focusing on developing models that generalize effectively across diverse scenarios. Despite the substantial domain gaps in RS images that are characterized by variabilities such as location, wavelength, and sensor type, research in this area remains underexplored: (1) Current cross-domain methods primarily focus on Domain Adaptation (DA), which adapts models to predefined domains rather than to unseen ones; (2) Few studies targeting the RSDG issue, especially for semantic segmentation tasks, where existing models are developed for specific unknown domains, struggling with issues of underfitting on other unknown scenarios; (3) Existing RS foundation models tend to prioritize in-domain performance over cross-domain generalization. To this end, we introduce the first vision foundation model for RSDG semantic segmentation, CrossEarth. CrossEarth demonstrates strong cross-domain generalization through a specially designed data-level Earth-Style Injection pipeline and a model-level Multi-Task Training pipeline. In addition, for the semantic segmentation task, we have curated an RSDG benchmark comprising 32 cross-domain settings across various regions, spectral bands, platforms, and climates, providing a comprehensive framework for testing the generalizability of future RSDG models. Extensive experiments on this benchmark demonstrate the superiority of CrossEarth over existing state-of-the-art methods.

遥感语义分割域泛化基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。