arXiv:2511.02831cs.LG2025-11被引 3

测试遥感模型跨卫星波段的泛化能力,发现现有模型表现不佳。

GeoCrossBench: Cross-Band Generalization for Remote Sensing

  • 构建新基准GeoCrossBench,评估模型在无波段重叠和新增波段下的表现
  • 模型在无波段重叠时性能下降2-4倍,新增波段导致平均5-25%性能下降
  • 仅微调最后一层线性层即可提升跨卫星一致性,表明模型仍有改进空间

随着遥感卫星数量与多样性增加,标注数据仍主要来自旧卫星。当地球观测基础模型规模扩大时,为支持新卫星而重新训练的成本也随之上升,因此模型对新卫星的泛化能力变得愈发关键。本文提出GeoCrossBench,扩展了流行的GeoBench基准,引入新评估协议:测试模型在分布内表现、无波段重叠卫星上的泛化能力,以及面对训练集未包含的新波段时的适应性。我们还开发了ChannelViT的自监督扩展模型ChiViT,以提升跨卫星性能。实验显示,即使最先进的遥感基础模型(DOFA、TerraFM)在分布内设置下也未能超越通用模型DINOv3;当泛化至无波段重叠的卫星时,所有模型性能下降2-4倍,而ChiViT显著优于次优模型DINOv3;在测试时引入额外波段的情况下,所有模型平均性能下降5-25%。此外,仅使用全波段真值标签微调最后一层线性层,即可获得相对稳定的跨卫星表现,表明当前基准尚未饱和。代码与数据集已公开,以推动更具未来适应性的遥感模型发展。

原文摘要 · Abstract (English)

The number and diversity of remote sensing satellites grows over time, while the vast majority of labeled data comes from older satellites. As the foundation models for Earth observation scale up, the cost of (re-)training to support new satellites grows too, so the generalization capabilities of the models towards new satellites become increasingly important. In this work we introduce GeoCrossBench, an extension of the popular GeoBench benchmark with a new evaluation protocol: it tests the in-distribution performance; generalization to satellites with no band overlap; and generalization to satellites with additional bands with respect to the training set. We also develop a self-supervised extension of ChannelViT, ChiViT, to improve its cross-satellite performance. First, we show that even the best foundation models for remote sensing (DOFA, TerraFM) do not outperform general purpose models like DINOv3 in the in-distribution setting. Second, when generalizing to new satellites with no band overlap, all models suffer 2-4x drop in performance, and ChiViT significantly outperforms the runner-up DINOv3. Third, the performance of all tested models drops on average by 5-25\% when given additional bands during test time. Finally, we show that fine-tuning just the last linear layer of these models using oracle labels from all bands can get relatively consistent performance across all satellites, highlighting that the benchmark is far from being saturated. We publicly release the code and the datasets to encourage the development of more future-proof remote sensing models with stronger cross-satellite generalization.

遥感跨卫星泛化能力自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。