arXiv:2605.02283cs.CVcs.AI2026-05被引 1

通用视觉模型在遥感检索中表现不逊于专用模型,且更稳定。

Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM

论文配图:Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM
图 1 · 摘自论文原文
  • 对比通用与遥感专用视觉模型的检索性能,控制数据集和评估流程。
  • 通用模型在多数场景下表现相当或更优,专用模型跨场景退化严重。
  • 适合关注遥感图像检索模型设计的研究者与开发者参考。

视觉基础模型因其能利用大规模无标注视觉数据而备受关注,这一优势在遥感领域尤为重要——数据获取成本高,标注常需专业知识。近期的电光视觉基础模型旨在从遥感图像中学习领域特定表征,但其在基于检索的评估中是否优于强通用视觉基础模型仍不明确。本研究对代表性遥感专用与通用视觉基础模型进行了受控对比,使用相同数据集、检索协议与评估指标,评估其域内性能与跨场景泛化能力。结果表明,强通用视觉基础模型在多数情况下表现不逊于甚至超越现有遥感专用模型;而遥感专用模型在跨场景评估中普遍出现显著性能下降,通用模型则展现出更稳定的迁移能力。这些发现表明,仅进行遥感预训练并不能保证更强的检索导向遥感表征。我们讨论了当前遥感专用预训练策略的局限性,并强调未来遥感视觉基础模型应更好地利用遥感图像的物理、空间、光谱与地理特征。

原文摘要 · Abstract (English)

Vision foundation models have attracted significant attention for their ability to leverage large-scale unlabeled visual data. This advantage is particularly important in remote sensing, where data acquisition is costly and annotation often requires expert knowledge. Recent electro-optical vision foundation models aim to learn domain-specific representations from remote sensing imagery, but it remains unclear whether they are more effective than strong generalist vision foundation models under retrieval-based evaluation. In this study, we conduct a controlled comparison between representative EO-specific and generalist vision foundation models for remote sensing image retrieval. Using the same datasets, retrieval protocol, and evaluation metric, we evaluate both in-domain performance and cross-scene generalization. Our results show that strong generalist vision foundation models are competitive with, and in some cases outperform, existing EO-specific models. Moreover, EO-specific models often suffer from substantial degradation under cross-scene evaluation, while generalist models show more stable transfer. These findings suggest that EO pretraining alone does not guarantee stronger retrieval-oriented remote sensing representations. We discuss the limitations of current EO-specific pretraining strategies and highlight the need for future EO vision foundation models to better exploit the physical, spatial, spectral, and geographic characteristics of remote sensing imagery.

遥感检索视觉模型跨场景泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。