用小模型组合提升遥感数据挖掘效果,更省资源。
Towards Scalable and Generalizable Earth Observation Data Mining via Foundation Model Composition
- 将多个预训练小模型在特征层融合,提升性能。
- 组合方案在11个数据集上媲美甚至超过大模型。
- 适合资源有限但需高效遥感分析的团队使用。
基础模型正快速改变遥感数据挖掘,为场景分类与语义分割等任务提供通用且可扩展的解决方案。尽管当前多数研究聚焦于从头训练大规模遥感模型,但一个未被充分探索的策略是复用并组合现有预训练模型。本研究探讨了在遥感与通用视觉数据上预训练的基础模型能否通过组合方式提升多种关键遥感任务的表现。基于GEO-Bench基准,我们在涵盖不同空间分辨率、传感器模态和任务类型的11个数据集上评估了Prithvi、Hiera和DOFA等主流模型。结果表明,小模型在特征层面的集成可达到甚至超越大型模型的性能,同时显著减少训练时间和计算资源消耗。此外,研究还展示了知识蒸馏技术可将集成模型的优势迁移至更紧凑的模型中,为实际部署基础模型提供了可行路径。
原文摘要 · Abstract (English)
Foundation models are rapidly transforming Earth Observation data mining by enabling generalizable and scalable solutions for key tasks such as scene classification and semantic segmentation. While most efforts in the geospatial domain have focused on developing large models trained from scratch using massive Earth Observation datasets, an alternative strategy that remains underexplored is the reuse and combination of existing pretrained models. In this study, we investigate whether foundation models pretrained on remote sensing and general vision datasets can be effectively combined to improve performance across a diverse set of key Earth Observation tasks. Using the GEO-Bench benchmark, we evaluate several prominent models, including Prithvi, Hiera, and DOFA, on eleven datasets covering a range of spatial resolutions, sensor modalities, and task types. The results show that feature-level ensembling of smaller pretrained models can match or exceed the performance of much larger models, while requiring less training time and computational resources. Moreover, the study highlights the potential of applying knowledge distillation to transfer the strengths of ensembles into more compact models, offering a practical path for deploying foundation models in real-world Earth Observation applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。