评测并提升遥感基础模型Prithvi的跨域适应能力,为未来模型设计提供关键洞见。
Geospatial foundation models for image analysis: evaluating and enhancing NASA-IBM Prithvi's domain adaptability
- 提出波段适配与多尺度特征生成策略,增强模型跨域性能。
- 在多个基准数据集上验证,显著提升模型在复杂场景下的预测准确率。
- 适合遥感图像分析、地理信息科学领域研究者参考使用。
地理空间基础模型(GFMs)在地理空间人工智能研究中成为热点,因其具备高泛化性和跨域适应能力,可降低研究人员的模型训练成本。与大型语言模型不同,构建用于图像分析的视觉基础模型,尤其是在遥感领域,面临将多样视觉任务统一到通用问题框架中的挑战。本文评估了近期发布的NASA-IBM GFM Prithvi在多个基准数据集上的高层图像分析任务表现。由于Prithvi是首个基于高分辨率时序遥感影像训练的开源地理空间基础模型,被选为评估对象。通过一系列实验,将其与其它预训练专用模型进行对比。本文引入波段适配、多尺度特征生成及微调技术,并集成至图像分析流程中,以增强Prithvi的域适应能力。深入分析揭示了Prithvi的优势与局限,为改进其性能以及未来地理空间视觉基础模型的研发提供了重要参考。
原文摘要 · Abstract (English)
Research on geospatial foundation models (GFMs) has become a trending topic in geospatial artificial intelligence (AI) research due to their potential for achieving high generalizability and domain adaptability, reducing model training costs for individual researchers. Unlike large language models, such as ChatGPT, constructing visual foundation models for image analysis, particularly in remote sensing, encountered significant challenges such as formulating diverse vision tasks into a general problem framework. This paper evaluates the recently released NASA-IBM GFM Prithvi for its predictive performance on high-level image analysis tasks across multiple benchmark datasets. Prithvi was selected because it is one of the first open-source GFMs trained on time-series of high-resolution remote sensing imagery. A series of experiments were designed to assess Prithvi's performance as compared to other pre-trained task-specific AI models in geospatial image analysis. New strategies, including band adaptation, multi-scale feature generation, and fine-tuning techniques, are introduced and integrated into an image analysis pipeline to enhance Prithvi's domain adaptation capability and improve model performance. In-depth analyses reveal Prithvi's strengths and weaknesses, offering insights for both improving Prithvi and developing future visual foundation models for geospatial tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。