arXiv:2607.12177cs.AIcs.CV2026-07

地理基础模型让专家用AI快速分析卫星图像,无需自己训练

The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning

论文配图:The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning
图 1 · 摘自论文原文
  • 用大规模地理数据预训练模型,专家可快速微调或提示使用
  • 视觉语言模型支持零样本识别,实现开放词汇图像分析
  • 未来可构建智能代理自动解析自然语言指令并执行分析

卫星与航拍影像分析迎来新阶段,得益于地理基础模型(GeoFMs)的出现。这类AI/ML模型通过多种方法在海量地理空间数据上进行预训练。本文阐述了其核心范式转变:将大规模模型训练交由专业机构完成,领域专家则专注于任务特定的微调或提示,从而实现先进AI的快速部署,同时保障下游任务的数据安全与保密性。文中区分了基于自监督方法(如掩码自编码)的可微调视觉模型,以及基于对比学习生成的视觉-语言模型,后者支持零样本任务,如开放词汇图像分析。接着探讨了实际应用中的性能-成本权衡及更广泛的MLOps生态系统,提出模型适应策略分类体系,并为领域专家提供选择最经济高效适配方案的框架。最后展望了智能体式地理推理的未来,即大型语言模型作为智能协调者,调用GeoFMs工具,以自然语言响应用户高阶查询,自动化复杂分析流程,推动领域从感知迈向认知。

原文摘要 · Abstract (English)

The analysis of satellite and aerial imagery has entered a new era with the advent of foundation models. This paper describes the concept of Geospatial Foundation Models (GeoFMs), which are artificial intelligence/machine learning (AI/ML) models pre-trained on massive geospatial datasets through varied methodologies. We first articulate the core paradigm shift that GeoFMs enable: a separation of duties, where large-scale model providers perform the computationally intensive pretraining, allowing domain experts to rapidly fine-tune or prompt these models for specific, mission-critical tasks. This approach democratizes access to state-of-the-art AI/ML while maintaining the security and confidentiality of the downstream task. We then explore the novel capabilities unlocked by different types of GeoFMs, distinguishing between the finetunable vision models produced by self-supervised techniques like masked auto-encoding, and the vision-language models produced by contrastive learning which enable zero-shot tasks like open-vocabulary image analysis. Next, we discuss the practical considerations for operationalizing GeoFMs, from performance-cost analysis to the broader MLOps ecosystem. To that end, we introduce a taxonomy of model adaptation strategies and propose a framework for domain experts to select the most cost-effective adaptation approach for their particular mission set. Finally, we present a forward-looking vision of Agentic Geospatial Reasoning, where Large Language Models act as intelligent orchestrators, leveraging GeoFMs as tools to answer high-level user queries in natural language and automate complex analytical workflows, moving the field from perception to cognition.

地理模型基础模型遥感智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。