arXiv:2602.17665cs.CV2026-02中稿 · ECCV被引 10

让AI读懂卫星图并自动执行地理分析任务,支持多工具协同推理。

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents

  • 构建统一工具库与轨迹策略学习框架,实现地理空间任务的模块化执行。
  • 在14,538个训练样本上完成超10万步推理,准确率显著优于基线模型。
  • 适合遥感、城市规划、灾害监测等领域研究人员快速部署智能分析系统。

多模态推理的进展使智能体能够解析图像、关联语言并执行结构化分析任务。将此类能力扩展至遥感领域仍具挑战性,因模型需在空间尺度、地理结构及多光谱指数间进行推理,并保持连贯的多步逻辑。为此,我们提出OpenEarthAgent,一个基于卫星影像、自然语言查询和结构化推理轨迹训练的工具增强型地理空间推理统一框架。该框架不仅作为基准,更建立了一套以统一可执行工具注册表和基于轨迹的策略学习为核心的智能体架构。其将异构的视觉、光谱、GIS及地理参考栅格操作标准化为一致调用接口,支持模块化编排与确定性执行。通过在结构化推理轨迹上进行监督微调,并采用确定性回放验证确保可执行性与空间正确性。配套语料库包含14,538个训练实例和1,169个评估实例,涵盖超过107,000个推理步骤,覆盖城市、环境、灾害与基础设施场景,整合了GIS操作与NDVI、NBR、NDBI等指数分析。基于显式推理轨迹,所学智能体展现出结构化推理能力、稳定的空间理解与可解释的工具驱动行为,在多种地球观测场景中表现优异,持续优于强基线模型,并达到近期开源与闭源模型的竞争力水平。代码、数据与训练模型均已公开:https://github.com/mbzuai-oryx/OpenEarthAgent

原文摘要 · Abstract (English)

Recent progress in multimodal reasoning has enabled agents that interpret imagery, connect it with language, and execute structured analytical tasks. Extending these capabilities to remote sensing remains challenging, as models must reason over spatial scale, geographic structures, and multispectral indices while maintaining coherent multi-step logic. To address this gap, we introduce \textit{OpenEarthAgent}, a unified framework for tool-augmented geospatial reasoning trained on satellite imagery, natural-language queries, and structured reasoning traces. Beyond serving as a benchmark, OpenEarthAgent establishes a cohesive agentic architecture built around a unified executable tool registry and trajectory-based policy learning. The framework standardizes heterogeneous visual, spectral, GIS, and georeferenced raster operations under a consistent callable schema, enabling modular orchestration and deterministic execution. Training is performed via supervised fine-tuning on structured reasoning trajectories with deterministic replay validation to ensure executability and spatial correctness. The accompanying corpus comprises 14,538 training and 1,169 evaluation instances with over 107K reasoning steps, spanning urban, environmental, disaster, and infrastructure domains and incorporating GIS operations alongside index analyses such as NDVI, NBR, and NDBI. Grounded in explicit reasoning traces, the learned agent demonstrates structured reasoning, stable spatial understanding, and interpretable tool-driven behaviour across diverse EO scenarios. We report consistent improvements over a strong baseline and competitive performance against recent open and closed-source models. Our code, data and trained models are publicly available: https://github.com/mbzuai-oryx/OpenEarthAgent

遥感智能工具增强地理推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。