构建地理遥感多模态评测基准与智能代理,提升专家级解析能力
GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing
- 设计跨领域、多传感器的问答评测集GeoMMBench
- 发现36个模型在地学知识与推理上存在系统性短板
- 提出工具增强型多智能体框架,显著超越单一大模型
多模态大语言模型虽推动了垂直领域AI发展,但在地质科学与遥感(RS)领域仍受限于学科知识广度、传感器模态异构及任务碎片化。为此,我们提出GeoMMBench,一个覆盖多元遥感学科、传感器与任务的综合性多模态问答评测基准,支持更广泛且严谨的评估。基于该基准,我们评估了36个开源与专有大语言模型,揭示其在领域知识、感知对齐与推理能力上的系统性不足,而这些正是实现专家级地理空间解析的关键。为突破瓶颈,我们进一步提出GeoMMAgent,一种通过领域特定遥感模型与工具协同检索、感知与推理的多智能体框架。实验表明,该框架显著优于独立大语言模型,凸显工具增强型智能体在动态应对复杂地学与遥感挑战中的关键作用。
原文摘要 · Abstract (English)
Recent advances in multimodal large language models (MLLMs) have accelerated progress in domain-oriented AI, yet their development in geoscience and remote sensing (RS) remains constrained by distinctive challenges: wide-ranging disciplinary knowledge, heterogeneous sensor modalities, and a fragmented spectrum of tasks. To bridge these gaps, we introduce GeoMMBench, a comprehensive multimodal question-answering benchmark covering diverse RS disciplines, sensors, and tasks, enabling broader and more rigorous evaluation than prior benchmarks. Using GeoMMBench, we assess 36 open-source and proprietary large language models, uncovering systematic deficiencies in domain knowledge, perceptual grounding, and reasoning--capabilities essential for expert-level geospatial interpretation. Beyond evaluation, we propose GeoMMAgent, a multi-agent framework that strategically integrates retrieval, perception, and reasoning through domain-specific RS models and tools. Extensive experimental results demonstrate that GeoMMAgent significantly outperforms standalone LLMs, underscoring the importance of tool-augmented agents for dynamically tackling complex geoscience and RS challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。