arXiv:2505.23522cs.CVcs.LG2025-05被引 13

首个覆盖地球六圈层的多模态评测基准,揭示大模型在地球系统认知上的显著短板。

OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data

  • 构建六圈层全覆盖的多模态评估体系,支持跨圈层交互分析。
  • 包含109项任务、2.9万条专家标注数据,测试显示顶尖模型准确率不足35%。
  • 适合地球科学与多模态大模型交叉研究者使用,推动系统认知能力提升。

现有地球科学多模态学习基准对地球圈层及其相互作用的覆盖有限,通常仅聚焦大气圈中的人类活动领域,且最多涵盖16项任务。其局限在于数据源单一、科学粒度受限、圈层扩展性差。为此,我们提出OmniEarth-Bench,首个系统覆盖大气圈、岩石圈、水圈、冰冻圈、生物圈及人类活动圈,并涵盖跨圈层交互的多模态评测基准。基于可扩展的模块化数据推理框架与原生多观测源、专家参与的标注流程,生成29,855条标准化、专家校验的标注数据。所有标注按四层层级(圈层、场景、能力、任务)组织,涵盖109项专家定义的评估任务。在9个前沿多模态大模型上进行实验发现,即使最先进的模型也难以应对本基准,无一达到35%准确率,暴露出地球系统认知能力的系统性差距。数据集与评估代码已公开于OmniEarth-Bench(https://anonymous.4open.science/r/OmniEarth-Bench-B1BD)。

原文摘要 · Abstract (English)

Existing benchmarks for multimodal learning in Earth science offer limited, siloed coverage of Earth's spheres and their cross-sphere interactions, typically restricting evaluation to the human-activity sphere of atmosphere and to at most 16 tasks. These limitations: narrow-source heterogeneity (single/few data sources), constrained scientific granularity, and limited-sphere extensibility. Therefore, we introduce OmniEarth-Bench, the first multimodal benchmark that systematically spans all six spheres: atmosphere, lithosphere, oceanosphere, cryosphere, biosphere, and human-activity sphere, and cross-spheres. Built with a scalable, modular-topology data inference framework and native multi-observation sources and expert-in-the-loop curation, OmniEarth-Bench produces 29,855 standardized, expert-curated annotations. All annotations are organized into a four-level hierarchy (Sphere, Scenario, Ability, Task), encompassing 109 expert-curated evaluation tasks. Experiments on 9 state-of-the-art MLLMs reveal that even the most advanced models struggle with our benchmarks, where none of them reach 35% accuracy, revealing systematic gaps in Earth-system cognitive ability. The dataset and evaluation code were released at OmniEarth-Bench (https://anonymous.4open.science/r/OmniEarth-Bench-B1BD).

地球科学多模态评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。