arXiv:2505.20740cs.AI2025-05ACL被引 5

构建首个地球科学多模态基准,助力大模型理解真实科研图示与推理。

MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs

  • 从高质量论文中提取28.9万张图表,融合上下文讨论生成精准描述。
  • 覆盖五大地球圈层,支持图文生成、问答与开放推理等多种任务。
  • 填补地球科学领域真实复杂推理数据的空白,适合科研级模型评估。

多模态大语言模型(MLLMs)的快速发展为复杂科学挑战带来新机遇,但其在地球科学——尤其是研究生层次——的应用仍因缺乏反映地质科学推理深度与复杂性的基准而受限。现有数据集多依赖合成数据或简单图文配对,难以捕捉真实场景中的精细推理需求。为此,我们提出MSEarth,一个从高质量开源论文中精心构建的多模态科学数据集与基准。该数据集涵盖地球科学五大领域:大气圈、冰冻圈、水圈、岩石圈和生物圈,包含超过28.9万张经上下文讨论和推理优化的图表及其精炼标注。基准支持科学图示描述、多项选择题和开放式推理等任务,提供可扩展、高保真的资源,用于开发与评估MLLM在科学推理中的表现。

原文摘要 · Abstract (English)

The rapid advancement of multimodal large language models (MLLMs) offers new opportunities for complex scientific challenges, yet their application in earth science-especially at the graduate level-remains underexplored due to a lack of benchmarks reflecting the depth and complexity of geoscientific reasoning. Existing datasets often rely on synthetic data or simple figure-caption pairs, failing to capture the nuanced reasoning required for real-world applications. To address this, we introduce MSEarth, a multimodal scientific dataset and benchmark curated from high-quality, open-access publications. Covering the five major spheres of Earth science-atmosphere, cryosphere, hydrosphere, lithosphere, and biosphere-MSEarth features over 289K figures with refined captions enriched by contextual discussions and reasoning from the original papers. The benchmark supports tasks such as scientific figure captioning, multiple choice questions, and open-ended reasoning, providing a scalable, high-fidelity resource for developing and evaluating MLLMs in scientific reasoning.

多模态地球科学大模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。