构建首个地球科学视觉推理评测基准,揭示当前AI在真实场景中推理能力严重不足。
GeoR-Bench: Evaluating Geoscience Visual Reasoning

- 设计视觉编辑任务评估模型对地球科学图像的推理能力
- 21个模型最高仅42.7%严格准确率,开源模型仅10.3%
- 输出图像质量高但科学准确性差,暴露浅层拟真问题
地球科学智能需理解、推理并预测地球系统变化,以支持灾害应对、气候适应与环境保护等关键决策。尽管现有研究在遥感解译、地理问答等特定任务上取得进展,但现有基准仍高度任务导向,难以反映真实世界开放性地球科学问题。为此,我们提出GeoR-Bench——一个通过推理驱动的视觉编辑任务评估地球科学视觉推理能力的基准。该基准包含440个精选样本,覆盖6类地球科学主题和24种任务类型,涵盖地球观测影像及地图、图表等结构化科学表征。我们在推理、一致性与质量三个维度评估模型输出。对21个闭源与开源多模态模型的测评显示,地球科学推理仍是核心瓶颈:最高性能模型整体严格准确率为42.7%,最佳开源模型仅为10.3%。值得注意的是,模型输出的视觉一致性和图像质量常远超其科学准确性。结果表明,当前模型生成的是表面看似合理但未捕捉地球科学本质过程的伪结论。
原文摘要 · Abstract (English)
Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains such as disaster response, climate adaptation and environmental protection. Although current research has shown promising progress on specific geoscience tasks, such as remote sensing interpretation, geographic question-answering, existing benchmarks remain largely task-specific which failing to capture the open-ended real world geoscience problems. As a result, it remains unclear how far current AI systems are from achieving genuine geoscience intelligence. To address this gap, we present \textbf{GeoR-Bench}, a \underline{Bench}mark for evaluating \underline{Geo}science visual \underline{R}easoning through reasoning informed visual editing tasks. GeoR-Bench contains 440 curated samples spanning 6 geoscience categories and 24 task types, covering earth observation imagery and structured scientific representations such as maps and diagrams. We evaluate outputs along three dimensions, including reasoning, consistency, and quality. Benchmark results of 21 closed- and open-source multimodal models reveal that geoscience reasoning remains a critical bottleneck. The highest-performing model achieves 42.7\% overall strict accuracy, while the best open-source models only get 10.3\%. Notably, the visual consistency and image quality of the outputs frequently surpass their scientific accuracy. Ultimately, these findings indicate that current models generate superficially plausible results but fail to capture underlying earth science processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。