arXiv:2605.07640cs.CVcs.AI2026-05

首个面向遥感岩性解读的多层级评测基准,助力大模型理解地质语义。

LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation

论文配图:LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation
图 1 · 摘自论文原文
  • 构建包含10,000条专家标注的岩性任务,分五认知层级设计评测
  • 大视觉语言模型在高阶解释与推理任务上表现显著不足
  • 融合专家反馈与结构化描述,提升评测的地质有效性

遥感岩性解读是地质调查、矿产勘探和区域地质制图的基础。不同于通用地表覆盖识别,岩性解读依赖专业知识,需从细微的视觉、光谱、纹理、地貌及上下文线索中推断岩性,自动化解读极具挑战。现有地质知识引导的大规模多模态模型缺乏可靠评估基准,尤其缺少岩性标注、多层级地质语义和专家评价体系。为此,我们提出LithoBench,一个用于评估遥感岩性解读中地质语义理解能力的多层级基准。该基准包含12种典型岩性类别,共10,000个专家标注实例,涵盖4,000道多选题和6,000道开放题,按五个认知层级组织:识别与描述、对比分析、机理解释、实际应用与综合推理。我们还开发了专家参与的半自动构建流程,结合结构化地质图像描述等多步骤,提升地质合理性与评估可靠性。对多个大型视觉-语言模型的实验显示,其在高阶解释、应用与推理任务中存在明显局限。

原文摘要 · Abstract (English)

Remote sensing lithology interpretation is fundamental to geological surveys, mineral exploration, and regional geological mapping. Unlike general land-cover recognition, lithology interpretation is a knowledge-intensive task that requires experts to infer rock types from various features, e.g., subtle visual, spectral, textural, geomorphological, and contextual cues, making reliable automated interpretation highly challenging. Geological knowledge-guided large multimodal models offer new opportunities, yet their evaluation remains constrained by the lack of benchmarks that capture lithological annotations, multi-level geological semantics, and expert-informed assessment. Here, we propose LithoBench, a multi-level benchmark for evaluating geological semantic understanding in remote sensing lithology interpretation. LithoBench contains 10,000 expert-annotated interpretation instances across 12 representative lithological categories, including 4,000 multiple-choice and 6,000 open-ended tasks organized into five cognitive levels: Identification and Description, Comparative Analysis, Mechanism Explanation, Practical Application, and Comprehensive Reasoning. We further develop an expert-in-the-loop, knowledge-grounded semi-automated construction pipeline, coupling multi sub-processes, e.g., structured geological image descriptions, to enhance geological validity and evaluation reliability. Experiments with multiple large vision-language models eveal substantial limitations in geological semantic understanding, particularly on higher-order explanation, application, and reasoning tasks.

遥感岩性多模态模型地质理解评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。