arXiv:2601.02158cs.CLcs.AI2026-01

评测大模型在石油地质领域的多选题能力,发现开源模型表现接近闭源。

FormationEval, an open multiple-choice benchmark for petroleum geoscience

  • 构建7个领域505道题的开放基准,避免抄袭并保留来源信息。
  • 顶级模型准确率超97%,开源模型如GLM-4.7达98.6%。
  • 适合评估地质、能源等领域模型,尤其关注开源方案表现。

本文提出FormationEval,一个面向石油地质与地下学科的语言模型多选题评测基准。数据集涵盖505道题,来自三个权威来源,覆盖测井学、石油地质学和储层工程等七个领域,采用基于概念的构建方法并借助推理模型生成,避免直接复制受版权保护内容。每道题附带来源元数据以确保可追溯性。评估涵盖72个主流模型,包括OpenAI、Anthropic、Google、Meta及开源模型。最优模型Gemini 3 Pro Preview准确率达99.8%,其他模型普遍超过97%;开源模型中GLM-4.7表现最佳(98.6%),多个DeepSeek、Llama、Qwen和Mistral模型准确率超93%。尽管存在领域差异与层级差距,但开源与闭源模型间的性能差距小于预期,部分低成本开源模型准确率超90%。测井学为所有模型中最难的领域,小模型表现波动较大。研究记录了数据集中正确答案偏长的残差长度偏差,并采取相应缓解策略。该基准、评估代码与结果已公开。

原文摘要 · Abstract (English)

This paper presents FormationEval, an open multiple-choice question benchmark for evaluating language models on petroleum geoscience and subsurface disciplines. The dataset contains 505 questions across seven domains including petrophysics, petroleum geology and reservoir engineering, derived from three authoritative sources using a reasoning model with detailed instructions and a concept-based approach that avoids verbatim copying of copyrighted text. Each question includes source metadata to support traceability and audit. The evaluation covers 72 models from major providers including OpenAI, Anthropic, Google, Meta and open-weight alternatives. The top performers achieve over 97% accuracy, with Gemini 3 Pro Preview reaching 99.8%, while tier and domain gaps persist. Among open-weight models, GLM-4.7 leads at 98.6%, with several DeepSeek, Llama, Qwen and Mistral models also exceeding 93%. The performance gap between open-weight and closed models is narrower than expected, with several lower-cost open-weight models exceeding 90% accuracy. Petrophysics emerges as the most challenging domain across all models, while smaller models show wider performance variance. Residual length bias in the dataset (correct answers tend to be longer) is documented along with bias mitigation strategies applied during construction. The benchmark, evaluation code and results are publicly available.

语言模型石油地质评测基准开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。