为石油工程领域大模型打造评估基准,测试8个主流模型表现。
PetroBench: A Benchmark for Large Language Models in Petroleum Engineering
- 构建三阶段数据流程,生成1200道高相关性专业题
- 多选题最高准确率65.3%,判断题74.3%,整体表现有限
- 中文模型在选择题上占优,国际模型在问答中更强
大语言模型在石油行业的应用日益广泛,亟需领域专用的评估框架。本研究构建了石油工程领域的基准评测体系,包含数据预处理、质量筛选和多模型验证三阶段流程。通过专家评审,建立具有强领域相关性和区分度的标准题库,覆盖生产、储层和钻井工程,共1200道题目,题型包括单选、判断、术语定义和简答。在统一API环境下对8个主流大模型进行评估。结果显示,模型在主观题上表现优于客观题,表明其事实知识辨识能力较弱。多选题最高准确率为65.3%,判断题为74.3%。Gemini-3-Pro、Kimi-K2.5和Claude-Opus-4.6-Thinking综合得分最高,达72%-74%。模型在生产工程任务中表现最好,在储层工程中最差。中文模型在多选题中占优,国际模型在简答题中略胜一筹。该基准为石油工程领域大模型的评估与部署提供了可复现、实用的参考。
原文摘要 · Abstract (English)
Large Language Models are increasingly applied in the petroleum industry, highlighting the need for a domain-specific evaluation framework. This study develops a benchmark for LLMs in petroleum engineering, including a three-stage process of data preprocessing, quality filtering, and multi-model validation. Using expert review, a standardized question bank with strong domain relevance and discriminative capability was constructed. The benchmark covers production, reservoir, and drilling engineering, with 1,200 questions across multiple-choice, true or false, term definition, and short-answer formats. Eight mainstream LLMs were evaluated under a unified API environment. Results show that models performed better on subjective than objective questions, indicating weaknesses in factual knowledge discrimination. The highest accuracies for multiple-choice and true or false questions were 65.3% and 74.3%, respectively. Gemini-3-Pro, Kimi-K2.5, and Claude-Opus-4.6-Thinking achieved the best overall scores of 72%-74%. Models performed best in production engineering and weakest in reservoir engineering. Chinese models showed advantages in multiple-choice questions, while international models performed slightly better in short-answer questions. The benchmark provides a reproducible and practical reference for evaluating and deploying LLMs in petroleum engineering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。