arXiv:2512.03672cs.CL2025-12

为评估大模型在水利科学与工程领域的知识水平,构建了包含4000道题的评测基准。

Evaluating Hydro-Science and Engineering Knowledge of Large Language Models

  • 构建覆盖9个子领域的4000道多选题评测集,涵盖概念、应用与计算能力
  • 商用大模型准确率0.74~0.80,小模型仅0.41~0.68,工程应用能力不足
  • 揭示大模型在行业标准和水工结构等专有知识上的短板,适合水利研究者参考

水利科学与工程(Hydro-SE)是保障人类供水、提供清洁水能、缓解洪旱灾害的关键领域,具有多工程目标,融合科学与工程知识,需专家协同决策,对智能系统构成挑战。随着大语言模型(LLMs)快速发展,其在水利领域的应用潜力日益受到关注,但其知识与应用能力尚未充分评估。为此,本文提出水利科学与工程大模型评测基准(Hydro-SE Bench),包含4,000道多选题,覆盖9个子领域,可评估基础概念、工程应用及推理计算能力。评测结果显示,商业大模型准确率在0.74至0.80之间,小参数模型在0.41至0.68之间。大模型在自然与物理科学相关子领域表现较好,但在行业标准、水工结构等领域知识掌握不足。模型规模主要提升推理与计算能力,但实际工程应用仍有巨大提升空间。本研究揭示了大模型在水利任务中的优劣势,为模型开发者提供明确训练方向,也为水利研究人员提供实用指导。

原文摘要 · Abstract (English)

Hydro-Science and Engineering (Hydro-SE) is a critical and irreplaceable domain that secures human water supply, generates clean hydropower energy, and mitigates flood and drought disasters. Featuring multiple engineering objectives, Hydro-SE is an inherently interdisciplinary domain that integrates scientific knowledge with engineering expertise. This integration necessitates extensive expert collaboration in decision-making, which poses difficulties for intelligence. With the rapid advancement of large language models (LLMs), their potential application in the Hydro-SE domain is being increasingly explored. However, the knowledge and application abilities of LLMs in Hydro-SE have not been sufficiently evaluated. To address this issue, we propose the Hydro-SE LLM evaluation benchmark (Hydro-SE Bench), which contains 4,000 multiple-choice questions. Hydro-SE Bench covers nine subfields and enables evaluation of LLMs in aspects of basic conceptual knowledge, engineering application ability, and reasoning and calculation ability. The evaluation results on Hydro-SE Bench show that the accuracy values vary among 0.74 to 0.80 for commercial LLMs, and among 0.41 to 0.68 for small-parameter LLMs. While LLMs perform well in subfields closely related to natural and physical sciences, they struggle with domain-specific knowledge such as industry standards and hydraulic structures. Model scaling mainly improves reasoning and calculation abilities, but there is still great potential for LLMs to better handle problems in practical engineering application. This study highlights the strengths and weaknesses of LLMs for Hydro-SE tasks, providing model developers with clear training targets and Hydro-SE researchers with practical guidance for applying LLMs.

大模型评测水利工程知识评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。