评测大模型在流体模拟中的科学能力,填补了其在复杂物理系统自动化上的空白。
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics
- 构建三类任务:流体知识、数值推理与代码实现,全面评估大模型能力
- 在真实流体模拟场景中验证,模型代码执行成功率与解的精度有显著差异
- 适合研究科学计算自动化、大模型在工程领域应用的学者使用
大型语言模型(LLMs)在通用自然语言任务中表现优异,但在自动化复杂物理系统数值实验——这一关键且耗时的任务——方面的应用仍待探索。作为过去几十年计算科学的核心工具,计算流体动力学(CFD)为评估大模型的科学能力提供了独特挑战。我们提出CFDLLMBench,一个包含三个互补组件的基准套件:CFDQuery、CFDCodeBench和FoamBench,旨在全面评估大模型在三个关键能力上的表现:研究生水平的CFD知识、CFD的数值与物理推理能力,以及上下文依赖的CFD工作流实现能力。该基准基于真实的CFD实践,结合详细的任务分类与严谨的评估框架,可重复地衡量模型在代码可执行性、解的准确性和数值收敛行为方面的性能。CFDLLMBench为发展和评估大模型驱动的复杂物理系统数值实验自动化奠定了坚实基础。代码与数据可在https://github.com/NREL-Theseus/cfdllmbench/获取。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated strong performance across general NLP tasks, but their utility in automating numerical experiments of complex physical system -- a critical and labor-intensive component -- remains underexplored. As the major workhorse of computational science over the past decades, Computational Fluid Dynamics (CFD) offers a uniquely challenging testbed for evaluating the scientific capabilities of LLMs. We introduce CFDLLMBench, a benchmark suite comprising three complementary components -- CFDQuery, CFDCodeBench, and FoamBench -- designed to holistically evaluate LLM performance across three key competencies: graduate-level CFD knowledge, numerical and physical reasoning of CFD, and context-dependent implementation of CFD workflows. Grounded in real-world CFD practices, our benchmark combines a detailed task taxonomy with a rigorous evaluation framework to deliver reproducible results and quantify LLM performance across code executability, solution accuracy, and numerical convergence behavior. CFDLLMBench establishes a solid foundation for the development and evaluation of LLM-driven automation of numerical experiments for complex physical systems. Code and data are available at https://github.com/NREL-Theseus/cfdllmbench/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。