评测大模型生成物理模拟代码的能力,填补科学建模评估空白。
FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
- 构建基于有限元法的结构化评测基准,要求生成符合物理规律的代码。
- 顶尖模型仅完成30/33任务一次,五次尝试中仅26/33稳定通过。
- 适合关注物理推理与科学计算的AI研究者,推动可验证的科学建模发展。
随着大模型在物理世界推理能力上的进步,缺乏对其生成科学有效物理模型能力的严格评测已成为关键短板。计算力学通过数学模型与数值方法预测物理系统在受力、变形和约束下的行为,为结构化科学推理评估提供了理想基础。该领域问题具有清晰的数学结构,强制执行严格的物理与数值约束,并支持客观验证。它要求显式构建物理系统模型,推理几何关系、空间位置及材料行为,直接对接当前人工智能在物理推理与世界建模方面的目标。本文提出FEM-Bench,一个用于评估大模型生成有限元法(FEM)及相关代码能力的计算力学基准。FEM-Bench 2025包含一套入门级但非平凡的任务,涵盖一门研究生计算力学课程的核心内容。这些任务捕捉了基本的数值与物理建模挑战,仅代表该学科复杂性的一小部分。尽管任务相对简单,当前最先进的大模型仍无法可靠完成全部任务。在五次尝试中,表现最佳的函数生成模型Gemini 3 Pro至少完成30/33个任务,五次均成功完成26/33个。在单元测试生成方面,表现最佳的GPT-5平均联合成功率达73.8%。其他主流模型性能差异显著。FEM-Bench建立了评估人工智能生成科学代码的结构化基础,未来版本将引入更复杂的任务,以追踪模型演进进展。
原文摘要 · Abstract (English)
As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to generate scientifically valid physical models has become a critical gap. Computational mechanics, which develops and applies mathematical models and numerical methods to predict the behavior of physical systems under forces, deformation, and constraints, provides an ideal foundation for structured scientific reasoning evaluation. Problems follow clear mathematical structure, enforce strict physical and numerical constraints, and support objective verification. The discipline requires constructing explicit models of physical systems and reasoning about geometry, spatial relationships, and material behavior, connecting directly to emerging AI goals in physical reasoning and world modeling. We introduce FEM-Bench, a computational mechanics benchmark designed to evaluate the ability of LLMs to generate correct finite element method (FEM) and related code. FEM-Bench 2025 contains a suite of introductory but nontrivial tasks aligned with material from a first graduate course on computational mechanics. These tasks capture essential numerical and physical modeling challenges while representing only a small fraction of the complexity present in the discipline. Despite their simplicity, state-of-the-art LLMs do not reliably solve all of them. In a five attempt run, the best performing model at function writing, Gemini 3 Pro, completed 30/33 tasks at least once and 26/33 tasks all five times. The best performing model at unit test writing, GPT-5, had an Average Joint Success Rate of 73.8%. Other popular models showed broad performance variation. FEM-Bench establishes a structured foundation for evaluating AI-generated scientific code, and future iterations will incorporate increasingly sophisticated tasks to track progress as models evolve.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。