arXiv:2504.06260cs.AIcs.CL2025-04中稿 · NeurIPS被引 24

测试大模型用有限元软件解多物理场问题的能力

FEABench: Evaluating Language Models on Multiphysics Reasoning Ability

  • 设计可与仿真软件交互的智能体,通过API调用求解
  • 最佳策略成功生成有效API调用率达88%
  • 适合研究具身智能与工程自动化的人看

构建真实世界的精确模拟并使用数值求解器回答定量问题是工程与科学的核心需求。我们提出FEABench,一个评估大语言模型(LLMs)及LLM智能体在有限元分析(FEA)框架下模拟和求解物理、数学与工程问题的能力的基准。我们设计了全面的评估方案,考察模型能否基于自然语言问题描述,通过操作COMSOL Multiphysics®软件完成端到端求解。此外,我们开发了一个具备通过API与软件交互、检查输出并迭代优化解决方案能力的语言模型智能体。最佳策略可实现88%的可执行API调用率。能够成功与FEA软件交互并求解复杂问题的模型将推动工程自动化的前沿发展,赋予语言模型数值求解的精度,加速自主系统在真实世界中应对复杂任务的进程。代码已开源。

原文摘要 · Abstract (English)

Building precise simulations of the real world and invoking numerical solvers to answer quantitative problems is an essential requirement in engineering and science. We present FEABench, a benchmark to evaluate the ability of large language models (LLMs) and LLM agents to simulate and solve physics, mathematics and engineering problems using finite element analysis (FEA). We introduce a comprehensive evaluation scheme to investigate the ability of LLMs to solve these problems end-to-end by reasoning over natural language problem descriptions and operating COMSOL Multiphysics$^\circledR$, an FEA software, to compute the answers. We additionally design a language model agent equipped with the ability to interact with the software through its Application Programming Interface (API), examine its outputs and use tools to improve its solutions over multiple iterations. Our best performing strategy generates executable API calls 88% of the time. LLMs that can successfully interact with and operate FEA software to solve problems such as those in our benchmark would push the frontiers of automation in engineering. Acquiring this capability would augment LLMs' reasoning skills with the precision of numerical solvers and advance the development of autonomous systems that can tackle complex problems in the real world. The code is available at https://github.com/google/feabench

大模型仿真工程自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。