评测大模型在机械与空间几何任务中的工程理解能力。
Large Language Models and their Awareness of Mechanics and Spatial Geometry

- 构建自动化基准MecEng,从文本生成多体动力学仿真模型。
- 开源模型在刚体任务中成功率86.0%,专用模型达91.4%。
- 适合关注大模型工程推理能力的研究者和开发者。
大型语言模型(LLMs)在代码生成和数学推理任务上表现优异,但在机械与空间几何领域(统称机械工程意识)的能力尚未系统量化。本文提出MecEng,一个全自动基准测试,评估LLMs从参数化文本描述生成多体仿真模型的能力。该基准包含84个通用任务,分三个难度等级,涵盖含关节与接触的刚体系统,以及需精确3D几何生成、四面体有限元网格划分和机器部件的Hurty-Craig-Bampton降阶建模的柔性多体系统。通过专用流水线,利用Netgen生成仿真就绪几何,并为Exudyn构建多体系统模型,再通过系统图同构性(含节点标注)、数值解及质量、几何、固有频率等部件特异性指标与专家真值对比验证。共评估32个开源与两个专有模型。在刚体任务中,最佳开源模型成功率达86.0%,最强专有模型为91.4%;柔性多体任务仍具挑战。额外研究分析了采样温度、推理策略、提示设计、模型规模与发布日期的影响。结果表明当前大模型的机械工程意识正快速提升,但仍有误判风险。
原文摘要 · Abstract (English)
Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in mechanics and spatial geometry, here denoted as mechanical engineering awareness, has not been quantified systematically. We present MecEng, a fully automated benchmark that evaluates LLMs on the creation of multibody simulation models from parameterized textual descriptions. The benchmark comprises 84 generic tasks on three difficulty levels, ranging from rigid-body systems with joints and contact to flexible multibody systems that require exact 3D geometry generation, tetrahedral finite-element meshing, and Hurty-Craig-Bampton model order reduction of machine parts. A dedicated pipeline with LLMs generates simulation-ready geometry from text using Netgen, and builds multibody system models for the code Exudyn, which are then verified against expert ground truth on several levels: system-graph isomorphism including graph node annotations, numerical solutions, and part-specific measures such as mass, geometry, and eigenfrequencies. In total, 32 open-weight and two proprietary LLMs are evaluated. On rigid-body tasks, the best open-weight model obtains an overall success rate of 86.0%, compared to 91.4% for the strongest proprietary model, while flexible multibody tasks remain considerably harder. Additional studies quantify the influence of sampling temperature, reasoning, prompt design, model size, and LLM-release date. The results indicate rapidly improving, but still error-prone, mechanical engineering awareness of current LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。