用真实工程场景评估大模型,发现其在抽象思维上表现弱。
Evaluating Large Language Models for Real-World Engineering Tasks
- 构建100+真实工程问题数据集,覆盖设计、诊断等核心能力。
- 大模型在基础时序和结构推理中表现较好,但抽象建模能力差。
- 适合关注AI工程应用局限性的研究人员和工程师参考。
大型语言模型(LLMs)不仅改变日常应用,也正在影响工程任务。然而当前对LLMs在工程领域的评估存在两大缺陷:(i) 依赖简化用例,多来自考试题,正确性易验证;(ii) 使用随意场景,无法充分反映关键工程能力。因此,大模型在复杂真实工程问题上的表现仍缺乏系统评估。本文通过构建包含100多个源自真实生产场景的工程问题数据集,系统覆盖产品设计、故障预测与诊断等核心能力,评估了四种先进LLM(含云端与本地部署实例)。结果表明,大模型在基础时序与结构推理方面表现良好,但在抽象推理、形式化建模及上下文敏感的工程逻辑方面存在显著不足。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are transformative not only for daily activities but also for engineering tasks. However, current evaluations of LLMs in engineering exhibit two critical shortcomings: (i) the reliance on simplified use cases, often adapted from examination materials where correctness is easily verifiable, and (ii) the use of ad hoc scenarios that insufficiently capture critical engineering competencies. Consequently, the assessment of LLMs on complex, real-world engineering problems remains largely unexplored. This paper addresses this gap by introducing a curated database comprising over 100 questions derived from authentic, production-oriented engineering scenarios, systematically designed to cover core competencies such as product design, prognosis, and diagnosis. Using this dataset, we evaluate four state-of-the-art LLMs, including both cloud-based and locally hosted instances, to systematically investigate their performance on complex engineering tasks. Our results show that LLMs demonstrate strengths in basic temporal and structural reasoning but struggle significantly with abstract reasoning, formal modeling, and context-sensitive engineering logic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。