arXiv:2505.10653cs.AI2025-05被引 2

提出工程通用智能评估框架,以衡量AI在真实工程设计中的综合能力。

On the Evaluation of Engineering Artificial General Intelligence

  • 基于布鲁姆分类法构建工程化评估体系,覆盖从知识到创造的全链条能力
  • 支持文本与结构化设计成果(如CAD、SysML)的多模态评估
  • 可定制化适配不同工程场景,推动AI工程应用落地

本文探讨了工程通用智能(eAGI)的评估挑战,并提出一个可扩展的评估框架。eAGI是通用智能的一种特殊形式,具备解决物理系统及其控制器设计中广泛问题的能力,但不包含软件工程内容。与人类工程师类似,eAGI应具备知识检索、工具熟悉度、对工业组件和典型设计范式的深入理解,以及跨情境创造性解决问题的能力。为应对这一综合性任务,我们提出将布鲁姆分类法(Bloom's taxonomy)特化并扎根于工程设计语境的评估框架。该框架在基准测试方面实现三方面突破:(a) 构建涵盖方法论知识至真实设计问题的丰富评估题库;(b) 支持文本及结构化设计成果(如CAD模型、SysML模型)的评估;(c) 提出可自动化定制的评估基准流程,以适应不同工程领域需求。

原文摘要 · Abstract (English)

We discuss the challenges and propose a framework for evaluating engineering artificial general intelligence (eAGI) agents. We consider eAGI as a specialization of artificial general intelligence (AGI), deemed capable of addressing a broad range of problems in the engineering of physical systems and associated controllers. We exclude software engineering for a tractable scoping of eAGI and expect dedicated software engineering AI agents to address the software implementation challenges. Similar to human engineers, eAGI agents should possess a unique blend of background knowledge (recall and retrieve) of facts and methods, demonstrate familiarity with tools and processes, exhibit deep understanding of industrial components and well-known design families, and be able to engage in creative problem solving (analyze and synthesize), transferring ideas acquired in one context to another. Given this broad mandate, evaluating and qualifying the performance of eAGI agents is a challenge in itself and, arguably, a critical enabler to developing eAGI agents. In this paper, we address this challenge by proposing an extensible evaluation framework that specializes and grounds Bloom's taxonomy - a framework for evaluating human learning that has also been recently used for evaluating LLMs - in an engineering design context. Our proposed framework advances the state of the art in benchmarking and evaluation of AI agents in terms of the following: (a) developing a rich taxonomy of evaluation questions spanning from methodological knowledge to real-world design problems; (b) motivating a pluggable evaluation framework that can evaluate not only textual responses but also evaluate structured design artifacts such as CAD models and SysML models; and (c) outlining an automatable procedure to customize the evaluation benchmark to different engineering contexts.

通用智能评估框架工程AI设计生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。