构建首个跨九领域工程设计评测基准,测试大模型真实设计能力。
Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
- 设计九类真实工程任务,强调知识融合与约束推理。
- 采用仿真验证替代静态答案判断,评估设计功能性。
- 适合关注工程AI、AGI落地的研究者和开发者。
现代工程涵盖电气、机械、航空航天、土木及计算机等九大领域,是人类文明基石。然而,工程设计对大语言模型(LLMs)构成与传统问答不同的挑战:需整合领域知识、权衡复杂取舍,并处理耗时繁琐的实际流程。现有基准多聚焦于事实记忆或问答,无法反映工程设计的真实需求。本文提出EngDesign,首个跨九领域工程设计评测基准,通过真实世界设计任务(含目标、约束与性能要求)评估LLMs的综合设计能力。其创新在于采用仿真驱动的动态验证范式,超越课本知识,实现功能性的实时验证,标志着迈向工程型通用人工智能(AGI)的关键一步。
原文摘要 · Abstract (English)
Modern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. However, engineering design poses a fundamentally different challenge for large language models (LLMs) compared with traditional textbook-style problem solving or factual question answering. Although existing benchmarks have driven progress in areas such as language understanding, code synthesis, and scientific problem solving, real-world engineering design demands the synthesis of domain knowledge, navigation of complex trade-offs, and management of the tedious processes that consume much of practicing engineers' time. Despite these shared challenges across engineering disciplines, no benchmark currently captures the unique demands of engineering design work. In this work, we introduce EngDesign, an Engineering Design benchmark that evaluates LLMs' abilities to perform practical design tasks across nine engineering domains. Unlike existing benchmarks that focus on factual recall or question answering, EngDesign uniquely emphasizes LLMs' ability to synthesize domain knowledge, reason under constraints, and generate functional, objective-oriented engineering designs. Each task in EngDesign represents a real-world engineering design problem, accompanied by a detailed task description specifying design goals, constraints, and performance requirements. EngDesign pioneers a simulation-based evaluation paradigm that moves beyond textbook knowledge to assess genuine engineering design capabilities and shifts evaluation from static answer checking to dynamic, simulation-driven functional verification, marking a crucial step toward realizing the vision of engineering Artificial General Intelligence (AGI).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。