构建工程领域大模型评估基准,覆盖9大专业55子领域
Engineering Reasoning and Instruction (ERI) Benchmark: A Large Taxonomy-driven Dataset for Foundation Models and Agents
- 按学科分类设计跨领域指令数据,含7类意图与3级难度
- 顶尖模型平均得分超4.3,研究生级问题表现明显下滑
- 采用多模型验证降低幻觉风险,仅1.7%存在误判
工程推理与指令(ERI)基准是一个基于分类体系的指令数据集,用于训练和评估具备工程能力的大语言模型(LLMs)与智能体。该数据集涵盖土木、机械、电气、化工、环境、航空航天、材料、消防和工业工程共9个工程领域及55个子领域,与7种意图类型(定义、解释、计算、比较、设计/合成、故障排查、代码相关)和3个难度层级(本科、研究生、专业级)交叉,生成57,750条带领域/子领域/意图/难度元信息的记录,并规范解决方案格式。我们通过7个LLM对ERI进行评估,发现显著的三级性能结构:前沿模型(GPT-5、Claude Sonnet 4、DeepSeek V3.1)在五分制下平均得分超过4.30,而中等和小型模型在研究生级问题上失败率更高且性能衰减更剧烈。为应对模型评估中的循环性问题,我们提出收敛性验证协议,结合跨提供商独立性、多评审员平均与前沿模型一致性分析,将幻觉风险实证控制在1.7%以内。ERI随同分类规范、验证脚本和评估工具包发布,支持指令微调、路由、检索增强评估及智能体工具使用流程的可复现对比与回归测试。
原文摘要 · Abstract (English)
The Engineering Reasoning and Instruction (ERI) benchmark is a taxonomy-driven instruction dataset designed to train and evaluate engineering-capable large language models (LLMs) and agents. This dataset spans nine engineering fields (namely: civil, mechanical, electrical, chemical, environmental, aerospace, materials, fire, and industrial engineering) and 55 subdomains, and is crossed with seven intent types (i.e., definition, explanation, calculation, comparison, design/synthesis, troubleshooting, and code-related) and three difficulty tiers (undergraduate, graduate, and professional), yielding 57,750 records with field/subdomain/type/difficulty metadata and solution formatting. We examined ERI via seven LLMs and report a statistically significant three-tier performance structure, with frontier models (GPT-5, Claude Sonnet 4, DeepSeek V3.1) achieving mean scores above 4.30 on a five-point scale, while mid-tier and smaller models exhibited progressively higher failure rates and steeper performance degradation on graduate-level questions. To address circularity concerns inherent in LLM benchmarks, we developed a convergent validation protocol that leverages cross-provider independence, multi-judge averaging, and frontier-model agreement analysis to empirically bound hallucination risk to 1.7%. ERI is released with taxonomy specifications, validation scripts, and an evaluation harness to enable reproducible comparisons and regression testing for instruction tuning, routing, retrieval-augmented evaluation, and agentic tool-use workflows in engineering settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。