arXiv:2511.01650cs.CLcs.AI2025-11被引 3

构建可验证的工程推理基准,测试大模型在真实物理约束下的逐步推理能力。

EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning

  • 基于90个参数化模板生成1350个抗污染问题,覆盖三大工程领域
  • 发现数值精度与推理过程一致性存在明显权衡,复杂任务出现能力断崖
  • 引入自动化流程校验+多智能体评审,实现对中间步骤的可验证评估

大型语言模型正日益应用于受严格量化标准和不可变物理定律约束的安全关键型工程流程,因此对其推理能力进行严格评估至关重要。然而,现有基准如MMLU、MATH和HumanEval仅评估孤立认知技能,无法捕捉工程中以物理为基础的推理核心——科学原理、定量建模与实际约束必须协同。为实现工程推理的可验证过程监督,我们提出EngTrace,一个基于90个参数化模板的符号基准,每个模板生成唯一且抗污染的问题实例,覆盖三个主要工程分支、九个核心领域和20个具体方向,共产生1,350个测试用例,全面检验模型在多样化物理场景下的泛化能力。超越仅比对最终答案,我们设计了一套可验证的两阶段评估框架,通过分层协议结合自动化流程检查与异构AI评审团,验证中间推理轨迹与最终答案。对27个领先大模型的评估揭示了数值精度与推理轨迹保真度之间的显著权衡,在复杂任务中出现能力断崖,表明抽象数学预训练难以转化为高级工程所需的整合性推理能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly entering specialized, safety-critical engineering workflows governed by strict quantitative standards and immutable physical laws, making rigorous evaluation of their reasoning capabilities imperative. However, existing benchmarks such as MMLU, MATH, and HumanEval assess isolated cognitive skills, failing to capture the physically grounded reasoning central to engineering, where scientific principles, quantitative modeling, and practical constraints must converge. To enable verifiable process supervision in engineering, we introduce EngTrace, a symbolic benchmark built on 90 parameterized templates, each generating unique, contamination-resistant problem instances, spanning three major engineering branches, nine core domains, and 20 distinct areas, yielding 1,350 test cases that stress-test generalization across diverse physical scenarios. Moving beyond outcome matching, we introduce a verifiable two-stage evaluation framework that uses a tiered protocol to validate intermediate reasoning traces alongside final answers through automated procedural checks and a heterogeneous AI Tribunal. Our evaluation of 27 leading LLMs reveals a distinct trade-off between numeric precision and trace fidelity, identifying a complexity cliff where abstract mathematical pre-training fails to translate into the integrative reasoning required for advanced engineering tasks.

工程推理可验证评估大模型评测符号基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。