构建工程推理新基准,用八阶段评估法揭示大模型在工程问题上的真实短板。
Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation

- 设计8阶段自动化评估框架,逐级检验工程解题过程
- 在696道跨5个工程领域的题目上,发现主流模型严重不足
- 人工评分与自动评分高度一致,验证评估体系可靠性
视觉语言模型在通用多模态推理任务中表现优异,但其工程推理能力仍待探索。与一般视觉问答不同,工程问题求解需解读技术图示、选择物理原理并保持多步推理的物理一致性。这类能力对工程教育、科研辅助和决策支持至关重要,因推理错误可能导致看似合理却物理无效的解。现有基准主要评估最终答案,缺乏对中间推理过程的细致分析。本文提出EngVQA,一个涵盖5个工程学科、包含696道题目的多模态基准,并设计8阶段自动化评估框架,可独立评价每一步解题过程,实现细粒度推理失败分析。我们在此框架上测试多个前沿开源与闭源VLM,揭示其工程推理能力存在显著局限。人工评估显示与自动评分高度一致(皮尔逊相关系数0.975,平均绝对误差0.67,在10分制下),证实该评估体系的可靠性。结果强调过程导向评估对可靠评估多模态工程推理系统的重要性。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) demonstrate strong performance on general multimodal reasoning benchmarks, yet their ability to perform engineering reasoning remains largely unexplored. Unlike general visual question answering, engineering problem solving requires interpreting technical diagrams, selecting governing physical principles, and maintaining physically consistent multi-step reasoning. These capabilities are increasingly important for AI systems used in engineering education, scientific assistance, and technical decision-making, where reasoning failures may produce physically invalid yet superficially plausible solutions. Existing benchmarks primarily evaluate final answers and provide limited assessment of intermediate reasoning processes. We introduce EngVQA, a multimodal benchmark for evaluating engineering reasoning across 5 engineering subjects containing 696 problems. We introduce an 8-stage automatic evaluation framework for assessing VLM-generated solutions. The framework independently evaluates each stage of the solution, enabling fine-grained analysis of reasoning failures. We benchmark multiple state-of-the-art open and closed source VLMs on our evaluation framework and demonstrate substantial limitations in current engineering reasoning capabilities. Human evaluation shows strong agreement with our automated framework, achieving a Pearson correlation of 0.975 and a mean absolute error of 0.67 on a 10-point grading scale. Our results highlight the importance of process-oriented evaluation for reliable assessment of multimodal engineering reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。