arXiv:2606.26103cs.CLcs.AI2026-06

用知识蒸馏法测试大模型解静力学题能力,发现图形和多步推理是难点

Investigating LLM's Problem Solving Capability -- a Study on Statics Questions

  • 通过蒸馏ChatGPT生成25道纯文本静力学题
  • 引入图表后准确率下降,多步推理错误增多
  • 核心瓶颈是逻辑连贯性而非图像识别能力

大型语言模型(LLMs)在教育领域影响广泛,因其在多个学科中完成作业和考试的能力而备受关注。然而,现有研究多依赖公开题库,缺乏针对特定题型的系统分析。在机械工程教育中,对LLM在特定问题类型上的表现研究仍不充分。本研究摒弃直接提问教材题目的方式,采用模型蒸馏方法评估LLM解决静力学问题的能力。通过蒸馏ChatGPT,提取出25道纯文本静力学问题,并进一步构建了两个新数据集:一个加入图表,另一个修改数值。实验结果显示,尽管LLM在纯文本问题上表现良好,但引入图表后准确率下降,尤其在需要多步推理的问题上。进一步分析表明,性能下降并非主要源于图像识别能力不足,而是难以维持多步骤推理中的逻辑一致性,以及未能在后续步骤中持续利用视觉信息。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have rapidly influenced many aspects of society, particularly education, due to their demonstrated ability to complete assignments and examinations across a wide range of subjects. Although prior studies have examined the educational impact of LLMs, much of the existing work relies on public or open problem datasets and lacks topic-specific analysis. In engineering education, especially within mechanical engineering, systematic investigations of LLM performance on specific problem types remain limited. Instead of using traditional methods that directly ask textbook questions to an LLM tool, our study adopts a model distillation process to evaluate LLM capabilities in solving statics problems. By distilling ChatGPT, we extracted 25 text-only statics questions and further constructed two additional datasets by adding diagrams and modifying their numerical values. Experimental results show that while LLMs perform well on text-only statics problems, their accuracy decreases when diagrams are introduced and the problems require multi-step reasoning. Further analysis suggests that this performance drop is not primarily caused by limitations in image recognition, but rather by difficulties in multi-step reasoning and in consistently applying extracted visual information across successive solution stages.

大模型静力学推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。