测试大模型在大学热力学中的推理能力,发现其表现远未达标。
From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics
- 构建50题热力学问答基准UTQA,覆盖理想气体、可逆性与图示理解
- 顶尖模型最高仅82%准确率,图像理解任务接近随机水平
- 当前大模型不适用于热力学的自主教学,尤其在非理想过程场景
大语言模型(LLMs)在科学教育中被视为潜在辅导工具,但其在无监督本科教学中的可用性仍不确定,因为可靠教学不仅需要流畅的知识复述,更需一致且基于原理的推理。热力学因其简洁定律和状态函数与路径函数、可逆性、熵等微妙区别,是评估此类能力的理想测试平台。本文提出UTQA,一个包含50道题的本科热力学问答基准,涵盖理想气体过程、可逆性及图示解读。2025年主流模型均未达到95%的胜任阈值:最佳模型准确率为82%,纯文本题目表现优于图像推理任务,后者常降至随机水平。提示词表述与句法复杂度对性能影响甚微。性能差距集中在有限速率/不可逆情景,以及将视觉特征与热力学意义关联的能力上,表明当前大模型尚不适合该领域的自主教学。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly considered as tutoring aids in science education. Yet their readiness for unsupervised use in undergraduate instruction remains uncertain, as reliable teaching requires more than fluent recall: it demands consistent, principle-grounded reasoning. Thermodynamics, with its compact laws and subtle distinctions between state and path functions, reversibility, and entropy, provides an ideal testbed for evaluating such capabilities. Here we present UTQA, a 50-item undergraduate thermodynamics question answering benchmark, covering ideal-gas processes, reversibility, and diagram interpretation. No leading 2025-era model exceeded our 95\% competence threshold: the best LLMs achieved 82\% accuracy, with text-only items performing better than image reasoning tasks, which often fell to chance levels. Prompt phrasing and syntactic complexity showed modest to little correlation with performance. The gap concentrates in finite-rate/irreversible scenarios and in binding visual features to thermodynamic meaning, indicating that current LLMs are not yet suitable for unsupervised tutoring in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。