顶级AI能解大学物理题,但图文题和难题表现差
Assessing AI in Introductory Physics Problem Solving
- 用OpenAI的o4-mini模型测试物理题求解能力
- 纯文字题准确率96%,图文混合题仅79%
- 难题准确率显著下降,适合评估AI物理推理边界
为探究推理型大模型在物理学问题求解中的能力,我们使用OpenAI的o4-mini模型,对霍尔iday与雷斯尼克《物理学基础》中的传统课后习题进行评估,覆盖本科物理核心内容。分析了模态与题目难度的影响。模型整体准确率约90%,但在纯文本题中达96%,而需图文协同理解的题目准确率降至79%。随着题目难度从低到高递增,准确率显著下降。结果表明,当前最先进的大模型虽可解决多数标准入门物理题,但其表现仍受题目呈现形式与难度制约。
原文摘要 · Abstract (English)
Reasoning or inference-scaling models are the new generation of Large Language Models (LLMs) capable of complex problem solving. To investigate their problem-solving capability in physics, we evaluated model o4-mini by OpenAI on solving traditional, end-of-chapter problems from Halliday and Resnick's "Fundamentals of Physics," spanning core topics in the undergraduate physics curriculum. Performance was analyzed across modality and problem difficulty. The model solved the problems with overall accuracy of about 90%, but performance depended strongly on representation: accuracy was much higher on text-only problems (96%) than on problems requiring coordinated interpretation of text and images (79%). Accuracy also declined significantly as the problem difficulty increased from low to medium to high. These results show that state-of-the-art LLMs can solve much of the standard introductory physics problems, but that their performance remains uneven and constrained by problem modality and problem difficulty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。