arXiv:2505.09438physics.ed-phcs.AI2025-05被引 22

GPT-4o和o1-preview在物理奥赛题上超越人类,引发教育评估新思考。

Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment

  • 对比GPT-4o与o1-preview解奥赛题表现
  • 两者平均得分高于德国奥赛选手
  • o1-preview稳定性更强,适合教育评估

大型语言模型(LLMs)已广泛普及,影响各教育阶段学习者。这引发了对模型使用可能绕过必要学习过程、破坏传统评估体系完整性的担忧。在以问题解决为核心的物理教育中,理解LLMs的物理问题求解能力至关重要,有助于制定负责任且符合教学规律的整合策略。本研究基于一套明确的奥赛题目,比较了通用模型GPT-4o(采用不同提示技巧)与推理优化模型o1-preview的表现,并与德国物理奥赛参赛者进行对比。除评估解题正确性外,还分析了模型解法的特征优势与局限。结果表明,两个测试模型在奥赛类物理问题上均展现出先进求解能力,平均表现优于人类参与者。提示技巧对GPT-4o影响较小,而o1-preview几乎始终优于GPT-4o及人类基准。研究据此探讨了物理教育中总结性与形成性评估的设计启示,包括如何维护评估完整性,以及如何引导学生批判性使用LLMs。

原文摘要 · Abstract (English)

Large language models (LLMs) are now widely accessible, reaching learners at all educational levels. This development has raised concerns that their use may circumvent essential learning processes and compromise the integrity of established assessment formats. In physics education, where problem solving plays a central role in instruction and assessment, it is therefore essential to understand the physics-specific problem-solving capabilities of LLMs. Such understanding is key to informing responsible and pedagogically sound approaches to integrating LLMs into instruction and assessment. This study therefore compares the problem-solving performance of a general-purpose LLM (GPT-4o, using varying prompting techniques) and a reasoning-optimized model (o1-preview) with that of participants of the German Physics Olympiad, based on a set of well-defined Olympiad problems. In addition to evaluating the correctness of the generated solutions, the study analyzes characteristic strengths and limitations of LLM-generated solutions. The findings of this study indicate that both tested LLMs (GPT-4o and o1-preview) demonstrate advanced problem-solving capabilities on Olympiad-type physics problems, on average outperforming the human participants. Prompting techniques had little effect on GPT-4o's performance, while o1-preview almost consistently outperformed both GPT-4o and the human benchmark. Based on these findings, the study discusses implications for the design of summative and formative assessment in physics education, including how to uphold assessment integrity and support students in critically engaging with LLMs.

物理教育大模型评测奥赛题评估设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。