arXiv:2505.13664cs.CYcs.CL2025-05被引 1

GPT-4o考试不及格,o1-preview超学生平均水平

Assessing GPT Performance in a Proof-Based University-Level Course Under Blind Grading

  • 用匿名评分法测试大模型解算法题表现
  • GPT-4o未达及格线,o1-preview超过学生中位数
  • 模型常犯无依据断言和误导性推理

随着大型语言模型(LLMs)的发展,其在高等教育中的应用,特别是在自由作答的问题求解方面,亟需深入评估。本研究在本科算法课程的真实教学情境下,评估了GPT-4o与o1-preview的表现。将匿名生成的解题答案交由不知来源的教学助理进行评分。分析涵盖粗粒度成绩(得分)与细粒度推理质量(错误模式)。结果显示,GPT-4o持续表现不佳,未能达到及格线;而o1-preview表现显著更好,不仅超过及格分数,部分题目甚至高于学生中位数。然而,两个模型均存在无依据断言和误导性论证问题。这些发现凸显了在教育中建立稳健评估策略与面向AI的评分政策的必要性。

原文摘要 · Abstract (English)

As large language models (LLMs) advance, their role in higher education, particularly in free-response problem-solving, requires careful examination. This study assesses the performance of GPT-4o and o1-preview under realistic educational conditions in an undergraduate algorithms course. Anonymous GPT-generated solutions to take-home exams were graded by teaching assistants unaware of their origin. Our analysis examines both coarse-grained performance (scores) and fine-grained reasoning quality (error patterns). Results show that GPT-4o consistently struggles, failing to reach the passing threshold, while o1-preview performs significantly better, surpassing the passing score and even exceeding the student median in certain exercises. However, both models exhibit issues with unjustified claims and misleading arguments. These findings highlight the need for robust assessment strategies and AI-aware grading policies in education.

大模型评估算法题教育应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。