用大模型辅助评估编程考题难度,提升考试公平性。
From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations
- 将大模型作为试题难度的外部证据源,结合学生表现等多维数据。
- 模型解题通过率与学生通过率相关性达0.866,验证其有效性。
- 适合教育研究者和考试设计者用于考题质量分析与对比。
平行课程的编程考试难度差异影响课程评估的公平性。本研究将大语言模型从评测目标转变为辅助证据来源,融合人工智能解题表现、学生作答数据、题目曝光度、在线判题过程记录及教师判断。首先,10个模型同步解答包含8道题的期末考,模型通过率与学生通过率呈正相关(Spearman rho = 0.866,p = 0.0119),基于解题的综合难度指数与通过率负相关(rho = -0.905,p = 0.0046)。随后,使用第三方兼容OpenAI接口的单结构化评审模型(gpt-5.6-sol)处理11套平行课程共79道题,模型整体难度与题目通过率相关性为-0.871,未作答率相关性为0.800;在26题的算法课程纵向样本中,相关性分别为-0.829与0.883。106题的入门课程(CS101)样本显示,题目级相关性降至-0.552,16次考试的总体相关性接近零,说明群体构成主导考试层级结果。暴露度折扣(0–0.40)与重复题扰动测试未改变趋势方向。因此,AI证据可作为题目验证、平行考试公平性讨论及长期质量追踪的外部参考,但模型身份边界、单评审设计与输出不稳定性意味着其难度评分不得用于个体学生评价或自动评分调整。
原文摘要 · Abstract (English)
Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions large language models from benchmark evaluation targets to auxiliary evidence sources for interpreting exam difficulty, combining AI evidence with aggregated student performance, item exposure, online-judge process data, and teacher interpretation. First, ten models solved an eight-problem final exam synchronously with 120 students: AI pass rate correlated positively with student pass rate (Spearman rho = 0.866, exact p = 0.0119), and a solving-based composite difficulty index correlated negatively with it (rho = -0.905, exact p = 0.0046). A single structured reviewer was then run via auditable API calls on a third-party OpenAI-compatible endpoint whose model label (gpt-5.6-sol) cannot authenticate an official OpenAI upstream model; call metadata and raw responses are archived. Across 79 problems from 11 parallel-class final exams, AI overall difficulty correlated with problem-level pass rate at rho = -0.871 and with non-attempt rate at rho = 0.800; in a 26-problem longitudinal Data Structures and Algorithms B sample, the correlations were -0.829 and 0.883. A 106-problem introductory-course (CS101) sample marks the boundary: the problem-level correlation weakened to rho = -0.552, and the exam-level correlation across 16 exams was near zero, with cohort composition dominating exam-level outcomes. Exposure-discount (0-0.40) and duplicate-problem perturbation tests did not change these directions. AI evidence can thus serve as an external reference for problem validation, parallel-class fairness discussion, and longitudinal quality tracking, while the model-identity boundary, single-reviewer design, and review-output instability set explicit limits: AI difficulty scales must not be used for individual student evaluation or automatic grade adjustment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。