AI辅导系统常误判错误推理得正确答案的情况,需警惕高准确率下的盲区。
Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning

- 分析真实学生答题数据,发现71%的推理误判集中在两类题型。
- 顶尖模型检测准确率仅57%,每检出1个错误会误报4次。
- 强调需人类判断弥补AI在推理评估中的根本缺陷。
智能辅导系统日益提供对学生作业的自动化反馈,但可靠反馈需评估推理过程而非仅最终答案。本文研究一种名为‘正确答案陷阱’(CAT)的失效模式:当学生通过错误推理得出正确答案时,模型难以识别其误解。基于Eedi数学平台的真实学生作答数据,我们发现71%的此类失败集中于两类具有共同结构的题目,其中错误推理恰好导致正确数值结果。对比微调后的T5与前沿大语言模型,发现更强能力虽可缓解问题(检测准确率84% vs 57%),但无法根除。即使表现最佳的模型也存在约四次误报对应一次真实检测,使独立筛查在真实班级规模下不可行。研究揭示:高总体准确率可能掩盖推理评估中的关键缺陷,因此对推理的细致分析仍需依赖人工判断。
原文摘要 · Abstract (English)
Intelligent tutoring systems increasingly provide automated feedback on student work, but robust feedback requires assessing reasoning, not only final answers. We study a failure mode we call the correct answer trap (CAT): models under-detect misconceptions when students reach a correct answer via flawed reasoning. Analysing real student responses from the Eedi mathematics platform, we show that 71% of these failures concentrate in just two question types, both sharing a common structure where flawed reasoning happens to produce the correct numerical answer. Comparing a fine-tuned T5 with a frontier large language model, we find that improved capabilities reduce but do not eliminate the problem (84% vs 57% detection accuracy). Even the best-performing model generates roughly four false alarms for every genuine detection, making stand-alone screening impractical at realistic class sizes. Our findings demonstrate that high overall accuracy can mask critical failures in reasoning assessment, and that careful analysis of student reasoning still benefits from human judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。