arXiv:2607.27586cs.CYcs.AI2026-07

让学生评估AI生成的解法,效果不比自己解题差,但需引导才能真学懂。

Is Solving Better Than Evaluating GenAI Solutions?

  • 对比学生自解题与评估AI解法,设计交叉实验
  • 评题组作业分更高,但考试成绩无显著差异
  • 适合想培养批判性思维的学生,需额外指导

随着生成式AI在解决计算任务上能力增强,教育界开始探索以评估、验证和批判为核心的教学方法。然而,这类评价导向任务对学生学习的影响仍缺乏证据,尤其在高年级理论课程中。本研究在一门大二算法课中开展随机化A/B交叉实验(N=220),比较学生直接解题与评估常含错误的AI生成解法的效果。六次作业中,小组角色中途对调。结果显示,两组在期中、期末、总成绩及与作业结构一致的考题上无统计显著差异;但评估组作业得分更高,该优势未转化为后续考核成绩提升。调查显示多数学生未改变学习习惯,但调整策略者认为评估任务更助学习。结果表明,评估AI解法将学生精力从构造转向验证与判断,但不会自动带来概念迁移。结论是:可在算法课程中融入生成式评估活动,但要实现深层学习,需设计有效支架引导学生超越简单纠错。

原文摘要 · Abstract (English)

As Generative AI (GenAI) tools become increasingly capable of generating solutions to computing assignments, the computing education community is exploring pedagogical approaches that emphasize solution evaluation, verification, and critique alongside traditional solution generation. However, evidence regarding the impact of such evaluation-centered tasks on student learning remains limited, particularly in upper-division, theory-heavy courses. We conducted a randomized A/B crossover study (N=220) in a junior-level algorithms course to compare evaluating GenAI-generated solutions with traditional problem solving. Across six assignments, student working groups either solved challenging algorithmic problems directly or evaluated often-flawed GenAI-generated solutions, with roles reversed midway through the semester. We found no statistically significant differences between groups in midterm scores, final exam scores, overall course grades, or exam problems structurally aligned with the homework interventions. Students received significantly higher homework scores when evaluating GenAI-generated solutions, but this localized advantage did not translate into downstream summative gains. Survey data further indicated that most students reported no change in study habits in response to the intervention; however, those who reported adapting their study strategies rated the GenAI-evaluation assignments as significantly more helpful. These findings suggest that GenAI evaluation redistributes student effort from open-ended solution construction toward verification, diagnosis, and judgment, but does not automatically produce stronger conceptual transfer. We conclude that GenAI-evaluation activities can be incorporated into algorithms coursework without broad performance losses, but meaningful learning gains may require deliberate scaffolding that pushes students beyond simple error diagnosis.

生成式AI算法教学学习评估教育实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。