大模型能生成符合学生错误思维的选项,关键在于先解对题再模拟错因。
Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation
- 先正确解题,再模拟多种可能错误,最后选合适干扰项
- 提供正确答案可使生成结果更接近人工标注,提升8%
- 适合教育AI、自动出题系统研究者参考
在教育人工智能中,建模合理的学生误解至关重要。本文研究大语言模型(LLMs)在生成多项选择题干扰项时对错误思路的建模能力,该任务需协调解题知识、模拟学生误解并评估合理性。我们提出一个分析框架,考察当前顶尖模型的推理策略,并与学习科学中的最佳实践对比。结构化分析发现,模型通常先正确解题,再生成并模拟多个潜在误解,最后筛选干扰项。失败主要源于未能恢复正确解或在候选项间选择失误,而非误解模拟或流程设计问题。结果显示,在提示中加入正确答案可使生成结果与人工干扰项的匹配度提高8%,凸显锚定正确解对生成合理错误思路的关键作用。整体而言,本研究为理解大模型建模学生错误思维的能力提供了可解释的分析视角。
原文摘要 · Abstract (English)
Modeling plausible student misconceptions is critical for AI in education. In this work, we examine how large language models (LLMs) reason about misconceptions when generating multiple-choice distractors, a task that requires modeling incorrect yet plausible answers by coordinating solution knowledge, simulating student misconceptions, and evaluating plausibility. We introduce a taxonomy for analyzing the strategies used by state-of-the-art LLMs, examining their reasoning procedures and comparing them to established best practices in the learning sciences. Our structured analysis reveals a surprising alignment between their processes and best practices: the models typically solve the problem correctly first, then articulate and simulate multiple potential misconceptions, and finally select a set of distractors. An analysis of failure modes reveals that errors arise primarily from failures in recovering the correct solution and selecting among response candidates, rather than simulating errors or structuring the process. Consistent with these results, we find that providing the correct solution in the prompt improves alignment with human-authored distractors by 8%, highlighting the critical role of anchoring to the correct solution when generating plausible incorrect student reasoning. Overall, our analysis offers a structured and interpretable lens into LLMs' ability to model incorrect student reasoning and produce high-quality distractors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。