LLMs会犯和学生一样的错,误选常见干扰项。
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
- 对比学生与LLM在多选题中的错误选项选择概率
- 两者对常见错误选项的偏好有中度相关性
- 小模型也可用于自动生成高质量干扰项
大型语言模型(LLMs)在教育任务中表现卓越,但其与人类学习模式的一致性,特别是对多选题(MCQs)中错误选项选择倾向的预测能力,仍不明确。本文构建了一个包含真实学生作答分布的多选题数据集,探究两个核心问题:(1)学生更常选的干扰项是否对应于LLM赋予更高生成概率的选项?(2)当LLM出错时,是否会倾向于选择大多数学生误选的选项?实验表明,LLM生成概率与学生选择模式在干扰项上存在中度相关性。此外,当LLM犯错时,其选择的错误答案与多数学生选择的选项高度一致,该现象在小型和大型语言模型中均成立。研究揭示尽管LLM在生成教育内容方面能力强,但其内在推理过程与人类认知在识别混淆干扰项方面仍存在差距。研究结果对教育评估设计具有重要启示:小型语言模型可高效用于自动干扰项生成,因其与大模型在识别易错选项上表现出相似模式。这种与学生误解模式的自然对齐,为生成高质量干扰项提供了新可能。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in various educational tasks, yet their alignment with human learning patterns, particularly in predicting which incorrect options students are most likely to select in multiple-choice questions (MCQs), remains underexplored. Our work investigates the relationship between LLM generation likelihood and student response distributions in MCQs with a specific focus on distractor selections. We collect a comprehensive dataset of MCQs with real-world student response distributions to explore two fundamental research questions: (1). RQ1 - Do the distractors that students more frequently select correspond to those that LLMs assign higher generation likelihood to? (2). RQ2 - When an LLM selects a incorrect choice, does it choose the same distractor that most students pick? Our experiments reveals moderate correlations between LLM-assigned probabilities and student selection patterns for distractors in MCQs. Additionally, when LLMs make mistakes, they are more likley to select the same incorrect answers that commonly mislead students, which is a pattern consistent across both small and large language models. Our work provides empirical evidence that despite LLMs' strong performance on generating educational content, there remains a gap between LLM's underlying reasoning process and human cognitive processes in identifying confusing distractors. Our findings also have significant implications for educational assessment development. The smaller language models could be efficiently utilized for automated distractor generation as they demonstrate similar patterns in identifying confusing answer choices as larger language models. This observed alignment between LLMs and student misconception patterns opens new opportunities for generating high-quality distractors that complement traditional human-designed distractors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。