用模型自身错误生成更贴近学生错法的数学选择题干扰项。
LookAlike: Consistent Distractor Generation in Math MCQs
- 通过挖掘模型生成不一致处构建偏好数据,自动优化干扰项。
- 在1400+道数学题上,干扰项生成准确率达51.6%,错因生成达57.2%。
- 无需人工标注,适合大规模生成教育场景下的真实错误干扰项。
大型语言模型(LLMs)被越来越多地用于生成多选题(MCQs)的干扰项,尤其是在数学教育领域。然而,现有方法难以确保生成的干扰项与学生常见错误保持一致。我们提出 LookAlike,一种通过偏好优化提升错误-干扰项一致性的方法。主要创新包括:(a) 从模型生成不一致性中挖掘合成偏好对;(b) 交替进行监督微调(SFT)与直接偏好优化(DPO),以稳定训练过程。与依赖启发式或人工标注偏好数据的方法不同,LookAlike 利用自身生成不一致性作为非优选样本,实现可扩展且稳定的训练。在包含1400+道数学多选题的真实数据集上评估,其干扰项生成准确率为51.6%,错误生成准确率为57.2%,优于现有最先进方法(45.6% / 47.7%)。结果表明,基于偏好的正则化与不一致性挖掘对大规模生成一致的数学多选题干扰项具有显著效果。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to generate distractors for multiple-choice questions (MCQs), especially in domains like math education. However, existing approaches are limited in ensuring that the generated distractors are consistent with common student errors. We propose LookAlike, a method that improves error-distractor consistency via preference optimization. Our two main innovations are: (a) mining synthetic preference pairs from model inconsistencies, and (b) alternating supervised fine-tuning (SFT) with Direct Preference Optimization (DPO) to stabilize training. Unlike prior work that relies on heuristics or manually annotated preference data, LookAlike uses its own generation inconsistencies as dispreferred samples, thus enabling scalable and stable training. Evaluated on a real-world dataset of 1,400+ math MCQs, LookAlike achieves 51.6% accuracy in distractor generation and 57.2% in error generation under LLM-as-a-judge evaluation, outperforming an existing state-of-the-art method (45.6% / 47.7%). These improvements highlight the effectiveness of preference-based regularization and inconsistency mining for generating consistent math MCQ distractors at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。