用强化学习让AI老师更会教数学,不直接给答案。
Towards Pedagogically Aligned LLM Tutors for Math Mistake Remediation

- 分两阶段对齐:先微调对话数据,再用偏好优化提升教学策略。
- 相比基础模型,答案更准确且教学过程更符合教育规律。
- 适合想开发透明可复现的智能辅导系统的研究人员。
大语言模型在智能辅导系统中潜力巨大,但常缺乏有效的教学策略,如引导学生而不直接给出答案。本文提出一种两阶段对齐流程,结合监督微调与直接偏好优化,用于数学错题纠正。构建了一个融合现有辅导语料和基于教学维度(如支架式支持、事实性)生成的合成数据的数据集,并研究了包含解题正确性和标准答案的不同输入配置。实验表明,该方法在事实准确性与教学质量上均优于基线模型及现有辅导模型。人工评估显示,最佳模型表现媲美强效专有基线,同时具备开放性、透明性和可复现性优势。结果验证了基于偏好的教学对齐有效性,也揭示了辅导质量可靠评估的挑战。
原文摘要 · Abstract (English)
Large language models have strong potential for use in intelligent tutoring systems, but they often fail to follow effective pedagogical strategies, such as guiding students without revealing final answers. We study the application of a two-stage alignment pipeline for math mistake remediation, combining supervised fine-tuning on tutoring dialogs with Direct Preference Optimization on synthetic preference pairs. We construct a dataset that integrates existing tutoring corpora with synthetic data generated along pedagogical dimensions, such as scaffolding and factuality, and study different input configurations that incorporate solution correctness and gold answers. Experiments show that this approach improves both factual accuracy and pedagogical quality over base models and existing tutoring models. Human evaluation further indicates that our best model is competitive with a strong proprietary baseline, while providing additional benefits in terms of openness, transparency, and reproducibility. Our results highlight the effectiveness of preference-based pedagogical alignment, while also revealing challenges in reliably evaluating tutoring quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。