为数学错题纠正AI导师设计可评估的奖励模型
Towards Reward Modeling for AI Tutors in Math Mistake Remediation
- 基于人类偏好构建教学层次结构,生成对比性答题对
- 仅用合成数据训练出0.69准确率的奖励模型,融合数据后达0.74
- 小模型胜过大模型,适合教育类AI系统优化
评估AI导师的教学质量仍具挑战:标准自然语言生成指标无法判断回答是否识别错误、引导思路或避免直接给答案。针对错题纠正任务,我们从人类对MRBench数据集的成对偏好中提炼出教学层面的层次结构,并合成在关键方面(如错误识别与定位、针对性、引导性、可操作性、清晰度和连贯性)存在最小差异的响应对。我们开发并发布了基于加权求和排名的布拉德利-特里偏好模型,这些排名由MRBench、合成数据及数据组合自动生成。仅使用合成数据,最佳模型在人类偏好测试中达到0.69的成对准确率;结合加权求和数据与特定合成组后,准确率提升至0.74,优于更大规模的通用奖励模型,且仅采用0.5B参数的骨干模型。
原文摘要 · Abstract (English)
Evaluating the pedagogical quality of AI tutors remains challenging: standard NLG metrics do not determine whether responses identify mistakes, scaffold reasoning, or avoid revealing the answers. For the task of mistake remediation, we derive a hierarchy of pedagogical aspects from human pairwise preferences on MRBench, and synthesize minimally contrastive response pairs that differ along key aspects (e.g., mistake identification and location, targetedness, scaffolding, actionability, clarity, and coherence). We develop and release Bradley-Terry preference models trained on weighted-sum rankings that we automatically create from MRBench, synthetic pairs, and data combinations. Using only synthetic data, our best model reaches 0.69 pairwise accuracy on a human preference test, and combining weighted-sum data with targeted synthetic groups improves accuracy to 0.74, outperforming larger general-purpose reward models while using only a 0.5B-parameter backbone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。