arXiv:2608.22806cs.CL2026-08中稿 · EMNLP

DIAG通过动态调整题目分布,提升数学推理模型训练效率。

DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation

论文配图:DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation
图 1 · 摘自论文原文
  • 根据学生错误痕迹生成针对性练习题
  • 自适应分配题目难度,提升有效偏好对产量
  • 适合数据稀缺下的数学推理模型优化

迭代偏好优化对对齐大语言模型的数学推理能力至关重要,但常因信号稀疏而效率低下:随着模型进步,静态题集逐渐与模型能力不匹配,导致生成结果过易或过难,无法提供有效偏好对。本文提出DIAG框架,通过自适应调整练习分布,聚焦于学生当前能力边界,以增加有信息量的监督。DIAG包含两阶段:(1) 利用经验贝叶斯收缩估计器诊断有效偏好对产出,校准探索-利用权衡并分配主题配额,优先高产出概念;(2) 教师基于学生错误轨迹生成针对性变体题。理论分析表明,DIAG可视为教师引导的KL正则化重加权近似,使练习分布逼近学生能力边界,最大化有效偏好对产量。实验显示,DIAG在各迭代中提升偏好对产量,并在等效训练预算下实现更强推理性能,证明其能更高效地提炼数学推理的有信息量偏好监督。

原文摘要 · Abstract (English)

Iterative preference optimization is essential for aligning Large Language Models on mathematical reasoning tasks, yet its efficiency is often throttled by signal scarcity: as the model improves, static problem sets become increasingly mismatched to the model's evolving competence, producing rollouts that are either too easy or too hard and therefore non-informative, which leads to a scarcity of valid preference pairs. We propose DIAG, a Diagnostic Iterative Alignment and Generation framework that adaptively reshapes the practice distribution to increase informative supervision and focus training near the student's current competence boundary. DIAG consists of two phases: (1) diagnosing valid preference-pair yield to calibrate the exploration-exploitation trade-off and allocate topic quotas via an Empirical Bayes shrinkage estimator, thereby prioritizing high-yield concepts; and (2) generating targeted practice, where a teacher synthesizes variants from the student's failure traces. We further provide a theoretical view interpreting DIAG as a teacher-mediated approximation to KL-regularized reweighting of the practice distribution toward the student's competence boundary, where valid preference-pair yield is maximized. Experiments show that DIAG boosts yield across iterations and delivers stronger reasoning performance under an iso-effective training budget, demonstrating that it can distill more informative preference supervision for mathematical reasoning.

数学推理偏好学习数据高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。