用低成本探索替代高开销挖掘,提升数学推理对齐效率
Mining or Synthesis? Rethinking Exploration Efficiency in Iterative Alignment of Mathematical Reasoning
- 用小规模采样(2≤N≤3)合成高质量偏好对,避免盲目搜寻
- 在1/5计算量下达到DPO-R1(N=16)性能,且抗标签噪声更强
- 适合追求高效训练与鲁棒性的数学推理模型优化者
迭代直接偏好优化(DPO)已成为对齐大语言模型进行推理任务的主流范式。现有方法通常依赖最佳-N采样(N≥8)从分布尾部挖掘正向轨迹。本文发现,在数学推理中,增大N带来边际收益递减,同时增加验证器引发的误报风险及策略更新所需分布偏移。为此,我们提出PACE(基于校正探索的近端对齐),一种生成式校正框架,将耗时的挖掘替换为低预算探索(2≤N≤3)。PACE不寻找越来越罕见的正样本,而是通过校正性回溯优化和验证引导过滤,从失败探索中合成高保真偏好对。实验表明,PACE在约1/5计算量下达到或超越DPO-R1(N=16)的性能,并在20%标签被污染情况下仍保持鲁棒,而高N基线显著暴露于噪声利用。
原文摘要 · Abstract (English)
Iterative Direct Preference Optimization (DPO) has emerged as a widely used paradigm for aligning Large Language Models on reasoning tasks. Existing approaches typically rely on Best-of-N sampling ($N\geq8$) to mine positive trajectories from the distribution tail. In this work, we show that in mathematical reasoning, increasing $N$ yields diminishing returns while increasing verifier-induced false-positive risk and the distribution shift required for policy updates. To address this, we introduce PACE (Proximal Alignment via Corrective Exploration), a generation-based corrective framework that replaces exhaustive mining with low-budget exploration ($2\leq N\leq3$). Rather than searching for increasingly rare positive samples, PACE synthesizes high-fidelity preference pairs from failed explorations through corrective hindsight refinement and verification-guided filtering. Empirically, PACE matches or exceeds the performance of DPO-R1 ($N=16$) while using about $1/5$ of the compute, and remains robust under 20\% label corruption, where high-$N$ baselines exhibit substantially higher noise exploitation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。