arXiv:2605.02435cs.LGstat.ML2026-05被引 7

提出新方法解决答案级微调中的系统性偏差问题,提升训练稳定性和效率。

Generalized Distributional Alignment Games for Unbiased Answer-Level Fine-Tuning

  • 用U统计量构建无偏估计器,针对多项式奖励的广义散度对齐游戏
  • 在标准KL散度游戏中设计最优极小极大多项式估计器,达Θ(1/K²)统计误差下限
  • 融合两种方法得新型优化程序,零额外计算开销,加速收敛并减少方差

分布对齐博弈框架为答案级微调(ALFT)提供了有力的变分视角。然而,现有算法依赖小批次对数奖励估计,因詹森不等式引入系统性偏差,可能导致训练不稳定。本文系统性地解决了这一结构性估计偏差:首先,将对齐游戏推广至任意Bregman散度,证明对于诱导多项式奖励的一类几何结构,可利用U统计量构造出无偏且精确的估计器;其次,针对标准KL散度游戏(精确解不可行),推导出全局鲁棒的极小极大多项式估计器,其达到理论最优,实现$Θ(1/K^2)$的统计误差下限,该结论通过Ditzian-Totik定理建立;最后,融合上述两法,提出新型方差最优增强型多项式优化程序(AQP)估计器,证明通过系统性降低方差,不仅实现最优偏差,还获得可证明的加速博弈收敛,带来更高效稳定的训练,且无需在线计算开销。

原文摘要 · Abstract (English)

The Distributional Alignment Game framework provides a powerful variational perspective on Answer-Level Fine-Tuning (ALFT). However, standard algorithms for these games rely on estimating logarithmic rewards from small batches, introducing a systematic bias due to Jensen's inequality that can destabilize training. In this paper, we systematically resolve this structural estimation bias. First, we generalize the alignment game to arbitrary Bregman divergences, showing that for a family of geometries inducing polynomial rewards, we can construct provably exact and unbiased estimators using U-statistics. Second, for the canonical KL divergence game where an exact solution is impossible, we derive a globally robust minimax polynomial estimator that is provably optimal, achieving the fundamental statistical error limit of $Θ(1/K^2)$, which we establish via the Ditzian-Totik theorem. Finally, we synthesize these two approaches to propose a novel Variance-Optimal Augmented Polynomial Optimization Program (AQP) Estimator, proving that by systematically reducing variance, our method achieves not only optimal bias but also provably accelerated game convergence, leading to more efficient and stable training with zero online computational overhead.

微调优化分布对齐无偏估计统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。