arXiv:2606.31524cs.LGcs.AI2026-06中稿 · UAI 2026被引 1

改进大模型对齐的自进化算法,提升收敛性与稳定性。

On the Convergence of Self-Improving Online LLM Alignment

论文配图:On the Convergence of Self-Improving Online LLM Alignment
图 1 · 摘自论文原文
  • 引入逆KL散度正则化,改善优化空间的凸性。
  • 证明新目标函数满足PL条件,实现近线性样本复杂度。
  • 在MuJoCo和大模型对齐任务中表现优于原始SAIL。

Self-Improving Alignment (SAIL) 算法通过将问题从双层优化简化为单层方法来应对分布偏移,实证表现良好。然而其收敛性质缺乏理论分析。我们发现标准SAIL目标函数因海森矩阵特性不保证强凹性。为此提出正则化目标 SAIL-RevKL,引入逆Kullback-Leibler (KL) 散度惩罚以优化优化景观。核心理论贡献是证明该正则化目标在有界参数空间内满足Polyak-Lojasiewicz (PL) 条件,建立全局收敛性,实现近线性样本复杂度。通过实证验证,SAIL-RevKL 在 MuJoCo 基准和大模型对齐任务上均优于原始SAIL,展现更强有效性和稳定性。

原文摘要 · Abstract (English)

The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient, single-level method. Empirically, SAIL has demonstrated strong performance on this task. However, a formal analysis of its convergence properties has been lacking. We identify a key theoretical challenge: the standard SAIL objective function is not guaranteed to be strongly concave due to unfavorable properties of its Hessian. To address this limitation, we propose a regularized objective, SAIL-RevKL, which incorporates a reverse Kullback-Leibler (KL) divergence penalty to improve the optimization landscape. Our central theoretical contribution is to prove that this regularized objective satisfies the Polyak-Lojasiewicz (PL) condition within a bounded parameter space. We establish global convergence guarantees, achieving a near-linear sample complexity. We further validate the effectiveness and stability of SAIL-RevKL through empirical evaluations, demonstrating that it outperforms the vanilla SAIL on both MuJoCo benchmarks and LLM alignment tasks.

大模型对齐优化理论收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。