提出新算法实现大模型对齐的稳定收敛,理论更可靠。
Provable Last-Iterate Convergence for Multi-Objective Safe LLM Alignment via Optimistic Primal-Dual
- 用乐观更新机制稳定多目标对齐的优化过程
- 证明最后迭代能收敛,即使在参数化策略下
- 适合关注对齐算法理论保障的研究者
基于人类反馈的强化学习(RLHF)在对齐大语言模型与人类偏好方面起关键作用。传统原始-对偶方法仅在分布空间中凸-凹情形下保证收敛,且在实际参数化策略下可能出现最后迭代不稳定或发散。本文提出一个统一的原始-对偶框架,涵盖安全RLHF、单次和多次方法。在此基础上,设计了包含预测更新的乐观原始-对偶(OPD)算法,稳定鞍点动态。理论上证明该方法在分布空间中可实现最后迭代收敛,并在参数化策略下收敛至与近似误差和偏差相关的邻域解。分析表明,乐观性有效抑制了约束对齐目标固有的振荡,填补了约束强化学习与实际RLHF之间的关键理论空白。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) plays a significant role in aligning Large Language Models (LLMs) with human preferences. While RLHF with expected reward constraints can be formulated as a primal-dual optimization problem, standard primal-dual methods only guarantee convergence with a distributional policy where the saddle-point problem is in convex-concave form. Moreover, standard primal-dual methods may exhibit instability or divergence in the last iterate under policy parameterization in practical applications. In this work, we propose a universal primal-dual framework for safe RLHF that unifies a broad class of existing alignment algorithms, including safe-RLHF, one-shot, and multi-shot based methods. Building on this framework, we introduce an optimistic primal-dual (OPD) algorithm that incorporates predictive updates for both primal and dual variables to stabilize saddle-point dynamics. We establish last-iterate convergence guarantees for the proposed method, covering both exact policy optimization in the distributional space and convergence to a neighborhood of the optimal solution whose gap is related to approximation error and bias under parameterized policies. Our analysis reveals that optimism plays a crucial role in mitigating oscillations inherent to constrained alignment objectives, thereby closing a key theoretical gap between constrained RL and practical RLHF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。