提出新方法让大模型在约束下更优对齐,理论证明可逼近最优解。
Alignment of large language models with constrained learning
- 基于拉格朗日对偶,交替优化模型策略与对偶变量。
- 在PKU-SafeRLHF和Anthropic HH-RLHF数据集上实现高对齐效果。
- 首次理论证明双方法可逼近最优,适用于安全可控生成任务。
我们研究了在约束对齐问题中计算最优大语言模型(LLM)策略的问题,目标是在最大化主奖励的同时满足次级效用的约束。尽管基于拉格朗日的方法在约束对齐中广受欢迎,但迭代式原-对偶方法常无法收敛,而非迭代对偶方法在LLM参数空间中无法达到最优。为解决这些问题,我们利用拉格朗日对偶性,提出一种迭代对偶对齐方法,通过拉格朗日最大化更新LLM策略,通过对偶下降更新对偶变量。理论上,我们刻画了分布空间中的原始值与LLM参数空间中的对偶值之间的原-对偶间隙,并量化了在近似最优对偶变量下,学习到的LLM策略在目标函数和约束函数上的最优性差距。结果证明,对偶方法可找到接近最优的约束型LLM策略,仅受LLM参数化能力限制。我们在PKU-SafeRLHF和Anthropic HH-RLHF数据集上进行了广泛实验,验证了该方法的有效性和优势。
原文摘要 · Abstract (English)
We study the problem of computing an optimal large language model (LLM) policy for the constrained alignment problem, where the goal is to maximize a primary reward objective while satisfying constraints on secondary utilities. Despite the popularity of Lagrangian-based LLM policy search in constrained alignment, iterative primal-dual methods often fail to converge, and non-iterative dual-based methods do not achieve optimality in the LLM parameter space. To address these challenges, we employ Lagrangian duality to develop an iterative dual-based alignment method that alternates between updating the LLM policy via Lagrangian maximization and updating the dual variable via dual descent. In theory, we characterize the primal-dual gap between the primal value in the distribution space and the dual value in the LLM parameter space. We further quantify the optimality gap of the learned LLM policies at near-optimal dual variables with respect to both the objective and the constraint functions. These results prove that dual-based alignment methods can find an optimal constrained LLM policy, up to an LLM parametrization gap. We demonstrate the effectiveness and merits of our approach through extensive experiments conducted on the PKU-SafeRLHF and Anthropic HH-RLHF datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。