为非可验证任务设计了基于推理链的偏好建模优化方法
Bradley-Terry Policy Optimization for Generative Preference Modeling
- 将推理链视为隐变量,重构偏好似然函数结构
- 提出蒙特卡洛梯度估计器,实现稳定训练
- 在多基准测试中优于现有启发式方法
强化学习(RL)在具有可验证答案的任务中已证明能有效扩展大语言模型的思维链(CoT)推理。然而,将基于RL的思维训练推广到更广泛的不可验证任务——仅通过成对人类偏好提供监督——仍具挑战性。现有方法通常以启发式方式将为可验证奖励设计的RL目标应用于偏好设置。本文表明,在偏好建模中引入推理链会从根本上改变Bradley-Terry(BT)似然结构,因推理过程必须作为隐变量处理。这导致偏好似然表达为随机生成轨迹期望的比值,无法使用Jensen型界或标准RL目标进行优化。为此,我们推导出该似然梯度的一致蒙特卡洛估计器,提出布拉德利-特里政策优化(BTPO)。实验表明,BTPO能够稳定有效地训练带有推理链的生成式偏好模型,在多个基准和模型规模下持续优于先前启发式方法。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has recently proven effective at scaling chain-of-thought (CoT) reasoning in large language models for tasks with verifiable answers. However, extending RL-based thought training to more general non-verifiable tasks-where supervision is provided only through pairwise human preferences-remains challenging. Existing approaches typically apply RL objectives designed for verifiable rewards to preference-based settings in a heuristic manner. In this work, we show that introducing CoT reasoning into preference modeling fundamentally changes the structure of the Bradley-Terry (BT) likelihood, as the reasoning process must be treated as a latent variable. This results in a preference likelihood expressed as a ratio of expectations over stochastic generation trajectories, which cannot be optimized using Jensen-style bounds or standard RL objectives. To address this challenge, we derive a consistent Monte Carlo estimator for the gradient of the resulting likelihood, leading to Bradley-Terry Policy Optimization (BTPO). Empirically, BTPO enables stable and effective training of generative preference models with CoT reasoning, consistently outperforming prior heuristic approaches across multiple benchmarks and model scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。