偏好数据限制强化学习优化效果,难以获得接近最优解。
The Limits of Preference Data for Post-Training
- 用投票理论证明偏好数据无法保证逼近最优解
- 即使理想偏好数据也无法克服根本性局限
- 适合关注人类反馈强化学习瓶颈的研究者
近期大语言模型能力的提升主要依赖于可自动验证结果的强化学习。然而,在需要人类反馈的领域(如深度研究、行程规划),结果评估是定性的,存在多种成功程度。偏好数据(成对或k-wise排序)是一种可扩展的人类反馈方式,但本文揭示其本质局限:即便拥有无限、无噪声、在线的理想偏好数据,基于序数反馈的强化学习仍无法获得近似最优解。通过类比模型响应与选民投票,我们形式化证明了这一不可能性。研究还发现,该局限主要抑制了强化学习后训练在推理行为(如回溯)上的表现,而对指令微调和安全训练影响较小,表明偏好数据限制了鲁棒策略的生成,而这类策略涵盖多数推理行为。
原文摘要 · Abstract (English)
Recent progress in strengthening the capabilities of large language models has stemmed from applying reinforcement learning to domains with automatically verifiable outcomes. A key question is whether we can similarly use RL to optimize for outcomes in domains where evaluating outcomes inherently requires human feedback; for example, in tasks like deep research and trip planning, outcome evaluation is qualitative and there are many possible degrees of success. One attractive and scalable modality for collecting human feedback is preference data: ordinal rankings (pairwise or $k$-wise) that indicate, for $k$ given outcomes, which one is preferred. In this work, we study a critical roadblock: preference data fundamentally and significantly limits outcome-based optimization. Even with idealized preference data (infinite, noiseless, and online), the use of ordinal feedback can prevent obtaining even approximately optimal solutions. We formalize this impossibility using voting theory, drawing an analogy between how a model chooses to answer a query with how voters choose a candidate to elect. This indicates that grounded human scoring and algorithmic innovations are necessary for extending the success of RL post-training to domains demanding human feedback. We also explore why these limitations have disproportionately impacted RLHF when it comes to eliciting reasoning behaviors (e.g., backtracking) versus situations where RLHF has been historically successful (e.g., instruction-tuning and safety training), finding that the limitations of preference data primarily suppress RLHF's ability to elicit robust strategies -- a class that encompasses most reasoning behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。