arXiv:2503.19523cs.LGcs.CV2025-03被引 3

统一强化学习与无强化学习的智能体对齐方法

One Framework to Rule Them All: Unifying RL-Based and RL-Free Methods in RLHF

  • 从神经结构化赌博机视角重构对齐算法
  • 证明标准RLHF目标等价于结构化赌博机预测
  • 提出GRO框架,融合有/无强化学习方法

本文系统分析了用于人类反馈强化学习(RLHF)和大推理模型(LRMs)的各种基于强化学习与非强化学习的方法。首先概述了典型的RLHF与LRMs流程,随后从神经结构化赌博机预测的角度重新诠释多种算法,揭示了看似不同的方法之间的深层关联。接着回顾强化学习核心原则,指出现有研究中常被忽视的方面,并在完整强化学习框架下推导出标准RLHF目标,证明其与神经结构化赌博机预测等价。最后,通过重新审视近端策略优化(PPO)的原理,识别出需调整的关键环节,提出广义强化优化(GRO)框架,实现了基于强化学习与非强化学习方法在RLHF中的无缝融合。期待社区对GRO进行实证验证并提出建设性反馈。

原文摘要 · Abstract (English)

In this article, we primarily examine a variety of RL-based and RL-free methods designed to address Reinforcement Learning from Human Feedback (RLHF) and Large Reasoning Models (LRMs). We begin with a concise overview of the typical steps involved in RLHF and LRMs. Next, we reinterpret several RL-based and RL-free algorithms through the perspective of neural structured bandit prediction, providing a clear conceptual framework that uncovers a deeper connection between these seemingly distinct approaches. Following this, we briefly review some core principles of reinforcement learning, drawing attention to an often-overlooked aspect in existing RLHF studies. This leads to a detailed derivation of the standard RLHF objective within a full RL context, demonstrating its equivalence to neural structured bandit prediction. Finally, by reinvestigating the principles behind Proximal Policy Optimization (PPO), we pinpoint areas needing adjustment, which culminates in the introduction of the Generalized Reinforce Optimization (GRO) framework, seamlessly integrating RL-based and RL-free methods in RLHF. We look forward to the community's efforts to empirically validate GRO and invite constructive feedback.

强化学习对齐方法大模型优化框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。