arXiv:2604.19024cs.LG2026-04被引 3

无需训练奖励模型,实现安全强化学习的全局收敛。

Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback

  • 基于原始-对偶方法,直接优化策略
  • 支持灵活轨迹长度,实现多项式收敛率
  • 首个在人类反馈下证明非渐近全局收敛的工作

从人类反馈中进行安全强化学习(Safe RLHF)近期在构建有益且无害的大语言模型方面取得实证成功,通过解耦人类对有益性和无害性的偏好。现有方法通常依赖于从人类反馈中拟合固定时域的奖励模型,且仅在经验上得到验证。本文将安全RLHF建模为无限时域折扣约束马尔可夫决策过程(CMDP),因为人类可能在持续交互序列中与模型互动,而非单个有限阶段内。我们提出两种无需奖励模型拟合的Safe RLHF算法,与以往假设固定长度轨迹的工作不同,支持灵活轨迹长度训练。两个算法均基于原始-对偶方法,在策略梯度迭代次数、轨迹样本长度和人类偏好查询次数上均实现全局收敛,且收敛速率为多项式。据我们所知,这是首次在人类反馈下研究无限时域折扣CMDP并建立全局、非渐近收敛性的工作。

原文摘要 · Abstract (English)

Safe Reinforcement Learning from Human Feedback (Safe RLHF) has recently achieved empirical success in developing helpful and harmless large language models by decoupling human preferences regarding helpfulness and harmlessness. Existing approaches typically rely on fitting fixed horizon reward models from human feedback and have only been validated empirically. In this paper, we formulate safe RLHF as an infinite horizon discounted Con- strained Markov Decision Process (CMDP), since humans may interact with the model over a continuing sequence of interactions rather than within a single finite episode. We propose two Safe RLHF algorithms that do not require reward model fitting and, in contrast to prior work assuming fixed-length trajectories, support flexible trajectory lengths for training. Both algo- rithms are based on the primal-dual method and achieve global convergence guarantees with polynomial rates in terms of policy gradient iterations, trajectory sample lengths, and human preference queries. To the best of our knowledge, this is the first work to study infinite horizon discounted CMDP under human feedback and establish global, non-asymptotic convergence.

强化学习安全学习人类反馈收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。