提出多目标奖励优化框架,实现安全强化学习与大模型对齐的理论突破。
Multi-Objective Reward and Preference Optimization: Theory and Algorithms
- 基于敏感度分析与信任域更新,稳定处理约束强化学习问题
- 在平均成本与有限时长远期任务中均达到最优性能并有理论保证
- 适合关注安全决策、大模型对齐及偏好学习的研究者
本论文构建了约束强化学习(RL)的理论框架与算法体系,涵盖控制、偏好学习与大语言模型对齐。首项贡献提出平均成本准则下的平均约束策略优化(ACPO),结合灵敏度分析与信任域更新,实现稳定约束处理,兼具理论保证与领先实证表现。其次,针对有限时长远期任务,提出e-COP,首个适用于周期性约束马尔可夫决策过程(episodic CMDPs)的策略优化方法,基于周期性策略差异引理,具备可证明性能、简洁性与可扩展性,适用于高安全性场景。论文进一步研究从人类偏好中学习:warmPref-PS引入后验采样策略,将异质评分者的离线偏好数据融入在线学习,显式建模评分者能力,显著降低遗憾并提升数据效率。PSPL算法通过成对轨迹比较联合采样奖励模型与转移动态,提供贝叶斯简单遗憾保证,稳健识别最优策略。最后,将上述方法应用于大规模模型对齐,提出多目标约束优化视角下的MOPO算法,具备闭式更新,可扩展至数十亿参数的语言模型,在多种对齐设置下保持鲁棒性。整体工作统一了平均成本、周期性与偏好驱动三类约束强化学习范式,推动安全与对齐决策的理论与实践进展。
原文摘要 · Abstract (English)
This thesis develops theoretical frameworks and algorithms that advance constrained reinforcement learning (RL) across control, preference learning, and alignment of large language models. The first contribution addresses constrained Markov Decision Processes (CMDPs) under the average-cost criterion through the Average-Constrained Policy Optimization (ACPO) algorithm. ACPO integrates sensitivity analysis with trust-region updates to ensure stable constraint handling, achieving state-of-the-art empirical performance with theoretical guarantees. Constrained RL is then extended to finite-horizon settings via e-COP, the first policy optimization method for episodic CMDPs. Built on an episodic policy difference lemma, e-COP offers provable performance, simplicity, and scalability in safety-critical environments. The thesis then investigates reinforcement learning from human preferences. warmPref-PS introduces a posterior sampling strategy for linear bandits that integrates offline preference data from heterogeneous raters into online learning. Explicit modeling of rater competence yields substantial regret reduction and more efficient data collection for RLHF. The PSPL algorithm further advances preference-based RL by jointly sampling reward models and transition dynamics from pairwise trajectory comparisons, providing Bayesian simple-regret guarantees and robust empirical identification of optimal policies. The final contribution applies these methods to large-scale model alignment. A multi-objective constrained optimization view yields MOPO, an iterative algorithm with closed-form updates that scales to multi-billion-parameter language models and remains robust across alignment settings. Collectively, the thesis unifies constrained RL across average-cost, episodic, and preference-driven paradigms, delivering theoretical advances and practical tools for safe and aligned decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。