arXiv:2607.29246cs.AI2026-07被引 1

通过分解策略提升多奖励强化学习的稳定性与可控性

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

  • 将策略分解为正向独立策略和全局负向策略,避免奖励冲突
  • 在科学推理、工具使用等任务中表现优于现有基线
  • 支持推理时灵活组合策略,实现偏好控制

现代大语言模型不仅需正确回答,还需适应不同人类价值观与应用场景。因此,多奖励强化学习(multi-reward RL)成为关键挑战,其中每个奖励代表一种期望行为维度。然而,多奖励优化易引发对齐税问题,不同目标间可能相互抵触,导致训练不稳定、效率低下。本文提出PRISM框架,基于策略空间分解与组合思想,不直接融合奖励,而是分别优化一组独立正向策略与一个全局负向策略。该方法缓解了多奖励优化中的潜在冲突,同时在推理阶段通过灵活策略组合实现可控性。在科学推理、工具使用推理及有用性-安全性对齐任务上的实验表明,PRISM持续优于现有多奖励RL基线,并具备推理时偏好调控能力。

原文摘要 · Abstract (English)

Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinforcement learning (RL) has become an increasingly important problem for LLMs, where each reward captures a different aspect of desired behavior. However, optimizing with multiple rewards suffers from a more severe alignment tax issue, where different optimization objectives can trade off or even conflict with each other, leading to unstable and inefficient post-training. In this work, we propose PRISM, a new multi-reward RL framework built upon the idea of policy-space decomposition and composition. Instead of compositing different rewards, PRISM optimizes a set of standalone positive policies and a global negative policy. This alleviates the potential conflict during multi-reward policy optimization, while enabling controllability during inference by flexible policy composition. Experiments on scientific reasoning, tool-use reasoning, and helpfulness-safety alignment show that PRISM consistently outperforms existing multi-reward RL baselines, with extra controllability for inference-time preference control.

强化学习多奖励策略分解大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。