解决多奖励强化学习中奖励相关性导致的训练不稳问题
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
- 用分位数归一化处理不同类型的奖励,稳定优势值分配
- 在每个活跃奖励子空间内进行马哈拉诺比去相关,减少冗余
- 提升指令遵循与复杂提示下的鲁棒性,适合长文本生成任务
复杂强化学习环境常采用多任务与混合奖励设置。此类场景下,异构奖励分布与相关奖励维度常导致标量优势值构建不稳定。为此,我们提出奖励去相关策略优化(RDPO),专门应对上述两种失效模式。RDPO首先使用幅度感知分位数归一化,在二元、分数与连续奖励间稳定提示级优势分配;随后在每个活跃奖励子空间内应用马哈拉诺比去相关,以减轻聚合前的相关冗余。在LongCat-Flash模型后训练阶段应用该方法,显著提升指令遵循能力、写作质量及对困难提示的鲁棒性,同时在推理与编码评估中保持广泛竞争力。
原文摘要 · Abstract (English)
Complex reinforcement learning environments frequently employ multi-task and mixed-reward formulations. In these settings, heterogeneous reward distributions and correlated reward dimensions often destabilize the construction of scalar advantages. To address these challenges, we propose Reward-Decorrelated Policy Optimization (RDPO), a reward-processing method designed to explicitly target both failure modes. RDPO first utilizes Magnitude-Aware Quantile normalization to stabilize prompt-level advantage allocation across binary, fractional, and continuous rewards. It then applies Mahalanobis whitening within each active reward subspace to mitigate correlation redundancy prior to aggregation. When applied during the post-training of LongCat-Flash, RDPO enhances instruction following, writing quality, and robustness to hard prompts while remaining broadly competitive on reasoning and coding evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。