通过回溯重标注提升多目标强化学习的偏好数据利用效率
Hindsight Preference Replay Improves Preference-Conditioned Multi-Objective Reinforcement Learning
- 用回溯重标注策略为历史数据赋予其他偏好标签,扩展监督信号
- 在六个任务中五项提升超体积指标,一项任务期望效用提升达5倍以上
- 无需改动原有模型,适合希望高效利用偏好数据的研究者
多目标强化学习(MORL)使智能体能在满足用户偏好条件下优化向量奖励。CAPQL是一种基于权重向量w的条件化演员-评论家方法,但仅使用特定偏好下收集的数据,导致其他偏好的离线数据被闲置。本文提出一种简单通用的重放增强策略——事后偏好重放(HPR),通过回溯重标注将存储的转移数据赋予替代偏好标签,从而在不改变CAPQL架构和损失函数的前提下,密化偏好单纯形上的监督信号。在六项MO-Gymnasium运动任务上,固定30万步预算下使用期望效用(EUM)、超体积(HV)和稀疏性评估,HPR-CAPQL在五个环境中提升超体积,在四个环境中提升期望效用。以mo-humanoid-v5为例,期望效用从$323\±125$升至$1613\u00b1464$,超体积从0.52M增至9.63M,统计显著。mo-halfcheetah-v5仍是挑战,此处CAPQL在相近期望效用下获得更高超体积。报告所有任务的最终结果与帕累托前沿可视化。
原文摘要 · Abstract (English)
Multi-objective reinforcement learning (MORL) enables agents to optimize vector-valued rewards while respecting user preferences. CAPQL, a preference-conditioned actor-critic method, achieves this by conditioning on weight vectors w and restricts data usage to the specific preferences under which it was collected, leaving off-policy data from other preferences unused. We introduce Hindsight Preference Replay (HPR), a simple and general replay augmentation strategy that retroactively relabels stored transitions with alternative preferences. This densifies supervision across the preference simplex without altering the CAPQL architecture or loss functions. Evaluated on six MO-Gymnasium locomotion tasks at a fixed 300000-step budget using expected utility (EUM), hypervolume (HV), and sparsity, HPR-CAPQL improves HV in five of six environments and EUM in four of six. On mo-humanoid-v5, for instance, EUM rises from $323\!\pm\!125$ to $1613\!\pm\!464$ and HV from 0.52M to 9.63M, with strong statistical support. mo-halfcheetah-v5 remains a challenging exception where CAPQL attains higher HV at comparable EUM. We report final summaries and Pareto-front visualizations across all tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。