arXiv:2503.00930cs.LGcs.AI2025-03被引 1

用偏好比较提升离线强化学习策略稳定性与性能

Behavior Preference Regression for Offline Reinforcement Learning

  • 基于配对样本比较重构策略优化目标,融合行为一致性约束
  • 在D4RL和V-D4RL上均达到当前最优性能,图像状态任务表现突出
  • 适合追求高稳定性和泛化能力的离线强化学习研究者

离线强化学习旨在仅使用固定数据集轨迹学习最优策略。策略约束方法将策略学习建模为最大化奖励与最小化偏离行为策略之间的权衡优化问题。该问题的闭式解可表示为加权行为克隆目标,理论上需计算难以处理的归一化常数。强化学习已用于语言建模中对齐模型与人类偏好:近期工作通过偏好模型对成对生成结果排序,并直接提高优选结果的似然。本文借鉴此思想,将配对样本优化问题重新表述,同时拟合Q函数的最大模态并最大化策略动作的行为一致性。由此提出离线强化学习中的行为偏好回归算法(BPR)。我们在广泛使用的D4RL运动与Antmaze数据集,以及更具挑战性的基于图像状态的V-D4RL套件上进行实验评估。BPR在所有领域均表现出当前最优性能。有监督实验表明,BPR能有效利用在线策略价值函数的稳定性,在运动类数据集上性能下降可忽略。

原文摘要 · Abstract (English)

Offline reinforcement learning (RL) methods aim to learn optimal policies with access only to trajectories in a fixed dataset. Policy constraint methods formulate policy learning as an optimization problem that balances maximizing reward with minimizing deviation from the behavior policy. Closed form solutions to this problem can be derived as weighted behavioral cloning objectives that, in theory, must compute an intractable partition function. Reinforcement learning has gained popularity in language modeling to align models with human preferences; some recent works consider paired completions that are ranked by a preference model following which the likelihood of the preferred completion is directly increased. We adapt this approach of paired comparison. By reformulating the paired-sample optimization problem, we fit the maximum-mode of the Q function while maximizing behavioral consistency of policy actions. This yields our algorithm, Behavior Preference Regression for offline RL (BPR). We empirically evaluate BPR on the widely used D4RL Locomotion and Antmaze datasets, as well as the more challenging V-D4RL suite, which operates in image-based state spaces. BPR demonstrates state-of-the-art performance over all domains. Our on-policy experiments suggest that BPR takes advantage of the stability of on-policy value functions with minimal perceptible performance degradation on Locomotion datasets.

强化学习离线学习偏好学习策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。