arXiv:2412.21001cs.LGcs.AI2024-12中稿 · IEEE Transactions …被引 1

用预训练模型生成偏好数据,高效实现离线强化学习。

LEASE: Offline Preference-based Reinforcement Learning with High Sample Efficiency

  • 用学习的转移模型生成无标签偏好数据,减少人工标注依赖。
  • 在仅需少量偏好数据情况下,性能媲美在线训练基线。
  • 引入置信度机制和理论保障,适合数据稀缺场景使用。

离线偏好强化学习(PbRL)可有效解决奖励设计困难和在线交互成本高的问题。然而,偏好标注需要实时人工反馈,获取足够标注数据仍具挑战。为此,本文提出高样本效率的离线偏好强化学习算法 LEASE,利用学习的转移模型生成未标注偏好数据。考虑到预训练奖励模型可能对未标注数据产生错误标签,我们设计了不确定性感知机制,仅选择高置信度、低方差的数据进行训练。此外,我们给出了奖励模型的泛化界,分析影响奖励准确性的因素,并证明了 LEASE 学习到的策略具有理论上的性能提升保证。该理论基于状态-动作对构建,可轻松与其它离线算法结合。实验表明,LEASE 在无需在线交互的情况下,仅用较少偏好数据即可达到与基线相当的性能。

原文摘要 · Abstract (English)

Offline preference-based reinforcement learning (PbRL) provides an effective way to overcome the challenges of designing reward and the high costs of online interaction. However, since labeling preference needs real-time human feedback, acquiring sufficient preference labels is challenging. To solve this, this paper proposes a offLine prEference-bAsed RL with high Sample Efficiency (LEASE) algorithm, where a learned transition model is leveraged to generate unlabeled preference data. Considering the pretrained reward model may generate incorrect labels for unlabeled data, we design an uncertainty-aware mechanism to ensure the performance of reward model, where only high confidence and low variance data are selected. Moreover, we provide the generalization bound of reward model to analyze the factors influencing reward accuracy, and demonstrate that the policy learned by LEASE has theoretical improvement guarantee. The developed theory is based on state-action pair, which can be easily combined with other offline algorithms. The experimental results show that LEASE can achieve comparable performance to baseline under fewer preference data without online interaction.

强化学习离线学习偏好学习样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。