arXiv:2501.07755cs.LGcs.AI2025-01中稿 · the Collaborative …被引 1

优化基于人类评分的强化学习,提升奖励函数推断效果。

Performance Optimization of Ratings-Based Reinforcement Learning

  • 通过最小化人类评分与估计评分间的交叉熵损失来推断奖励函数。
  • 发现超参数对RbRL性能影响显著,需系统性调参以获得稳定结果。
  • 为用户提供实用超参数选择指南,适合研究人机协同强化学习者。

本文探讨了多种优化方法以提升基于评分的强化学习(RbRL)的性能。RbRL通过人类评分来推断奖励函数,用于无奖励环境下的策略学习,其核心是使人类评分与由推断奖励生成的估计评分之间的一致性最大化。具体而言,该方法最小化衡量两者差异的交叉熵损失,损失越低,一致性越高。尽管形式简单,但RbRL存在多个超参数且对各类因素敏感。因此,本文开展全面实验,分析不同超参数对性能的影响,目前仍为进行中工作,旨在为用户提供在使用RbRL时选择超参数的通用指导原则。

原文摘要 · Abstract (English)

This paper explores multiple optimization methods to improve the performance of rating-based reinforcement learning (RbRL). RbRL, a method based on the idea of human ratings, has been developed to infer reward functions in reward-free environments for the subsequent policy learning via standard reinforcement learning, which requires the availability of reward functions. Specifically, RbRL minimizes the cross entropy loss that quantifies the differences between human ratings and estimated ratings derived from the inferred reward. Hence, a low loss means a high degree of consistency between human ratings and estimated ratings. Despite its simple form, RbRL has various hyperparameters and can be sensitive to various factors. Therefore, it is critical to provide comprehensive experiments to understand the impact of various hyperparameters on the performance of RbRL. This paper is a work in progress, providing users some general guidelines on how to select hyperparameters in RbRL.

强化学习人类评分超参数优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。