自动挑选最优奖励函数,让强化学习更快更省力。
ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
- 把奖励设计当成在线选择问题,自动挑出好奖励。
- 比之前方法节省8倍计算时间,数据效率提升超50%。
- 无需人工干预,效果接近专家设计的奖励函数。
奖励塑造在强化学习中至关重要,尤其在奖励稀疏的复杂任务中。然而,从一组奖励函数中高效选择有效塑造奖励仍是一个开放挑战。我们提出在线奖励选择与策略优化(ORSO),将塑造奖励函数的选择建模为在线模型选择问题。ORSO无需人工干预即可自动识别性能优异的塑造奖励函数,并具备可证明的后悔边界。我们在多个连续控制任务上验证了其有效性:相比先前方法,ORSO显著减少了评估塑造奖励函数所需的数据量,实现了更高的数据效率和计算时间降低(最高达8倍)。ORSO始终能识别出优于以往方法50%以上的高质量奖励函数,并平均达到与领域专家手动设计奖励函数所训练策略相当的性能。
原文摘要 · Abstract (English)
Reward shaping is critical in reinforcement learning (RL), particularly for complex tasks where sparse rewards can hinder learning. However, choosing effective shaping rewards from a set of reward functions in a computationally efficient manner remains an open challenge. We propose Online Reward Selection and Policy Optimization (ORSO), a novel approach that frames the selection of shaping reward function as an online model selection problem. ORSO automatically identifies performant shaping reward functions without human intervention with provable regret guarantees. We demonstrate ORSO's effectiveness across various continuous control tasks. Compared to prior approaches, ORSO significantly reduces the amount of data required to evaluate a shaping reward function, resulting in superior data efficiency and a significant reduction in computational time (up to 8 times). ORSO consistently identifies high-quality reward functions outperforming prior methods by more than 50% and on average identifies policies as performant as the ones learned using manually engineered reward functions by domain experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。