用半监督学习让智能体在稀疏奖励中表现更好
Shaping Sparse Rewards in Reinforcement Learning: A Semi-supervised Approach
- 结合半监督学习与新数据增强,从零奖励轨迹中学习表示
- 在稀疏环境里性能翻倍,最佳得分提升15.8%
- 适合奖励稀少的机器人控制与游戏任务
在许多真实场景中,智能体的奖励信号极度稀疏,导致难以学习有效的奖励函数。本文提出一种新方法,不仅利用非零奖励转移,还通过半监督学习(SSL)与新型双熵数据增强技术,从大量零奖励转移中学习轨迹空间表示,从而提升奖励塑造效果。在Atari和机器人操作任务中的实验表明,该方法在奖励推断上优于监督基线,尤其在更稀疏的奖励环境中,最高得分可达监督方法的两倍。所提双熵数据增强使最佳得分相较其他增强方法提升15.8%。
原文摘要 · Abstract (English)
In many real-world scenarios, reward signal for agents are exceedingly sparse, making it challenging to learn an effective reward function for reward shaping. To address this issue, the proposed approach in this paper performs reward shaping not only by utilizing non-zero-reward transitions but also by employing the \emph{Semi-Supervised Learning} (SSL) technique combined with a novel data augmentation to learn trajectory space representations from the majority of transitions, {i.e}., zero-reward transitions, thereby improving the efficacy of reward shaping. Experimental results in Atari and robotic manipulation demonstrate that our method outperforms supervised-based approaches in reward inference, leading to higher agent scores. Notably, in more sparse-reward environments, our method achieves up to twice the peak scores compared to supervised baselines. The proposed double entropy data augmentation enhances performance, showcasing a 15.8\% increase in best score over other augmentation methods
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。