用人类经验生成对抗性估计,仅5分钟数据提升强化学习采样效率
Search-Based Adversarial Estimates for Improving Sample Efficiency in Off-Policy Reinforcement Learning
- 基于少量人类轨迹进行隐空间相似度搜索,生成对抗性奖励信号
- 在稀疏奖励环境下,算法收敛速度显著快于原始版本
- 适用于极稀疏奖励场景,适合研究高效强化学习的学者
深度强化学习中的样本效率低仍是长期挑战,尤其在奖励稀疏或延迟的环境中。本文提出一种新的、简单且高效的对抗性估计方法,用于改进一类基于反馈的深度强化学习算法。该方法利用仅5分钟的人类记录轨迹,通过从少量人类轨迹中进行潜在空间相似性搜索来增强学习效果。实验表明,采用对抗性估计训练的算法比原始版本收敛更快。此外,我们探讨了该方法如何使基于反馈的算法在极端稀疏奖励场景下仍可实现有效学习。
原文摘要 · Abstract (English)
Sample inefficiency is a long-lasting challenge in deep reinforcement learning (DRL). Despite dramatic improvements have been made, the problem is far from being solved and is especially challenging in environments with sparse or delayed rewards. In our work, we propose to use Adversarial Estimates as a new, simple and efficient approach to mitigate this problem for a class of feedback-based DRL algorithms. Our approach leverages latent similarity search from a small set of human-collected trajectories to boost learning, using only five minutes of human-recorded experience. The results of our study show algorithms trained with Adversarial Estimates converge faster than their original version. Moreover, we discuss how our approach could enable learning in feedback-based algorithms in extreme scenarios with very sparse rewards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。