arXiv:2410.20487cs.LGcs.AI2024-10IJCAI被引 9

用多样性重放缓冲区,让强化学习在复杂环境中学得更快更好

Efficient Diversity-based Experience Replay for Deep Reinforcement Learning

  • 基于确定性点过程衡量样本多样性,优先回放差异大的经验
  • 在MuJoCo、Atari和Habitat上提升学习效率,尤其在高维空间表现优异
  • 适合需要高效训练的真实世界机器人任务,如复杂环境操控

经验回放广泛用于强化学习中以提升学习效率,但现有方法(无论均匀或优先采样)在高维状态空间的真实场景中常效率低下。为此,我们提出一种新方法——高效多样性经验回放(EDER),利用确定性点过程建模样本间多样性,并据此优先回放差异较大的经验。为应对大规模状态空间,引入乔列斯基分解;同时采用拒绝采样筛选高多样性样本,进一步提升学习效果。在MuJoCo的机器人操控任务、Atari游戏以及Habitat中的真实室内环境上进行了大量实验。结果表明,该方法显著提升学习效率,并在高维、现实环境中实现更优性能。

原文摘要 · Abstract (English)

Experience replay is widely used to improve learning efficiency in reinforcement learning by leveraging past experiences. However, existing experience replay methods, whether based on uniform or prioritized sampling, often suffer from low efficiency, particularly in real-world scenarios with high-dimensional state spaces. To address this limitation, we propose a novel approach, Efficient Diversity-based Experience Replay (EDER). EDER employs a determinantal point process to model the diversity between samples and prioritizes replay based on the diversity between samples. To further enhance learning efficiency, we incorporate Cholesky decomposition for handling large state spaces in realistic environments. Additionally, rejection sampling is applied to select samples with higher diversity, thereby improving overall learning efficacy. Extensive experiments are conducted on robotic manipulation tasks in MuJoCo, Atari games, and realistic indoor environments in Habitat. The results demonstrate that our approach not only significantly improves learning efficiency but also achieves superior performance in high-dimensional, realistic environments.

强化学习经验回放机器人控制多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。