分层强化学习解决稀疏奖励长时序任务探索难题
Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning

- 两级架构:高层规划+底层SAC连续控制,熵正则化优化策略
- 在SAR-2数据集上成功率达89.3%,覆盖效率提升42%
- 适合长时序、奖励稀疏的机器人导航等复杂任务
稀疏奖励的长时序任务中探索面临重大挑战。为此,我们提出一种两级分层强化学习(HRL)框架:高层负责高层次战略规划,低层采用连续控制的软演员-评论家(SAC)算法,两者均使用熵正则化策略优化。该框架在Search-and-Rescue-2(SAR-2)数据集上进行训练与评估。结果显示,HRL-SAC有效应对延迟奖励和连续控制下的长时序搜索问题,在成功率、覆盖效率和收敛速度上均优于扁平SAC基线。这些发现表明,分层熵正则化策略是解决长时序稀疏奖励强化学习任务的有前景方案。
原文摘要 · Abstract (English)
Exploration in sparse-reward long-horizon tasks poses significant challenges for reinforcement learning. To address these challenges, we propose a two-level Hierarchical Reinforcement Learning (HRL) framework. The first level handles high-level strategic planning, while the low-level uses the continuous-control Soft Actor-Critic (SAC) algorithm, and they utilize entropy-regularized policy optimization. The proposed framework was trained and evaluated using the Search-and-Rescue-2 (SAR-2) dataset. HRL-SAC effectively addresses sparse-reward long-horizon search problems characterized by delayed rewards and continuous control, and its outperforming the flat SAC baseline reinforcement learning in terms of success rates, coverage efficiency, and convergence. These findings indicate that hierarchical entropy-regularized policies are a promising solution to tackle long-horizon sparse-reward reinforcement learning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。