SLOPE通过构建乐观势能场,提升稀疏奖励下的强化学习探索效率。
SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning
- 用乐观分布回归构建高置信度上界势能场,增强稀疏成功信号
- 在5个基准的30多个任务中,稀疏/半稀疏/密集奖励下均超越基线
- 适合需要高效探索的机器人控制等实际应用
基于模型的强化学习(MBRL)虽样本效率高,但在稀疏奖励场景下表现受限。根本瓶颈在于稀疏奖励下标准奖励模型常导致平坦的梯度景观,难以指导规划。为此,我们提出一种新框架SLOPE:通过将奖励建模从预测稀疏标量转变为构建有信息量的势能场。SLOPE采用乐观分布回归估计高置信度上界,放大罕见的成功信号,确保充分探索梯度。在5个基准上的30多个任务及真实机器人部署中,SLOPE在完全稀疏、半稀疏和密集奖励设置下均持续优于领先基线。
原文摘要 · Abstract (English)
Model-based reinforcement learning (MBRL) is sample-efficient but struggles in sparse reward settings. A critical bottleneck arises from the lack of informative gradients in sparse settings, where standard reward models often yield flat landscapes that struggle to guide planning. To address this challenge, we propose Shaping Landscapes with Optimistic Potential Estimates (SLOPE), a novel framework that shifts reward modeling from predicting sparse scalars to constructing informative potential landscapes. SLOPE employs optimistic distributional regression to estimate high-confidence upper bounds, which amplifies rare success signals and ensures sufficient exploration gradients. Evaluations on 30+ tasks across 5 benchmarks and real-world robotic deployments, demonstrate that SLOPE consistently outperforms leading baselines in fully sparse, semi-sparse, and dense rewards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。