dense reward下仍可保持强化学习的度量结构,提升训练效率。
Quasimetric Value Functions with Dense Rewards
- 利用密集奖励保持值函数的拟度量性质
- 12个基准任务中密集奖励训练效果优于稀疏奖励
- 适合需要高效样本学习的机器人控制场景
作为强化学习向可参数化目标扩展的通用方法,目标条件强化学习(GCRL)在机器人等复杂任务中应用广泛。近期研究发现,GCRL的最优值函数 $Q^\ast(s,a,g)$ 具有拟度量结构,由此设计出符合该结构的神经架构。然而,相关分析均基于稀疏奖励设定——这会加剧样本复杂性问题。本文证明,支撑拟度量的核心性质(三角不等式)在密集奖励设置下依然成立。与此前认为密集奖励对GCRL有害的观点相反,我们识别出三角不等式得以保持的关键条件:满足该条件的密集奖励函数只会改善,绝不会恶化样本复杂性。这为结合密集奖励训练高效神经架构开辟了新路径,进一步提升样本效率。我们在12个标准的GCRL基准环境(涵盖挑战性连续控制任务)中验证该方法,实证结果表明,在密集奖励下训练拟度量值函数显著优于稀疏奖励设置。
原文摘要 · Abstract (English)
As a generalization of reinforcement learning (RL) to parametrizable goals, goal conditioned RL (GCRL) has a broad range of applications, particularly in challenging tasks in robotics. Recent work has established that the optimal value function of GCRL $Q^\ast(s,a,g)$ has a quasimetric structure, leading to targetted neural architectures that respect such structure. However, the relevant analyses assume a sparse reward setting -- a known aggravating factor to sample complexity. We show that the key property underpinning a quasimetric, viz., the triangle inequality, is preserved under a dense reward setting as well. Contrary to earlier findings where dense rewards were shown to be detrimental to GCRL, we identify the key condition necessary for triangle inequality. Dense reward functions that satisfy this condition can only improve, never worsen, sample complexity. This opens up opportunities to train efficient neural architectures with dense rewards, compounding their benefits to sample complexity. We evaluate this proposal in 12 standard benchmark environments in GCRL featuring challenging continuous control tasks. Our empirical results confirm that training a quasimetric value function in our dense reward setting indeed outperforms training with sparse rewards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。