arXiv:2501.02709cs.LGcs.AI2025-01ICLR被引 14

让智能体学会从近处目标推导远距离目标的策略,提升强化学习泛化能力

Horizon Generalization in Reinforcement Learning

  • 通过规划不变性设计策略,使近距训练成果可推广至远距目标
  • 理论证明在合理假设下,远距离目标可达性可实现
  • 为跨领域泛化技术迁移提供新思路,适合强化学习研究者

我们从泛化的角度研究目标导向的强化学习,但不采用传统随机增强或域随机化。相反,目标是学习能够适应时间跨度(horizon)泛化的策略:在训练时仅需到达邻近目标(易学),却能在实际中成功抵达遥远目标(难学)。正如归一化层使网络对尺度变化保持不变从而泛化到不同输入,我们发现这种时间跨度泛化与规划不变性密切相关——一个向目标前进的策略,其行动选择应与经过途中某路点导航一致。因此,仅训练近处目标的策略,也能成功到达任意远的目标。理论分析证明,在一定假设下,时间和规划的泛化皆可能实现。本文还呈现新实验结果,并回顾已有工作支持上述理论。整体表明,可借鉴其他机器学习领域的不变性与泛化技术,推动该理想属性的发展。

原文摘要 · Abstract (English)

We study goal-conditioned RL through the lens of generalization, but not in the traditional sense of random augmentations and domain randomization. Rather, we aim to learn goal-directed policies that generalize with respect to the horizon: after training to reach nearby goals (which are easy to learn), these policies should succeed in reaching distant goals (which are quite challenging to learn). In the same way that invariance is closely linked with generalization is other areas of machine learning (e.g., normalization layers make a network invariant to scale, and therefore generalize to inputs of varying scales), we show that this notion of horizon generalization is closely linked with invariance to planning: a policy navigating towards a goal will select the same actions as if it were navigating to a waypoint en route to that goal. Thus, such a policy trained to reach nearby goals should succeed at reaching arbitrarily-distant goals. Our theoretical analysis proves that both horizon generalization and planning invariance are possible, under some assumptions. We present new experimental results and recall findings from prior work in support of our theoretical results. Taken together, our results open the door to studying how techniques for invariance and generalization developed in other areas of machine learning might be adapted to achieve this alluring property.

强化学习泛化能力目标导向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。