arXiv:2605.23565cs.LGcs.AI2026-05

研究强化学习训练顺序如何影响新环境下的目标泛化能力

Understanding Goal Generalisation in Sequential Reinforcement Learning

论文配图:Understanding Goal Generalisation in Sequential Reinforcement Learning
图 1 · 摘自论文原文
  • 通过模拟低维潜在变量演化预测泛化行为
  • 早期学得的目标会持续影响后续学习
  • 方法可解释且适用于未见过的训练流程

强化学习智能体在分布外环境中常表现出意外的目标导向行为,但我们对这类智能体如何根据训练历史泛化到新环境仍缺乏系统理解。本文研究了超过100条顺序训练路径,在250多个分布外环境中评估行为表现。发现显著特征驱动泛化,且早期学习的目标会持续影响后期学习。为此提出潜空间策略梯度(latent policy gradients)方法,通过模拟训练中低维潜在变量的演化,基于简单行为映射模型预测高奖励路径。该方法具有强预测准确性,能泛化至未见训练流程,并具备可解释性。研究揭示:尽管分布外行为依赖完整训练历程,但其背后存在可捕捉的结构,为从发展视角理解目标泛化奠定基础。

原文摘要 · Abstract (English)

Reinforcement learning agents often exhibit unintended goal-directed behaviour outside their training distribution, but we currently lack a principled understanding of how such agents will generalise to novel environments based on their training history. We address this gap for agents trained sequentially on one or more tasks. We study over 100 sequential training pipelines, evaluating behaviour across over 250 out-of-distribution environments. We find that salient features drive generalisation, and that goals learnt early in training can persist and influence those acquired later. To explain these phenomena, we introduce latent policy gradients, a method that predicts what out-of-distribution behaviour a training pipeline will likely induce. Our method simulates the evolution of low-dimensional latent variables during training according to what would achieve high reward on the training objective with respect to a simple model of how the latent variables map to behaviour. It achieves strong predictive accuracy, generalises to unseen types of training pipeline, and is interpretable. Our findings demonstrate that while out-of-distribution RL agent behaviour is dependent on the whole training pipeline, this dependence has an underlying structure we can capture, laying groundwork for understanding goal generalisation from a developmental perspective.

强化学习目标泛化训练顺序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。