arXiv:2602.12753cs.LG2026-02

提出分层预测表示,让智能体在复杂环境中更稳定地迁移学习。

Hierarchical Successor Representation for Robust Transfer

  • 用时序抽象构建分层预测表示,提升状态特征稳定性。
  • 结合非负矩阵分解,实现低秩稀疏表征,支持高效任务迁移。
  • 发现可解释的拓扑结构,适合复杂环境下的自主探索与迁移。

successor representation(SR)能解耦预测动态与奖励,实现奖励配置下的快速泛化。但经典SR受策略依赖性限制:策略随持续学习、环境非平稳性和任务需求变化而改变,导致已有预测表示失效。此外,在拓扑复杂的环境中,SR存在谱扩散问题,导致特征密集重叠且扩展性差。本文提出分层成功者表示(HSR),通过将时序抽象融入预测表示构建,学习对任务引发的策略变化具有鲁棒性的稳定状态特征。对HSR应用非负矩阵分解(NMF)得到稀疏、低秩的状态表示,显著提升多隔间环境中的新任务样本效率。进一步分析表明,HSR-NMF可发现可解释的拓扑结构,提供无策略依赖的分层地图,有效衔接模型无关最优性与模型驱动灵活性。除支持任务迁移外,还证明其时序扩展的预测结构可用于高效探索,可扩展至大规模程序生成环境。

原文摘要 · Abstract (English)

The successor representation (SR) provides a powerful framework for decoupling predictive dynamics from rewards, enabling rapid generalisation across reward configurations. However, the classical SR is limited by its inherent policy dependence: policies change due to ongoing learning, environmental non-stationarities, and changes in task demands, making established predictive representations obsolete. Furthermore, in topologically complex environments, SRs suffer from spectral diffusion, leading to dense and overlapping features that scale poorly. Here we propose the Hierarchical Successor Representation (HSR) for overcoming these limitations. By incorporating temporal abstractions into the construction of predictive representations, HSR learns stable state features which are robust to task-induced policy changes. Applying non-negative matrix factorisation (NMF) to the HSR yields a sparse, low-rank state representation that facilitates highly sample-efficient transfer to novel tasks in multi-compartmental environments. Further analysis reveals that HSR-NMF discovers interpretable topological structures, providing a policy-agnostic hierarchical map that effectively bridges model-free optimality and model-based flexibility. Beyond providing a useful basis for task-transfer, we show that HSR's temporally extended predictive structure can also be leveraged to drive efficient exploration, effectively scaling to large, procedurally generated environments.

强化学习分层表示迁移学习探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。