arXiv:2605.12206cs.LG2026-05被引 1

揭示了多稳态是实现长时序强化学习泛化的核心机制

On the Importance of Multistability for Horizon Generalization in Reinforcement Learning

论文配图:On the Importance of Multistability for Horizon Generalization in Reinforcement Learning
图 1 · 摘自论文原文
  • 提出多稳态为长时序泛化的必要条件,需结合瞬态动态
  • 发现现代并行化RNN因单稳态设计无法跨时序泛化
  • 为可扩展长时序RL提供新架构设计方向

在部分可观测马尔可夫决策过程(POMDP)中,强化学习代理需依赖记忆(通常由循环神经网络RNN编码)整合历史观测信息。长时序POMDP尤其困难:训练面临泛化差、样本效率低和探索成本高的问题。理想情况下,短时序训练的代理应能保持在任意更长时序下的最优行为,但当前缺乏形式化框架描述其可行性。本文形式化定义了时间时序泛化——策略对所有时序均保持最优的性质,推导出其充要条件,并实验评估了非线性及可并行化RNN变体的实现能力。分析表明,多稳态是实现时序泛化的必要条件,在简单任务中亦充分;复杂任务还需瞬态动力学。相反,现代可并行化架构(如状态空间模型与门控线性RNN)因构造上为单稳态,无法实现跨时序泛化。结论指出,多稳态与瞬态动态是时序泛化的两个关键互补动力学模式,而现有并行化RNN均不具备二者。因此,设计兼具这两种模式的可并行化架构,成为实现可扩展长时序强化学习的关键方向。

原文摘要 · Abstract (English)

In reinforcement learning (RL), agents acting in partially observable Markov decision processes (POMDPs) must rely on memory, typically encoded in a recurrent neural network (RNN), to integrate information from past observations. Long-horizon POMDPs, in which the relevant observation and the optimal action are separated by many time steps (called the horizon), are particularly challenging: training suffers from poor generalization, severe sample inefficiency, and prohibitive exploration costs. Ideally, an agent trained on short horizons would retain optimal behavior at arbitrarily longer ones, but no formal framework currently characterizes when this is achievable. To fill this gap, we formalized temporal horizon generalization, the property that a policy remains optimal for all horizons, derived a necessary and sufficient condition for it, and experimentally evaluated the ability of nonlinear and parallelizable RNN variants to achieve it. This paper presents the resulting theoretical framework, the empirical evaluation, and the dynamical interpretation linking RNN behavior to temporal horizon generalization. Our analyses reveal that multistability is necessary for temporal horizon generalization and, in simple tasks, sufficient; more complex tasks further require transient dynamics. In contrast, modern parallelizable architectures, namely state space models and gated linear RNNs, are monostable by construction and consequently fail to generalize across temporal horizons. We conclude that multistability and transient dynamics are two essential and complementary dynamical regimes for horizon generalization, and that no current parallelizable RNN exhibits both. Designing parallelizable architectures that combine these regimes thus emerges as a key direction for scalable long-horizon RL.

强化学习时序泛化RNN动力学多稳态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。