让循环强化学习的隐藏状态有物理意义,提升鲁棒性。
Neural Co-state Policies: Structuring Hidden States in Recurrent Reinforcement Learning

- 将隐藏状态与最优控制中的协态变量对齐,用协态损失显式结构化
- 在部分可观测环境上性能持平或超越基线,对传感器遮蔽零样本鲁棒
- 适合关注连续控制可解释性与鲁棒性的研究者
智能体在部分可观测环境下的决策能力至关重要:即使观测不完整,也能有效推理与行动。尽管基于循环神经网络的策略可通过编码历史来缓解此问题,但其内部动态仍为不可解释的黑箱。本文首次建立隐状态与最优控制中庞特里亚金极小值原理(PMP)之间的形式联系,证明标准循环架构的隐表示直接对应于PMP协态变量,且读出层可被理解为执行哈密顿量最小化。由于标准奖励最大化无法自然发现这一对齐,我们引入基于PMP的协态损失以显式结构化内部动态。实验表明,该方法在部分可观测的DMControl任务上表现相当或更优,并对零样本分布外传感器遮蔽具有鲁棒性。通过将循环网络视为受极小值原理支配的动力系统,我们提供了一种设计稳健连续控制策略的理论基础。
原文摘要 · Abstract (English)
A key capability of intelligent agents is operating under partial observability: reasoning and acting effectively despite missing or incomplete state observations. While recurrent (memory-based) policies learned via reinforcement learning address this by encoding history into latent state representations, their internal dynamics remain uninterpretable black boxes. This paper establishes a formal link between these hidden states and the Pontryagin minimum principle (PMP) from optimal control. We demonstrate that for standard recurrent architectures, latent representations map directly to PMP co-states, which allows the readout layer to be interpreted as performing Hamiltonian minimization. Because standard reward maximization does not naturally discover this alignment, we introduce a PMP-derived co-state loss to explicitly structure the internal dynamics. Empirically, this approach matches or improves performance on partially observable DMControl tasks, and is robust against zero-shot out-of-distribution sensor masking. By framing recurrent networks as dynamic processes governed by the minimum principle, we provide a principled approach to designing robust continuous control policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。