发现循环神经策略中隐藏的周期性动态结构,解释其泛化优势。
Unraveling the Hidden Dynamical Structure in Recurrent Neural Policies
- 将策略与环境视为联合动力系统,发现稳定环状轨迹
- 环状结构几何特征与行为关系对应,提升适应能力
- 适合研究强化学习动力学机制的研究者参考
循环神经策略在部分可观测控制和元强化学习任务中广泛应用,因其能维持内部记忆并快速适应未知场景而表现卓越。然而,其优越泛化与鲁棒性的内在机制仍不明确。本研究分析了多种训练方法、模型架构和任务下学习到的循环策略隐状态空间,发现与环境交互过程中始终存在稳定的周期性结构。这些结构在动力系统理论中类似于极限环,若将策略与环境视为联合混合动力系统。进一步发现,极限环的几何特性与策略行为具有结构化对应关系。该发现为循环策略诸多优良特性提供了新解释:极限环稳定了策略内部记忆与任务相关的环境状态,抑制了由环境不确定性带来的冗余变化;同时,环的几何结构编码了行为间的关联关系,使面对非平稳环境时更易实现技能迁移。
原文摘要 · Abstract (English)
Recurrent neural policies are widely used in partially observable control and meta-RL tasks. Their abilities to maintain internal memory and adapt quickly to unseen scenarios have offered them unparalleled performance when compared to non-recurrent counterparts. However, until today, the underlying mechanisms for their superior generalization and robustness performance remain poorly understood. In this study, by analyzing the hidden state domain of recurrent policies learned over a diverse set of training methods, model architectures, and tasks, we find that stable cyclic structures consistently emerge during interaction with the environment. Such cyclic structures share a remarkable similarity with \textit{limit cycles} in dynamical system analysis, if we consider the policy and the environment as a joint hybrid dynamical system. Moreover, we uncover that the geometry of such limit cycles also has a structured correspondence with the policies' behaviors. These findings offer new perspectives to explain many nice properties of recurrent policies: the emergence of limit cycles stabilizes both the policies' internal memory and the task-relevant environmental states, while suppressing nuisance variability arising from environmental uncertainty; the geometry of limit cycles also encodes relational structures of behaviors, facilitating easier skill adaptation when facing non-stationary environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。