用多时间尺度的预测表征提升强化学习在持续变化环境中的稳定性
Balancing Plasticity and Stability with Fast and Slow Successor Features

- 采用神经启发的突触固化机制,以成功特征(SFs)为优化目标
- 在连续变化环境中,多时间尺度的SF固化使性能优于传统方法20%以上
- 适合关注持续学习与稳定适应的智能体设计者
智能体需在非平稳环境中持续适应,但深度强化学习常在此类场景中表现不佳。现有研究多通过突变设置非平稳性,而真实环境更常见渐进式演化。为此,我们改造3D Miniworld和MuJoCo环境,引入自然化的持续非平稳性,评估稳定性与适应性对性能的影响。结果表明,强调稳定性的方法(如突触固化)优于侧重可塑性的方法(如参数重置)。受此启发,结合成功特征(SFs)可降低干扰的特性,我们发现将突触固化应用于SFs能显著提升性能。尤其当SFs在多个时间尺度上被稳定时,效果最佳,因这些尺度分别捕捉了环境变化的不同特征。研究表明,在渐进式变化中,稳定性更重要,且多时间尺度的预测表征固化是有效策略。
原文摘要 · Abstract (English)
A hallmark of intelligence is the ability to adapt in non-stationary environments, yet deep Reinforcement Learning (RL) agents often struggle in such settings. Prior studies introduce non-stationarity through abrupt shifts in features or dynamics, whereas real-world environments often evolve gradually through continual drift. This distinction has important implications for the "stability-plasticity dilemma" in RL, as abrupt task changes may demand more plasticity than naturalistic settings. To address this, we modify existing 3D Miniworld and MuJoCo environments to incorporate naturalistic, continual non-stationarity, and use them to examine how stability and adaptation affect performance under continuous environmental change. We find that methods favoring stability, such as synaptic consolidation, outperform approaches focused on plasticity, such as parameters resetting. Motivated by this result, and prior evidence that Successor Features (SFs) reduce interference, we investigate whether SFs are better consolidation targets than Q-values. Across both environments, applying neuro-inspired synaptic consolidation to SFs yields superior performance on continually changing settings. Moreover, consolidation is most effective when SFs are stabilized across multiple timescales, which capture complementary aspects of gradual environmental change. Together, these results suggest that stability is more critical in continual learning when changes are gradual, and that multi-timescale consolidation of predictive representations is an effective approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。