arXiv:2503.01468cs.LG2025-03被引 3

用不确定性感知的策略优化,让智能体在动态变化环境中更快适应。

Overcoming Non-stationary Dynamics with Evidential Proximal Policy Optimization

  • 引入证据理论构建概率化价值网络,实时捕捉环境变化带来的分布偏移。
  • 在连续控制任务中,相比现有方法提升任务回报和整体性能。
  • 适合需要快速适应环境变化的机器人控制与自动驾驶场景。

非平稳环境下的连续控制是深度强化学习的一大挑战,状态转移动态的时间依赖性加剧了无模型深度演员-评论家架构的稳定性问题。我们提出两个关键属性有助于克服动态非平稳性:(i) 保持评论家网络的可塑性,(ii) 实现有方向性的探索以快速适应动态变化。通过在策略学习中采用证据论评论家,我们实现了这两点。证据论设计能快速准确地逼近状态值周围的不确定性,从而通过检测动态变化引起的分布偏移来维持评论家的可塑性。同时,概率化评论家使演员训练目标成为随机变量,间接支持有方向的探索策略。我们将该算法命名为「证据论近端策略优化(EPPO)」,因其在策略评估与改进阶段均依赖证据论不确定性量化。在环境动态以固定周期变化的非平稳连续控制任务中,实验表明该算法在任务特定回报和总体回报上均优于当前最先进的在策略强化学习方法。

原文摘要 · Abstract (English)

Continuous control of non-stationary environments is a major challenge for deep reinforcement learning algorithms. The time-dependency of the state transition dynamics aggravates the notorious stability problems of model-free deep actor-critic architectures. We posit that two properties will play a key role in overcoming non-stationarity in transition dynamics: (i)~preserving the plasticity of the critic network and (ii) directed exploration for rapid adaptation to changing dynamics. We show that performing on-policy reinforcement learning with an evidential critic provides both. The evidential design ensures a fast and accurate approximation of the uncertainty around the state value, which maintains the plasticity of the critic network by detecting the distributional shifts caused by changes in dynamics. The probabilistic critic also makes the actor training objective a random variable, enabling the use of directed exploration approaches as a by-product. We name the resulting algorithm \emph{Evidential Proximal Policy Optimization (EPPO)} due to the integral role of evidential uncertainty quantification in both policy evaluation and policy improvement stages. Through experiments on non-stationary continuous control tasks, where the environment dynamics change at regular intervals, we demonstrate that our algorithm outperforms state-of-the-art on-policy reinforcement learning variants in both task-specific and overall return.

强化学习非平稳环境不确定性建模策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。