提出自适应不确定性评价值集成框架,提升对抗强化学习稳定性与鲁棒性。
UACER: An Uncertainty-Adaptive Critic Ensemble Framework for Robust Adversarial Reinforcement Learning
- 用多个评价值网络并行估计,降低值函数方差。
- 根据认知不确定性动态调节探索,训练更稳定且性能更优。
- 适合高维复杂环境下的鲁棒决策任务,如自动驾驶、机器人控制。
鲁棒对抗强化学习已成为应对真实环境中不确定干扰的有效范式,广泛应用于自动驾驶、机器人控制等序列决策场景。该范式通常将智能体训练建模为主角与对手之间的零和马尔可夫博弈,以增强策略鲁棒性。然而,对手的可训练特性导致学习动态非平稳,加剧了训练不稳定性与收敛困难,尤其在高维复杂环境中。本文提出一种新方法——不确定性自适应评价值集成框架(UACER),包含两个组件:1)多样化评价值集成:并行使用 K 个评价值网络,稳定鲁棒对抗强化学习中的 Q 值估计,相比传统单评价值设计显著降低方差、提升鲁棒性;2)时变衰减不确定性(TDU)机制:超越简单线性加权,提出基于方差的 Q 值聚合策略,显式引入认知不确定性,自适应调节探索与利用权衡,同时稳定训练过程。在多个挑战性的 MuJoCo 控制任务上的综合实验表明,UACER 在整体性能、训练稳定性与效率上均优于现有先进方法。
原文摘要 · Abstract (English)
Robust adversarial reinforcement learning has emerged as an effective paradigm for training agents to handle uncertain disturbance in real environments, with critical applications in sequential decision-making domains such as autonomous driving and robotic control. Within this paradigm, agent training is typically formulated as a zero-sum Markov game between a protagonist and an adversary to enhance policy robustness. However, the trainable nature of the adversary inevitably induces non-stationarity in the learning dynamics, leading to exacerbated training instability and convergence difficulties, particularly in high-dimensional complex environments. In this paper, we propose a novel approach, Uncertainty-Adaptive Critic Ensemble for robust adversarial Reinforcement learning (UACER), which consists of two components: 1) Diversified critic ensemble: A diverse set of K critic networks is employed in parallel to stabilize Q-value estimation in robust adversarial reinforcement learning, reducing variance and enhancing robustness compared to conventional single-critic designs. 2) Time-varying Decay Uncertainty (TDU) mechanism: Moving beyond simple linear combinations, we propose a variance-derived Q-value aggregation strategy that explicitly incorporates epistemic uncertainty to adaptively regulate the exploration-exploitation trade-off while stabilizing the training process. Comprehensive experiments across several challenging MuJoCo control problems validate the superior effectiveness of UACER, outperforming state-of-the-art methods in terms of overall performance, stability, and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。