提出新方法平衡强化学习中最优性与抗干扰能力的矛盾
On the Tension Between Optimality and Adversarial Robustness in Policy Optimization
- 设计双层框架BARPO,动态调节对抗强度以调和冲突
- 实验证明相比传统方法,新模型在保持鲁棒性同时提升收益
- 适合追求高可靠性的强化学习应用,如自动驾驶、机器人控制
深度强化学习中,最优性与对抗鲁棒性长期被视为相互冲突的目标。尽管近期理论研究(CAR)揭示二者存在潜在统一可能,但如何在实践中实现仍不明确。本文通过对比标准策略优化(SPO)与对抗鲁棒策略优化(ARPO),发现两者虽具理论一致性,但在实际梯度方法中仍存在根本性权衡:SPO易收敛至性能强但脆弱的一阶驻点(FOSPs),而ARPO虽更鲁棒,却导致回报下降。我们归因于ARPO中最强对抗者带来的重塑效应,其使全局优化景观复杂化,产生误导性‘粘滞’FOSPs,虽增强鲁棒性,却阻碍导航。为此,提出BARPO——一种通过调节对抗强度统一SPO与ARPO的双层框架,既提升可导航性,又保留全局最优解。大量实验表明,BARPO持续优于原始ARPO,为理论与实际性能的协调提供了可行路径。
原文摘要 · Abstract (English)
Achieving optimality and adversarial robustness in deep reinforcement learning has long been regarded as conflicting goals. Nonetheless, recent theoretical insights presented in CAR suggest a potential alignment, raising the important question of how to realize this in practice. This paper first identifies a key gap between theory and practice by comparing standard policy optimization (SPO) and adversarially robust policy optimization (ARPO). Although they share theoretical consistency, a fundamental tension between robustness and optimality arises in practical policy gradient methods. SPO tends toward convergence to vulnerable first-order stationary policies (FOSPs) with strong natural performance, whereas ARPO typically favors more robust FOSPs at the expense of reduced returns. Furthermore, we attribute this tradeoff to the reshaping effect of the strongest adversary in ARPO, which significantly complicates the global landscape by inducing deceptive sticky FOSPs. This improves robustness but makes navigation more challenging. To alleviate this, we develop the BARPO, a bilevel framework unifying SPO and ARPO by modulating adversary strength, thereby facilitating navigability while preserving global optima. Extensive empirical results demonstrate that BARPO consistently outperforms vanilla ARPO, providing a practical approach to reconcile theoretical and empirical performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。