提出对抗性强化学习框架,提升连续控制任务的鲁棒性。
Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives
- 用极小极大优化平衡智能体与干扰者策略
- 在MuJoCo中对扰动和参数变化的鲁棒性显著提升
- 适合需要高可靠性的机器人控制场景
强化学习在各类控制与决策任务中取得显著进展,但在遭遇意外外部扰动和模型不确定性时,性能常出现不稳定或退化。本文提出极小极大深度确定性策略梯度(MMDDPG),用于连续控制任务中学习抗干扰策略。训练过程建模为用户策略与对抗性扰动策略间的极小极大优化:用户策略最小化目标函数,而对抗策略最大化该函数。为稳定此交互,引入分数目标函数,平衡任务表现与扰动强度,避免过度剧烈的扰动,促进稳健学习。在MuJoCo环境中的实验表明,所提方法在应对外部力扰动和模型参数变化方面均表现出显著增强的鲁棒性。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has achieved remarkable success in a wide range of control and decision-making tasks. However, RL agents often exhibit unstable or degraded performance when deployed in environments subject to unexpected external disturbances and model uncertainties. Consequently, ensuring reliable performance under such conditions remains a critical challenge. In this paper, we propose minimax deep deterministic policy gradient (MMDDPG), a framework for learning disturbance-resilient policies in continuous control tasks. The training process is formulated as a minimax optimization problem between a user policy and an adversarial disturbance policy. In this problem, the user learns a robust policy that minimizes the objective function, while the adversary generates disturbances that maximize it. To stabilize this interaction, we introduce a fractional objective that balances task performance and disturbance magnitude. This objective prevents excessively aggressive disturbances and promotes robust learning. Experimental evaluations in MuJoCo environments demonstrate that the proposed MMDDPG achieves significantly improved robustness against both external force perturbations and model parameter variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。