让强化学习控制更稳定,避免初始微小变化导致系统崩溃
GIFT: Global stabilisation via Intrinsic Fine Tuning

- 用定制奖励函数优化现有策略的全局稳定性
- 在保持任务性能的同时显著提升系统长期稳定性
- 适合需要可靠控制的机器人、自动驾驶等真实场景
深度强化学习策略在具有非线性接触力的复杂连续控制环境中表现优异。然而,这些策略常产生混沌状态动态,初始条件的微小变化会显著影响系统的长期行为。这种对初始条件的高度敏感性限制了深度强化学习在真实世界控制系统中的应用,后者通常要求性能和稳定性保障。为此,我们提出一种通用训练框架GIFT(Global stabilisation via Intrinsic Fine Tuning),通过自定义奖励函数直接优化现有高性能深度强化学习策略的全局稳定性。实验表明,GIFT在维持相近任务性能的同时显著提升了控制交互的稳定性,从而增强了深度强化学习策略在真实世界控制系统中的适用性。
原文摘要 · Abstract (English)
Deep reinforcement learning policies achieve strong performance in complex continuous control environments with nonlinear contact forces. However, these policies often produce chaotic state dynamics, with trivially small changes to the initial conditions significantly impacting the long-term behaviour of the control system. This high sensitivity to initial conditions limits the application of Deep RL to real-world control systems where performance and stability guarantees are often required. To address this issue, we propose Global stabilisation via Intrinsic Fine Tuning (GIFT), a general-purpose training framework which directly optimises the global stability of existing high-performing deep RL policies using a custom reward function. We demonstrate that GIFT increase the stability of the control interaction while maintaining comparable task performance, thereby improving the suitability of deep RL policies for real-world control systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。