arXiv:2503.00282cs.RO2025-03被引 1

提出成本感知框架,让无人机在变风环境下长期学习不崩溃。

Maintaining Plasticity in Reinforcement Learning: A Cost-Aware Framework for Aerial Robot Control in Non-stationary Environments

  • 用回顾性成本机制动态调整学习率,平衡奖励与损失。
  • 在变风环境中训练悬停策略,比标准PPO少11.29%的死单元。
  • 适合需要长期适应非平稳环境的无人机控制场景。

强化学习(RL)在短时训练中能保持飞行机器人控制策略的可塑性,但在非平稳环境中的长期学习中易出现可塑性丧失。例如,标准近端策略优化(PPO)策略在长期训练中会失效,导致控制性能显著下降。为此,本文提出一种成本感知框架,引入回顾性成本机制(RECOM),通过奖励与损失之间的成本梯度关系,在非平稳环境中平衡奖励与损失。该框架利用成本梯度动态调整学习率,主动适应扰动风环境。实验结果表明,该框架在变风条件下学习到的悬停策略未发生策略坍塌,相比使用L2正则化的PPO,死单元比例降低11.29%。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has demonstrated the ability to maintain the plasticity of the policy throughout short-term training in aerial robot control. However, these policies have been shown to loss of plasticity when extended to long-term learning in non-stationary environments. For example, the standard proximal policy optimization (PPO) policy is observed to collapse in long-term training settings and lead to significant control performance degradation. To address this problem, this work proposes a cost-aware framework that uses a retrospective cost mechanism (RECOM) to balance rewards and losses in RL training with a non-stationary environment. Using a cost gradient relation between rewards and losses, our framework dynamically updates the learning rate to actively train the control policy in a disturbed wind environment. Our experimental results show that our framework learned a policy for the hovering task without policy collapse in variable wind conditions and has a successful result of 11.29% less dormant units than L2 regularization with PPO.

强化学习无人机控制非平稳环境可塑性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。