arXiv:2409.08938cs.ROcs.LG2024-09被引 3

用平均奖励熵最大化强化学习,让欠驱动双摆自动完成摆起与稳定。

Average-Reward Maximum Entropy Reinforcement Learning for Underactuated Double Pendulum Tasks

  • 结合平均奖励与最大熵思想设计新型强化学习算法
  • 在双摆任务中性能与鲁棒性优于已有方法
  • 无需复杂奖励函数或系统模型,适合仿真环境应用

本文针对AI奥林匹克大赛在IROS 2024中的机械臂双摆摆起与稳定任务,提出一种基于平均奖励熵优势策略优化(AR-EAPO)的无模型强化学习方法。该方法融合平均奖励强化学习与最大熵强化学习,无需依赖复杂的奖励函数或系统模型。实验结果表明,在acrobot与pendubot仿真环境中,所提控制器在性能与鲁棒性上均优于现有基准方法。

原文摘要 · Abstract (English)

This report presents a solution for the swing-up and stabilisation tasks of the acrobot and the pendubot, developed for the AI Olympics competition at IROS 2024. Our approach employs the Average-Reward Entropy Advantage Policy Optimization (AR-EAPO), a model-free reinforcement learning (RL) algorithm that combines average-reward RL and maximum entropy RL. Results demonstrate that our controller achieves improved performance and robustness scores compared to established baseline methods in both the acrobot and pendubot scenarios, without the need for a heavily engineered reward function or system model. The current results are applicable exclusively to the simulation stage setup.

强化学习双摆控制无模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。