arXiv:2409.05811cs.RO2024-09被引 4

用模型强化学习控制欠驱动双摆,参赛即胜

Learning control of underactuated double pendulum with Model-Based Reinforcement Learning

  • 基于MC-PILCO模型强化学习算法构建控制器
  • 在复杂动态系统中实现稳定控制与高精度轨迹跟踪
  • 适合机器人控制与强化学习竞赛场景

本报告介绍了我们在2024年IROS举办的第二届AI奥林匹克竞赛中提出的解决方案。该方案基于一种近期的基于模型的强化学习算法MC-PILCO。除了简要回顾该算法外,还讨论了在当前任务中实现MC-PILCO的关键因素。实验表明,该方法在欠驱动双摆系统的控制任务中表现出优异的稳定性与轨迹跟踪能力,在竞赛中取得了领先结果。

原文摘要 · Abstract (English)

This report describes our proposed solution for the second AI Olympics competition held at IROS 2024. Our solution is based on a recent Model-Based Reinforcement Learning algorithm named MC-PILCO. Besides briefly reviewing the algorithm, we discuss the most critical aspects of the MC-PILCO implementation in the tasks at hand.

强化学习机器人控制模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。