针对双摆任务的强化学习控制器,提升稳定性和适应性。
Average-Reward Maximum Entropy Reinforcement Learning for Global Policy in Double Pendulum Tasks
- 基于平均奖励熵优势策略优化,改进控制算法。
- 在新竞赛框架下实现双摆的摆起与稳定,表现稳健。
- 适合关注复杂系统控制与强化学习应用的研究者。
本报告提出了一种基于强化学习的方法,用于解决acrobot和pendubot的摆起与稳定任务,特别针对ICRA 2025第三届人工智能奥林匹克竞赛的最新指南进行优化。在先前提出的平均奖励熵优势策略优化(AR-EAPO)算法基础上,我们进一步完善了控制方案,以应对新的竞赛场景与评估指标。大量仿真结果表明,该控制器在更新后的任务框架中表现出良好的鲁棒性与有效性,具备较强的适应能力。
原文摘要 · Abstract (English)
This report presents our reinforcement learning-based approach for the swing-up and stabilisation tasks of the acrobot and pendubot, tailored specifcially to the updated guidelines of the 3rd AI Olympics at ICRA 2025. Building upon our previously developed Average-Reward Entropy Advantage Policy Optimization (AR-EAPO) algorithm, we refined our solution to effectively address the new competition scenarios and evaluation metrics. Extensive simulations validate that our controller robustly manages these revised tasks, demonstrating adaptability and effectiveness within the updated framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。