arXiv:2410.20096cs.RO2024-10被引 3

用历史信息增强强化学习,让机器人更稳地控制不完全可控系统。

Velocity-History-Based Soft Actor-Critic Tackling IROS'24 Competition "AI Olympics with RealAIGym"

  • 引入卷积网络编码状态历史作为上下文向量
  • 在双轨比赛中均取得高分与良好鲁棒性
  • 适合需要实时抗干扰的机器人控制场景

「AI Olympics with RealAIGym」竞赛要求参赛者使用先进控制算法稳定混沌的欠驱动动力学系统。本文提交至IROS'24竞赛的新方法基于流行的无模型熵正则强化学习算法Soft Actor-Critic(SAC)。通过在状态中加入一个‘上下文’向量,该向量利用卷积神经网络(CNN)编码最近的历史信息,以补偿真实系统中的未建模效应。该方法在竞赛的两个赛道——Pendubot和Acrobot上均取得了高性能得分与具有竞争力的鲁棒性得分。

原文摘要 · Abstract (English)

The ``AI Olympics with RealAIGym'' competition challenges participants to stabilize chaotic underactuated dynamical systems with advanced control algorithms. In this paper, we present a novel solution submitted to IROS'24 competition, which builds upon Soft Actor-Critic (SAC), a popular model-free entropy-regularized Reinforcement Learning (RL) algorithm. We add a `context' vector to the state, which encodes the immediate history via a Convolutional Neural Network (CNN) to counteract the unmodeled effects on the real system. Our method achieves high performance scores and competitive robustness scores on both tracks of the competition: Pendubot and Acrobot.

强化学习机器人控制状态历史SAC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。