arXiv:2501.10698cs.ROcs.LG2025-01被引 4

提出可解释的神经控制网络,让机器人用更少样本高效学会走路

An Interpretable Neural Control Network with Adaptable Online Learning for Sample Efficient Robot Locomotion Learning

  • 分三层设计可解释的神经网络,逐层生成运动状态与动作指令
  • 减少40%训练样本,性能提升150%,物理机器人仅需10分钟学习
  • 适合关注模型可解释性与高效强化学习的研究者

基于强化学习的机器人行走训练存在样本效率低和黑箱问题。本文提出SME-AGOL框架:首先,序列运动执行器(SME)为三层可解释神经网络,第一层生成顺序传播的隐状态,第二层构建干扰小的三角基底,第三层将基底映射为电机指令;其次,自适应梯度加权在线学习(AGOL)算法优先更新高相关性参数。该框架使每个隐状态/基底对应一个关键姿态或机器人构型,具备可分析性。相比现有方法,SME-AGOL在模拟六足机器人上减少40%样本、最终奖励提升150%,且在真实六足机器人上仅需10分钟即可从零开始完成学习。本工作不仅实现高效可解释的行走学习,还揭示了可解释性对提升样本效率与性能的潜力。

原文摘要 · Abstract (English)

Robot locomotion learning using reinforcement learning suffers from training sample inefficiency and exhibits the non-understandable/black-box nature. Thus, this work presents a novel SME-AGOL to address such problems. Firstly, Sequential Motion Executor (SME) is a three-layer interpretable neural network, where the first produces the sequentially propagating hidden states, the second constructs the corresponding triangular bases with minor non-neighbor interference, and the third maps the bases to the motor commands. Secondly, the Adaptable Gradient-weighting Online Learning (AGOL) algorithm prioritizes the update of the parameters with high relevance score, allowing the learning to focus more on the highly relevant ones. Thus, these two components lead to an analyzable framework, where each sequential hidden state/basis represents the learned key poses/robot configuration. Compared to state-of-the-art methods, the SME-AGOL requires 40% fewer samples and receives 150% higher final reward/locomotion performance on a simulated hexapod robot, while taking merely 10 minutes of learning time from scratch on a physical hexapod robot. Taken together, this work not only proposes the SME-AGOL for sample efficient and understandable locomotion learning but also emphasizes the potential exploitation of interpretability for improving sample efficiency and learning performance.

机器人行走强化学习可解释性样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。