用时间感知网络解决低成本四足机器人延迟问题,实现稳定行走。
Reinforcement Learning on Cost-Constrained Quadrupedal Hardware

- 用前向模型+时间感知神经网络处理电机延迟与噪声反馈
- 在+320毫秒延迟下仍能保持稳定步态,平均延迟超50毫秒
- 发现类似脊椎动物的自维持节律生成机制,适合低预算硬件部署
在低成本机器人平台部署学习到的控制策略时,传输延迟和噪声电机反馈会显著扩大仿真到现实的差距。以Mini Pupper 2为例,实测>50毫秒的传输延迟使运动任务从标准马尔可夫决策过程变为部分可观测问题。本文采用生物启发方法应对延迟与噪声反馈,从而缩小仿真与真实间的鸿沟。通过低成本四足平台实验发现,结合平均执行器延迟的前向模型与时间感知神经网络,可实现鲁棒行走。此外,该网络自主学习出一种中央模式发生器(CPG)——一种自维持的周期性步态,对+320毫秒延迟扰动仍具鲁棒性,类似于脊椎动物脊髓中的生理机制。我们推测,时间上的自组织可能是低成本运动控制的通用策略。
原文摘要 · Abstract (English)
Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-to-real gap. The chasm of simulation to deployment in hardware lies in the delay of the actuator reaching the commanded position. On platforms such as the Mini Pupper 2, a measured >50 ms transport delay transforms the locomotion task from a standard Markov decision process into a partially observable one. In this paper, we take a biologically inspired approach of handling noisy and delayed feedback to close the sim-to-real gap, thereby expanding the capability of reinforcement learning on cost-constrained hardware. Using a low-cost quadrupedal hardware platform, we find that using a forward model of the average actuator delay, paired with a time-aware neural network results in robust locomotion. Additionally, our time-aware neural network learned a central pattern generator (CPG): a self-sustaining rhythmic gait that is robust to +320 ms latency perturbations, mirroring the CPGs found in the spinal cords of vertebrates. We posit that temporal self-organization may be a general strategy for cost-constrained locomotion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。