arXiv:2510.07094cs.RO2025-10被引 1

通过优化采样策略,让一只机器狗学会适应多种硬件配置。

Sampling Strategies for Robust Universal Quadrupedal Locomotion Policies

  • 用三种方法采样关节增益参数,提升训练泛化能力。
  • 在模拟中训练的单一策略可直接部署到真实机器狗上。
  • 适合做机器人控制、强化学习落地的开发者参考。

本研究聚焦四足机器人运动策略的配置变异采样策略,旨在生成具有鲁棒性的通用运动控制策略。通过对比三种基本的关节增益采样方法:(1) 质量-增益间线性与多项式映射采样,(2) 基于性能的自适应过滤,(3) 均匀随机采样,验证其对策略泛化的影响。通过引入名义先验和参考模型对配置进行偏差引导,显著提升策略鲁棒性。所有训练在RaiSim仿真环境中完成,在多种异构四足机器人上进行仿真测试,并零样本部署至ANYmal四足机器人硬件。相比多个基线方法,结果表明:必须对关节控制器增益进行充分随机化,才能有效缩小仿真到现实的差距。

原文摘要 · Abstract (English)

This work focuses on sampling strategies of configuration variations for generating robust universal locomotion policies for quadrupedal robots. We investigate the effects of sampling physical robot parameters and joint proportional-derivative gains to enable training a single reinforcement learning policy that generalizes to multiple parameter configurations. Three fundamental joint gain sampling strategies are compared: parameter sampling with (1) linear and polynomial function mappings of mass-to-gains, (2) performance-based adaptive filtering, and (3) uniform random sampling. We improve the robustness of the policy by biasing the configurations using nominal priors and reference models. All training was conducted using the RaiSim simulation environment, tested in simulation on a range of diverse quadrupeds, and zero-shot deployed onto hardware using the ANYmal quadruped robot. Compared to multiple baseline implementations, our results demonstrate the need for significant joint controller gains randomization for robust closing of the sim-to-real gap.

四足机器人强化学习仿真实战鲁棒控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。