arXiv:2603.10306cs.RO2026-03

让仿人机器人稳稳端盘子,靠的是分步学习稳定与行走。

SteadyTray: Learning Object Balancing Tasks in Humanoid Tray Transport via Residual Reinforcement Learning

  • 分层强化学习:行走和托盘稳定分开训练,互不干扰。
  • 仿真中96.9%成功完成变速追踪,74.5%抗外力扰动。
  • 直接部署到G1机器人,无需额外训练,通用性强。

在非结构化环境中,稳定动态双足行走带来的负载振荡仍是仿人机器人的一大工程难题。为此,我们提出ReST-RL,一种分层强化学习架构,明确将行走与负载稳定解耦,并通过SteadyTray基准进行评估。不同于单一的端到端学习,该框架结合稳健的基础行走策略与动态残差模块,主动在末端执行器处抵消步态引起的扰动。这种架构分离确保了托盘运输的稳定性,同时不损害底层双足平衡性。仿真结果显示,残差设计显著优于端到端基线,在步态平滑性和姿态精度上表现更佳,实现96.9%的变速度追踪成功率和74.5%的外部扰动鲁棒性。成功部署于Unitree G1仿人机器人硬件,展现出对多种物体及外部扰动的高度可靠零样本模拟到现实泛化能力。

原文摘要 · Abstract (English)

Stabilizing unsecured payloads against the inherent oscillations of dynamic bipedal locomotion remains a critical engineering bottleneck for humanoids in unstructured environments. To solve this, we introduce ReST-RL, a hierarchical reinforcement learning architecture that explicitly decouples locomotion from payload stabilization, evaluated via the SteadyTray benchmark. Rather than relying on monolithic end-to-end learning, our framework integrates a robust base locomotion policy with a dynamic residual module engineered to actively cancel gait-induced perturbations at the end-effector. This architectural separation ensures steady tray transport without degrading the underlying bipedal stability. In simulation, the residual design significantly outperforms end-to-end baselines in gait smoothness and orientation accuracy, achieving a 96.9% success rate in variable velocity tracking and 74.5% robustness against external force disturbances. Successfully deployed on the Unitree G1 humanoid hardware, this modular approach demonstrates highly reliable zero-shot sim-to-real generalization across various objects and external force disturbances.

仿人机器人强化学习托盘搬运模拟到现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。