让人形机器人上半身任务空间轨迹异步跟踪更准更稳
Learning Asynchronous Upper-body Task-space Trajectory Tracking Policy for Humanoid Robots

- 用缓存未来轨迹+执行时间索引,训练异步跟踪策略
- 低更新率下轨迹漂移减少,硬件实验表现优于同步基线
- 适合需要低频规划却高精度执行的机器人控制场景
高层人形机器人规划器常输出稀疏、低频的任务空间轨迹,而全身控制器以高频运行,导致规划与执行在时间上不同步,且全身控制结构不完整。本文提出一种异步上半身任务空间轨迹跟踪框架。通过教师-学生蒸馏初始化学生策略,条件输入为完整的未来轨迹缓存和执行时间索引,并采用滑动窗口全局奖励进行训练,以减少帧漂移,无需显式帧估计。针对特定任务,通过模型预测控制(MPC)模块将稀疏参考补全为浮点基底和上半身引导,同时在动作和正向运动学层面引入自指导机制,抑制策略漂移。仿真与Unitree G1硬件实验表明,在低更新率下追踪性能更优,显著优于同步与解耦基线方法,并能更安全地适应分布外运动。
原文摘要 · Abstract (English)
High-level humanoid planners often output sparse task-space, low-rate trajectories, whereas whole-body controllers run at high frequency. This creates temporal asynchrony between the planning and execution, and structural incompleteness for full-body control. We propose an asynchronous upper body task-space tracking framework for humanoids. A student policy is initialized by teacher-student distillation, conditioned on the full cached future trajectory and an execution-time index, and trained with a sliding-window global reward to reduce frame drift without explicit frame estimation. For task-specific post-training, an MPC module completes sparse references into floating-base and upper-body guidance, while action- and FK level self-guidance constrain policy drift. Simulation and Unitree G1 hardware experiments show improved tracking under low update rates, stronger performance than synchronous and decoupled baselines, and safer adaptation to out-of-distribution motions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。