arXiv:2606.23680cs.ROcs.AI2026-06被引 5

让高自由度人形机器人边走边灵巧操作,突破传统停顿式操控限制。

CoorDex: Coordinating Body and Hand Priors for Continuous Dexterous Humanoid Loco-Manipulation

论文配图:CoorDex: Coordinating Body and Hand Priors for Continuous Dexterous Humanoid Loco-Manipulation
图 1 · 摘自论文原文
  • 用身体与手部先验生成协同潜空间,实现运动与操作的联合控制
  • 在相同奖励预算下,传统方法失败,本方案实现连续瓶抓、开门等操作
  • 适合研究高自由度机器人灵巧操作与强化学习策略设计的学者

人形机器人通常采用停顿式操作:走到物体前停下,完成操作后再继续行走。同时普遍依赖低自由度末端执行器,仅能实现开闭式抓取。本文提出CoorDex,一种将高维全身与灵巧手控制转化为协同潜空间残差控制的学习框架,实现移动中的高自由度灵巧操作。基于全身体与手部模拟演示,训练出躯干与灵巧手的高级运动追踪教师模型,将其知识蒸馏为感知条件下的潜先验,并作为下游残差强化学习的动作空间。协调的潜空间残差策略通过共享任务上下文和独立的躯干-手部残差头,保留自然全身运动的同时提升指尖接触可靠性。CoorDex使配备20自由度WUJI手的Unitree G1机器人在行进中完成非中断瓶抓与携带、移动中开门及方块抓转等任务。消融实验表明,在相同奖励预算下,关节空间PPO、关节空间手控及单体潜空间预测均失败,而潜先验接口与协调残差结构使高维接触密集型操作可训练。

原文摘要 · Abstract (English)

Humanoid loco-manipulation is often simplified into a stop-and-go process: walking to an object, stopping to manipulate it, and then resuming locomotion. It also commonly relies on low degree-of-freedom (DoF) end effectors that behave like an open-close grasp primitive. We introduce CoorDex, a learning pipeline that converts high-dimensional body and dexterous hand control into coordinated latent residual control, enabling high-DoF dexterous loco-manipulation on the move. Starting from simulated whole-body and hand demonstrations, CoorDex trains privileged motion tracking teachers for the humanoid body and dexterous hand, distills them into proprioception-conditioned latent priors, and uses the frozen priors as the action space for downstream residual reinforcement learning. A coordinated latent residual policy composes these priors through shared task context and separate body-hand residual heads, preserving natural whole-body motion while improving finger-level contact reliability. CoorDex enables a Unitree G1 humanoid with a 20-DoF WUJI hand to execute dexterous manipulation while in motion, including non-stop bottle grasping and carrying, fridge door opening on the move, and cube pick-and-turn. Ablations on the walk-grasp-carry task show that joint-space PPO, joint-space hand control, and monolithic latent prediction all fail under the same reward budget, while the latent-prior interface and coordinated residual structure make high-dimensional contact-rich loco-manipulation trainable. Project Page: https://skevinci.github.io/coordex/

人形机器人灵巧操作强化学习协同控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。