用可组合的操控策略,让机器人更稳定地完成复杂人机交互任务。
Coordinated Humanoid Manipulation with Choice Policies
- 将人形机器人控制拆解为手眼协同、抓取等模块,提升数据采集效率。
- 提出多候选动作评分机制,实现快速推理与多模态行为建模。
- 在洗碗和擦白板任务中表现优于扩散模型和传统模仿学习,适合长时序任务。
人形机器人在人类环境中的应用前景广阔,但实现头部、双手与双腿间的鲁棒全身协调仍是一大挑战。本文提出一种结合模块化遥操作界面与可扩展学习框架的系统。遥操作设计将人形机器人控制分解为直观的子模块,包括手眼协同、抓取原语、手臂末端跟踪和行走。这种模块化结构使高质量示范数据的采集更加高效。在此基础上,我们引入选择策略(Choice Policy),一种模仿学习方法,能够生成多个候选动作并学习对其评分,该架构支持快速推理并有效建模多模态行为。我们在两个真实任务上验证了该方法:洗碗装载和全身心体协同白板擦拭。实验表明,选择策略显著优于扩散策略和标准行为克隆。此外,结果表明手眼协同对长时程任务的成功至关重要。本工作为非结构化环境中人形机器人协调操作的可扩展数据收集与学习提供了实用路径。
原文摘要 · Abstract (English)
Humanoid robots hold great promise for operating in human-centric environments, yet achieving robust whole-body coordination across the head, hands, and legs remains a major challenge. We present a system that combines a modular teleoperation interface with a scalable learning framework to address this problem. Our teleoperation design decomposes humanoid control into intuitive submodules, which include hand-eye coordination, grasp primitives, arm end-effector tracking, and locomotion. This modularity allows us to collect high-quality demonstrations efficiently. Building on this, we introduce Choice Policy, an imitation learning approach that generates multiple candidate actions and learns to score them. This architecture enables both fast inference and effective modeling of multimodal behaviors. We validate our approach on two real-world tasks: dishwasher loading and whole-body loco-manipulation for whiteboard wiping. Experiments show that Choice Policy significantly outperforms diffusion policies and standard behavior cloning. Furthermore, our results indicate that hand-eye coordination is critical for success in long-horizon tasks. Our work demonstrates a practical path toward scalable data collection and learning for coordinated humanoid manipulation in unstructured environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。