仅用一次人类示范,让机器人学会像人一样灵活操作工具。
A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation

- 将人体动作重定向为机器人可执行的参考轨迹,保留手物空间与接触关系。
- 在仿真中训练残差强化学习策略,精准追踪物体关键点,零样本迁移到真实硬件。
- 首次系统分析复杂接触场景下的仿真到现实迁移关键因素,适合机器人操控研究者。
近期在人形全身控制中取得成功的简单范式:将人体运动重定向为机器人运动参考,再通过强化学习(RL)训练策略进行跟踪。但该方法能否应用于灵巧操作尚不明确,因操作涉及复杂的、富含接触的动力学,需精细调控接触模式与力。本文提出REGRIND,一种极简的重定向引导式强化学习流程,仅需一次人类示范即可学习灵巧操作策略。REGRIND将人体手物运动重定向为保留手物空间与接触关系的机器人参考轨迹,在仿真中训练残差强化学习策略,以追踪物体中心关键点,并通过精细系统辨识实现零样本迁移至真实硬件。所生成策略在两种不同多指机械手上均表现出流畅、类人的行为,完成剪刀操作、螺丝刀旋转等高接触任务。通过系统的硬件实验,我们识别并分析了灵巧操作中仿真到现实迁移的关键因素,为基于重定向的学习提供了实用指导。视频与代码见https://yunhaifeng.com/REGRIND。
原文摘要 · Abstract (English)
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manipulation policies from a single human demonstration. REGRIND retargets human hand-object motion to a robot reference that preserves hand-object spatial and contact relationships, trains a residual RL policy in simulation to track object-centric keypoints along that reference, and transfers the resulting policy zero-shot to hardware with careful system identification. The resulting policies produce fluid, human-like behavior on two different multi-fingered hands across contact-rich tool-use tasks, including operating a pair of scissors and turning a screwdriver. Through systematic hardware experiments, we identify and analyze the key factors that govern sim-to-real transfer in dexterous manipulation, offering practical guidance for retargeting-based learning in contact-rich settings. Videos and code are available at https://yunhaifeng.com/REGRIND.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。