让机器人学会精准穿鞋带,成功率83.3%。
GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- 用强化学习过滤并增强人类示范数据,提升动作质量。
- 在穿鞋带任务中实现83.3%成功率,需毫米级精度与长时序推理。
- 适合追求高精度、长时程机械操作的机器人研究者。
我们提出GR-RL,一种将通用视觉-语言-动作(VLA)策略转化为高度专业化的长时程灵巧操作专家的框架。现有VLA策略假设人类示范最优,但在高度灵巧的任务中,人类示范往往存在噪声且非最优。GR-RL采用多阶段训练流程,通过强化学习对示范数据进行筛选、增强与强化。首先,基于离线强化学习与稀疏奖励,学习视觉-语言条件下的任务进展函数,利用所得$Q$值作为鲁棒的进展评估指标,筛选出对任务进展有正向贡献的动作轨迹。其次,引入形态对称增强策略,显著提升模型泛化能力与性能。最后,为实现高精度控制,通过在线强化学习学习隐空间噪声预测器,使VLA策略与实际执行行为对齐。该框架在穿鞋带任务中实现了83.3%的成功率,需完成多眼孔穿绳、毫秒级精度与柔体交互,是首个基于学习的自主穿鞋带系统。我们希望此工作推动通用机器人基础模型向真实世界专家演进。
原文摘要 · Abstract (English)
We present GR-RL, a robotic learning framework that turns a generalist vision-language-action (VLA) policy into a highly capable specialist for long-horizon dexterous manipulation. Assuming the optimality of human demonstrations is core to existing VLA policies. However, we claim that in highly dexterous and precise manipulation tasks, human demonstrations are noisy and suboptimal. GR-RL proposes a multi-stage training pipeline that filters, augments, and reinforces the demonstrations by reinforcement learning. First, GR-RL learns a vision-language-conditioned task progress, filters the demonstration trajectories, and only keeps the transitions that contribute positively to the progress. Specifically, we show that by directly applying offline RL with sparse reward, the resulting $Q$-values can be treated as a robust progress function. Next, we introduce morphological symmetry augmentation that greatly improves the generalization and performance of GR-RL. Lastly, to better align the VLA policy with its deployment behaviors for high-precision control, we perform online RL by learning a latent space noise predictor. With this pipeline, GR-RL is, to our knowledge, the first learning-based policy that can autonomously lace up a shoe by threading shoelaces through multiple eyelets with an 83.3% success rate, a task requiring long-horizon reasoning, millimeter-level precision, and compliant soft-body interaction. We hope GR-RL provides a step toward enabling generalist robot foundation models to specialize into reliable real-world experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。