arXiv:2503.20208cs.ROcs.AI2025-03被引 5

仅需一次示范,机器人就能学会自适应抓取。

Learning Adaptive Dexterous Grasping from Single Demonstrations

  • 从单次人类示范中学习多种抓取技能,通过轨迹奖励提升样本效率。
  • 使用视觉语言模型根据指令选择最适抓取策略,实现零样本跨域迁移。
  • 在真实机械手上达成90%成功率,显著优于基线方法。

如何让机器人高效学习灵巧抓取技能,并根据用户指令自适应应用?本文提出AdaDexGrasp框架,从每种技能的单一人类示范中学习抓取库,并利用视觉语言模型(VLM)根据用户指令选择最合适的技能。为提高样本效率,设计了轨迹跟随奖励,引导强化学习(RL)向人类示范状态靠近,同时保持探索灵活性。为突破单次示范限制,采用课程学习,逐步增加物体姿态变化以增强鲁棒性。部署时,VLM基于用户指令检索匹配技能,实现低层技能与高层意图的衔接。在仿真和真实场景中验证表明,该方法显著提升RL效率,可学习类人抓取策略。最终在真实PSYONIC Ability手上实现零样本迁移,对各类物体抓取成功率达90%,远超基线。

原文摘要 · Abstract (English)

How can robots learn dexterous grasping skills efficiently and apply them adaptively based on user instructions? This work tackles two key challenges: efficient skill acquisition from limited human demonstrations and context-driven skill selection. We introduce AdaDexGrasp, a framework that learns a library of grasping skills from a single human demonstration per skill and selects the most suitable one using a vision-language model (VLM). To improve sample efficiency, we propose a trajectory following reward that guides reinforcement learning (RL) toward states close to a human demonstration while allowing flexibility in exploration. To learn beyond the single demonstration, we employ curriculum learning, progressively increasing object pose variations to enhance robustness. At deployment, a VLM retrieves the appropriate skill based on user instructions, bridging low-level learned skills with high-level intent. We evaluate AdaDexGrasp in both simulation and real-world settings, showing that our approach significantly improves RL efficiency and enables learning human-like grasp strategies across varied object configurations. Finally, we demonstrate zero-shot transfer of our learned policies to a real-world PSYONIC Ability Hand, with a 90% success rate across objects, significantly outperforming the baseline.

灵巧抓取少样本学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。