arXiv:2509.22149cs.RO2025-09被引 12

仅凭一次抓取示范,就能让机械手通用抓取新物体。

DemoGrasp: Universal Dexterous Grasping from a Single Demonstration

  • 通过编辑示范轨迹的腕部姿态和手指角度实现泛化抓取
  • 仿真中对300+物体成功率达95%,真实世界抓110种新物体
  • 支持视觉输入、语言指令,适应光照/背景变化,适合实际部署

多指灵巧手的通用抓取是机器人操作的核心挑战。现有方法虽借助强化学习(RL)学习闭环抓取策略,但高维、长时程探索导致奖励与课程设计复杂,常难以在多样物体上获得最优解。本文提出DemoGrasp,从单个特定物体的成功抓取示范轨迹出发,通过修改轨迹中的机器人动作实现对新物体和姿态的适应:调整腕部姿态决定抓取位置,改变手指关节角度决定抓取方式。将此轨迹编辑建模为单步马尔可夫决策过程(MDP),利用RL并行优化跨数百物体的通用策略,仅使用二值成功奖励与机器人-桌面碰撞惩罚。仿真中,使用Shadow Hand在DexGraspNet数据集上达到95%成功率,优于现有最先进方法;在六组未见物体数据集上,对六种不同灵巧手形态平均成功率达84.6%,训练仅用175个物体。通过视觉模仿学习,该策略在真实世界成功抓取110种未见过的物体,包括细小薄片类物品,具备空间、背景、光照变化鲁棒性,支持RGB与深度输入,并可扩展至杂乱场景下的语言引导抓取。

原文摘要 · Abstract (English)

Universal grasping with multi-fingered dexterous hands is a fundamental challenge in robotic manipulation. While recent approaches successfully learn closed-loop grasping policies using reinforcement learning (RL), the inherent difficulty of high-dimensional, long-horizon exploration necessitates complex reward and curriculum design, often resulting in suboptimal solutions across diverse objects. We propose DemoGrasp, a simple yet effective method for learning universal dexterous grasping. We start from a single successful demonstration trajectory of grasping a specific object and adapt to novel objects and poses by editing the robot actions in this trajectory: changing the wrist pose determines where to grasp, and changing the hand joint angles determines how to grasp. We formulate this trajectory editing as a single-step Markov Decision Process (MDP) and use RL to optimize a universal policy across hundreds of objects in parallel in simulation, with a simple reward consisting of a binary success term and a robot-table collision penalty. In simulation, DemoGrasp achieves a 95% success rate on DexGraspNet objects using the Shadow Hand, outperforming previous state-of-the-art methods. It also shows strong transferability, achieving an average success rate of 84.6% across diverse dexterous hand embodiments on six unseen object datasets, while being trained on only 175 objects. Through vision-based imitation learning, our policy successfully grasps 110 unseen real-world objects, including small, thin items. It generalizes to spatial, background, and lighting changes, supports both RGB and depth inputs, and extends to language-guided grasping in cluttered scenes.

灵巧抓取视觉模仿泛化能力真实世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。