arXiv:2510.12724cs.RO2025-10被引 12

用图扩散模型实现多机械手灵巧抓取,速度快且通用性强。

T(R,O) Grasp: Efficient Graph Diffusion of Robot-Object Spatial Transformation for Cross-Embodiment Dexterous Grasping

  • 构建机器人-物体空间变换图,统一建模抓取关系与几何特征。
  • 平均成功率94.83%,单次推理仅0.21秒,每秒生成41次抓取。
  • 适合需要高速闭环控制的灵巧操作场景,内存占用低。

灵巧抓取因状态与动作空间高维复杂仍是机器人领域核心挑战。我们提出T(R,O) Grasp,一种基于扩散模型的框架,可高效生成跨多种机械手的准确且多样化的抓取策略。其核心是T(R,O)图,统一建模机器人与物体间的空间变换关系,并编码其几何属性。结合高效的逆运动学求解器,该框架支持无条件与有条件抓取生成。在多样化灵巧手上的大量实验表明,T(R,O) Grasp在NVIDIA A100 40GB GPU上达到平均成功率为94.83%,单次推理耗时0.21秒,吞吐量达每秒41次抓取,显著优于现有基线方法。此外,该方法在不同机械手间具有强鲁棒性与泛化能力,内存消耗大幅降低。更重要的是,其高推理速度使闭环灵巧操作成为可能,凸显其作为灵巧抓取基础模型的潜力。

原文摘要 · Abstract (English)

Dexterous grasping remains a central challenge in robotics due to the complexity of its high-dimensional state and action space. We introduce T(R,O) Grasp, a diffusion-based framework that efficiently generates accurate and diverse grasps across multiple robotic hands. At its core is the T(R,O) Graph, a unified representation that models spatial transformations between robotic hands and objects while encoding their geometric properties. A graph diffusion model, coupled with an efficient inverse kinematics solver, supports both unconditioned and conditioned grasp synthesis. Extensive experiments on a diverse set of dexterous hands show that T(R,O) Grasp achieves average success rate of 94.83%, inference speed of 0.21s, and throughput of 41 grasps per second on an NVIDIA A100 40GB GPU, substantially outperforming existing baselines. In addition, our approach is robust and generalizable across embodiments while significantly reducing memory consumption. More importantly, the high inference speed enables closed-loop dexterous manipulation, underscoring the potential of T(R,O) Grasp to scale into a foundation model for dexterous grasping.

灵巧抓取扩散模型图神经网络机器人操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。