arXiv:2606.00998cs.RO2026-06被引 1

让机器人抓取模型跨肢体通用,适配新夹爪和新物体。

GraspGen-X: Cross-Embodiment 6-DOF Diffusion-based Grasping

论文配图:GraspGen-X: Cross-Embodiment 6-DOF Diffusion-based Grasping
图 1 · 摘自论文原文
  • 用路径体积法编码夹爪,结合扩散模型生成6自由度抓取动作。
  • 在20亿次模拟抓取数据上训练,零样本泛化能力优于基线方法。
  • 适合需要快速适配新机械臂的工业场景,尤其擅长真实夹爪迁移。

本文研究跨肢体6自由度机器人抓取问题。与以往工作不同,要求模型不仅泛化到新物体/场景,还需适应新夹爪形态与物理抓取过程。提出将扩散模型生成式6自由度抓取方法扩展为可条件于夹爪表示的新范式,设计基于扫掠体积的夹爪编码方法,并在包含20亿次抓取的大规模程序化夹爪数据集上进行训练。仿真实验表明,该模型在零样本条件下对新型真实夹爪与物体的泛化性能优于基线方法;同时作为微调初始化也表现良好。消融实验验证了扫掠体积编码与程序化夹爪数据集的有效性。最后,在真实世界中实现对新型夹爪的零样本6自由度抓取,跨肢体泛化能力超越现有方法。

原文摘要 · Abstract (English)

We study cross-embodiment 6-DOF robot grasping. Unlike prior works, we require the model not only to generalize to novel objects / scenes but also to novel gripper morphologies and physical grasping processes. Our method extends diffusion model based generative 6-DOF grasping models to condition on the additional gripper's representation. We propose a swept-volume heuristic for encoding the gripper. We train our cross-embodiment model with procedural grippers and a large-scale dataset of 2 Billion grasps. In simulation experiments, our model has the best zero-shot generalization to novel real-world grippers and objects over baseline methods. Our model also serves as a good initialization for fine-tuning to adapt to novel grippers. In ablations, we demonstrate the efficiency of our sweep-volume gripper representation and our procedural gripper training dataset. Last, we show zero-shot generalization to real-world novel grippers for 6-DOF grasping, surpassing baselines in cross-embodiment generalization.

机器人抓取扩散模型跨肢体泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。