arXiv:2510.27420cs.RO2025-10被引 1

用几何推理实现跨机械臂抓取,数据效率高且推理快。

Towards a Multi-Embodied Grasping Agent

  • 基于流模型与等变架构,仅凭机械臂和场景几何推导抓取动作。
  • 支持2000+种不同自由度的夹爪,25,000场景下生成2000万次抓取。
  • 全JAX实现,可批量处理多场景、多夹爪,训练更快、推理更高效。

多具身抓取旨在开发能在多种夹爪设计间泛化的通用抓取方法。现有方法通常隐式学习机器人运动学结构,受限于大规模数据获取困难。本文提出一种数据高效的流模型等变抓取合成架构,可处理不同自由度的夹爪类型,并仅从夹爪与场景几何中推导全部必要信息。与以往等变抓取方法不同,我们从零开始在JAX中实现所有模块,支持对场景、夹爪和抓取动作的批量处理,带来更平滑的学习过程、更高的性能表现和更快的推理速度。数据集涵盖从人形手到平行旋转夹爪的各类夹爪,包含25,000个场景和2000万次抓取。

原文摘要 · Abstract (English)

Multi-embodiment grasping focuses on developing approaches that exhibit generalist behavior across diverse gripper designs. Existing methods often learn the kinematic structure of the robot implicitly and face challenges due to the difficulty of sourcing the required large-scale data. In this work, we present a data-efficient, flow-based, equivariant grasp synthesis architecture that can handle different gripper types with variable degrees of freedom and successfully exploit the underlying kinematic model, deducing all necessary information solely from the gripper and scene geometry. Unlike previous equivariant grasping methods, we translated all modules from the ground up to JAX and provide a model with batching capabilities over scenes, grippers, and grasps, resulting in smoother learning, improved performance and faster inference time. Our dataset encompasses grippers ranging from humanoid hands to parallel yaw grippers and includes 25,000 scenes and 20 million grasps.

抓取生成等变网络流模型多具身

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。