arXiv:2507.13097cs.ROcs.AI2025-07被引 47

用扩散模型生成6自由度抓取,提升机器人跨设备泛化能力

GraspGen: A Diffusion-based Framework for 6-DOF Grasping with On-Generator Training

  • 基于扩散过程建模抓取生成,结合高效判别器筛选抓取姿态
  • 在5300万条模拟抓取数据上训练,实测性能优于现有方法
  • 适用于多种夹爪和真实场景,尤其适合视觉噪声环境

抓取是机器人的一项基础技能,尽管研究进展显著,但现有的基于学习的6-DOF抓取方法仍难以即插即用,在不同机器人形态和真实环境中泛化能力不足。我们借鉴物体中心抓取生成过程作为迭代扩散过程的成功经验,提出GraspGen框架:采用DiffusionTransformer架构增强抓取生成,并配备高效判别器对采样抓取进行评分与过滤。引入一种新颖且高效的判别器在线生成训练策略。为实现对物体和夹爪的可扩展性,我们发布了包含超过5300万条抓取的新型仿真数据集。实验表明,GraspGen在单个物体上的仿真测试中优于先前方法,于FetchBench抓取基准上达到当前最优性能,并在具备噪声视觉观测的真实机器人上表现良好。

原文摘要 · Abstract (English)

Grasping is a fundamental robot skill, yet despite significant research advancements, learning-based 6-DOF grasping approaches are still not turnkey and struggle to generalize across different embodiments and in-the-wild settings. We build upon the recent success on modeling the object-centric grasp generation process as an iterative diffusion process. Our proposed framework, GraspGen, consists of a DiffusionTransformer architecture that enhances grasp generation, paired with an efficient discriminator to score and filter sampled grasps. We introduce a novel and performant on-generator training recipe for the discriminator. To scale GraspGen to both objects and grippers, we release a new simulated dataset consisting of over 53 million grasps. We demonstrate that GraspGen outperforms prior methods in simulations with singulated objects across different grippers, achieves state-of-the-art performance on the FetchBench grasping benchmark, and performs well on a real robot with noisy visual observations.

6-DOF抓取扩散模型机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。