arXiv:2602.22862cs.ROcs.CV2026-02中稿 · CVPR被引 1

用潜在扩散模型提升抓取策略的精度与泛化能力。

GraspLDP: Towards Generalizable Grasping Policy via Latent Diffusion

  • 在扩散模型中引入抓取姿态先验,引导动作生成更符合可行抓取配置。
  • 通过自监督重建任务学习抓取度特征,在反向扩散步骤中重构相机图像。
  • 在仿真和真实机器人上均表现优越,适合动态抓取场景应用。

本文致力于提升基于模仿学习的机器人操作策略在抓取任务中的精度与泛化能力。近年来,基于扩散模型的策略学习已成为机器人操作任务的主流方法。由于抓取是操作中的关键子任务,模仿学习策略在执行精确且泛化的抓取时尤为关键。现有抓取模仿学习方法常面临抓取不精准、空间泛化能力有限及对象泛化性差等问题。为此,我们将在扩散策略框架中融入抓取先验知识。具体而言,采用潜在扩散策略,利用抓取姿态先验指导动作块解码,确保生成的动作轨迹紧密贴合可行抓取构型。此外,在扩散过程中引入自监督重建目标:在每个反向扩散步骤中,从中间表示中反投影抓取度,并重建腕部相机图像。仿真与真实机器人实验表明,该方法显著优于基线方法,展现出强大的动态抓取能力。

原文摘要 · Abstract (English)

This paper focuses on enhancing the grasping precision and generalization of manipulation policies learned via imitation learning. Diffusion-based policy learning methods have recently become the mainstream approach for robotic manipulation tasks. As grasping is a critical subtask in manipulation, the ability of imitation-learned policies to execute precise and generalizable grasps merits particular attention. Existing imitation learning techniques for grasping often suffer from imprecise grasp executions, limited spatial generalization, and poor object generalization. To address these challenges, we incorporate grasp prior knowledge into the diffusion policy framework. In particular, we employ a latent diffusion policy to guide action chunk decoding with grasp pose prior, ensuring that generated motion trajectories adhere closely to feasible grasp configurations. Furthermore, we introduce a self-supervised reconstruction objective during diffusion to embed the graspness prior: at each reverse diffusion step, we reconstruct wrist-camera images back-projected the graspness from the intermediate representations. Both simulation and real robot experiments demonstrate that our approach significantly outperforms baseline methods and exhibits strong dynamic grasping capabilities.

抓取策略扩散模型模仿学习机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。