arXiv:2506.02489cs.ROcs.CV2025-06NeurIPS被引 3

让不同机械手学会用视觉模仿抓取,无需配对数据或仿真。

Grasp2Grasp: Vision-Based Dexterous Grasp Translation via Schrödinger Bridges

  • 用量子桥接方法建模抓取分布间的随机传输。
  • 在多种手-物组合上生成稳定且物理合理的等效抓取。
  • 适合需要跨机械手抓取迁移的机器人研究者使用。

我们提出一种基于视觉的灵巧抓取迁移新方法,旨在将抓取意图从一个机械手形态迁移到另一个不同形态的机械手。给定源机械手抓取物体的视觉观测,目标是合成目标机械手的功能等效抓取,无需成对演示或手特异性仿真。我们将该问题建模为基于薛定谔桥形式的抓取分布间随机传输,通过条件于视觉观测的得分匹配与流匹配学习源与目标潜空间之间的映射。为引导此迁移,引入融合物理信息的成本函数,编码基姿态、接触图、力矩空间和可操作性的对齐。在多样化的手-物组合上实验表明,该方法能生成稳定且物理合理的抓取,具备强泛化能力。本工作实现了异构操作器间的语义抓取迁移,并连接了视觉抓取与概率生成建模。更多信息见 https://grasp2grasp.github.io/

原文摘要 · Abstract (English)

We propose a new approach to vision-based dexterous grasp translation, which aims to transfer grasp intent across robotic hands with differing morphologies. Given a visual observation of a source hand grasping an object, our goal is to synthesize a functionally equivalent grasp for a target hand without requiring paired demonstrations or hand-specific simulations. We frame this problem as a stochastic transport between grasp distributions using the Schrödinger Bridge formalism. Our method learns to map between source and target latent grasp spaces via score and flow matching, conditioned on visual observations. To guide this translation, we introduce physics-informed cost functions that encode alignment in base pose, contact maps, wrench space, and manipulability. Experiments across diverse hand-object pairs demonstrate our approach generates stable, physically grounded grasps with strong generalization. This work enables semantic grasp transfer for heterogeneous manipulators and bridges vision-based grasping with probabilistic generative modeling. Additional details at https://grasp2grasp.github.io/

抓取迁移视觉感知生成模型机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。