arXiv:2511.09558cs.ROcs.AI2025-11被引 1

用互联网级视觉模型+仿真生成力闭合抓取,实现无需人工数据的精准抓握。

IFG: Internet-Scale Guidance for Functional Grasping Generation

  • 结合大模型语义理解与仿真生成力闭合抓取策略
  • 在真实相机点云上实时运行,抓取成功率超90%
  • 无需人工标注数据,适合复杂场景机器人抓取

在互联网规模数据上训练的大视觉模型展现出强大的物体部件分割与语义理解能力,即使在杂乱拥挤场景中亦然。然而,这类模型仅能引导机械臂定位到物体大致区域,缺乏精确控制灵巧手进行3D抓取所需的几何理解。为此,我们提出利用仿真环境中的力闭合抓取生成流程,该流程可理解手与物体局部几何关系。由于该流程耗时且需真值观测,我们将其生成的数据蒸馏为一个扩散模型,可在真实相机点云上实时运行。通过融合互联网级模型的全局语义理解与基于仿真的局部感知力闭合机制,本方法在无任何人工收集训练数据的情况下,实现了高性能的语义抓取。可视化效果请访问 https://ifgrasping.github.io/

原文摘要 · Abstract (English)

Large Vision Models trained on internet-scale data have demonstrated strong capabilities in segmenting and semantically understanding object parts, even in cluttered, crowded scenes. However, while these models can direct a robot toward the general region of an object, they lack the geometric understanding required to precisely control dexterous robotic hands for 3D grasping. To overcome this, our key insight is to leverage simulation with a force-closure grasping generation pipeline that understands local geometries of the hand and object in the scene. Because this pipeline is slow and requires ground-truth observations, the resulting data is distilled into a diffusion model that operates in real-time on camera point clouds. By combining the global semantic understanding of internet-scale models with the geometric precision of a simulation-based locally-aware force-closure, \our achieves high-performance semantic grasping without any manually collected training data. For visualizations of this please visit our website at https://ifgrasping.github.io/

机器人抓取扩散模型力闭合语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。