arXiv:2507.02747cs.CVcs.RO2025-07ICCV被引 36

用17000万数据训练大模型,让机械手按语言指令精准抓取物体

DexVLG: Dexterous Vision-Language-Grasp Model at Scale

  • 基于170万抓取姿态数据,用视觉语言模型预测符合指令的抓取姿势
  • 零样本测试中执行成功率超76%,在仿真和真实场景均实现精准部位抓取
  • 适合需要高精度人形机械手操作的智能机器人研究与应用

随着大模型发展,视觉-语言-动作系统使机器人能完成更复杂任务。但受限于数据采集难度,现有研究多集中于简单夹持器控制,对类人灵巧手的功能性抓取研究较少。本文提出DexVLG,一个面向灵巧手抓取姿态预测的大规模视觉-语言-抓取模型,支持单视角RGBD输入并响应语言指令。为此,我们构建了包含17000万抓取姿态、覆盖17.4万个物体的仿真数据集DexGraspNet 3.0,每件物体配有部件级语义描述。利用该数据集训练视觉语言模型与基于流匹配的姿态预测头,可生成与指令对齐的抓取姿势。通过物理仿真基准和真实实验评估,DexVLG展现出强大零样本泛化能力:仿真中零样本执行成功率超过76%,达到当前最优部件级抓取准确率;真实世界中亦成功实现物体部件对齐抓取。

原文摘要 · Abstract (English)

As large models gain traction, vision-language-action (VLA) systems are enabling robots to tackle increasingly complex tasks. However, limited by the difficulty of data collection, progress has mainly focused on controlling simple gripper end-effectors. There is little research on functional grasping with large models for human-like dexterous hands. In this paper, we introduce DexVLG, a large Vision-Language-Grasp model for Dexterous grasp pose prediction aligned with language instructions using single-view RGBD input. To accomplish this, we generate a dataset of 170 million dexterous grasp poses mapped to semantic parts across 174,000 objects in simulation, paired with detailed part-level captions. This large-scale dataset, named DexGraspNet 3.0, is used to train a VLM and flow-matching-based pose head capable of producing instruction-aligned grasp poses for tabletop objects. To assess DexVLG's performance, we create benchmarks in physics-based simulations and conduct real-world experiments. Extensive testing demonstrates DexVLG's strong zero-shot generalization capabilities-achieving over 76% zero-shot execution success rate and state-of-the-art part-grasp accuracy in simulation-and successful part-aligned grasps on physical objects in real-world scenarios.

灵巧手视觉语言抓取预测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。